Complete Revision & Study Guide
A topic-by-topic guide covering dimension reduction, clustering, regression, regularization,
classification, support vector machines, tree-based models, and neural networks.
Original study notes — independently written summary and explanation
,Contents
• 1. Foundations of Machine Learning
• 2. Dimension Reduction: PCA and Kernel PCA
• 3. Dissimilarity and Clustering
• 4. Regression and Model Evaluation
• 5. Nonlinear Regression and Regularization
• 6. Logistic Regression and Classification Evaluation
• 7. Support Vector Machines
• 8. Tree-Based Models and Ensembles
• 9. Neural Networks
, 1. Foundations of Machine Learning
1.1 What is machine learning?
Machine Learning (ML) aims to automatically detect patterns in data in order to predict future
outcomes of interest and support decision-making.
Any statements produced by a trained model are only applicable to the population actually
represented by the training data — a model should never be assumed to generalise beyond
that population without justification. This matters directly for bias: if a model is trained on
biased data, it will learn and perpetuate those biases. Collecting data and designing sample
surveys appropriately is therefore critical to producing sensible, defensible analysis
downstream.
Before applying any ML technique, it is standard practice to explore the data first —
exploratory data analysis (EDA). Summaries and charts should be tailored to the type of
data being examined (numerical vs categorical).
1.2 Data pre-processing
Data cleaning may include handling:
• Duplicates.
• Inconsistencies (typos, differing labels for the same category).
• Outliers (which may need to be removed, but carefully — not automatically).
• Missing values (either dropping the affected observations, or dropping the affected
feature).
• Variable types that need adjusting.
Beyond cleaning, pre-processing also covers data integration (combining data from multiple
sources), data reduction (reducing volume or dimensionality), and data partitioning (splitting
into training, validation, and test sets).
1.3 Feature engineering
Feature engineering is the process of extracting real-valued features of a common
dimension from raw inputs. A key risk to keep in mind throughout is the curse of
dimensionality: adding more features tends to worsen model performance unless those
features are genuinely relevant.
A dataset is typically represented as a design matrix X, with N rows (observations) and D
columns (features), where each row xₙ ∈ ℝᴰ.
Encoding categorical variables
• One-hot encoding (dummy expansion) — transforms a categorical variable into
numerical indicator variables. In a linear regression with an intercept, one category
must be fixed as the reference category, so only C−1 dummy variables are introduced
for a variable with C categories — including all C would make the design matrix rank-
deficient (the columns become linearly dependent, so infinitely many weight vectors
would fit equally well).
• Ordinal encoding — replaces ordered categories with their rank, e.g. mapping C
ordered categories to the values (c − 0.5)/C for c = 1, ..., C, which spreads the
categories evenly across the interval (0, 1).
• Target encoding — replaces each category with the mean of the target variable within
that category; useful when there are many unordered categories. Care is needed to