USAAIO Lesson 69, from Week 24 of Phase 4. It builds kNN majority-vote classification from scratch, covers the L2, L1, and cosine distance metrics, and works the curse of dimensionality, where the ratio of the maximum to the minimum distance converges to 1. It then covers the KD-tree and ball-tree for efficient lookup, weighted kNN using 1/distance, and the bias-variance trade-off as k changes. Every number was verified with numpy and sklearn in June 2026. The lesson runs to 32 slides.
Subject: Machine Learning · 63 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 69 · Week 24
The simplest possible model — and the hardest possible caveat. Majority vote, distance metrics, KD-trees, and why kNN quietly breaks in high dimensions.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 69: kNN and the Curse of Dimensionality: without looking back, what was the main idea of AdaBoost, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
AdaBoost from scratch — sample-weight updates, the alpha formula, exponential-loss minimization, gradient descent in function space, and convergence theory. Implement decision-stump AdaBoost, verify it matches sklearn, and compare it to a single deep tree.
Section
Part 1 of 4
Concept
kNN stores the full training set. At prediction time it computes the distance from the query to every training point, takes the k closest, and returns the majority class.
k=1 is the most sensitive (nearest neighbor wins alone). Larger k smooths the decision boundary and reduces variance at the cost of bias (Lesson 32).
Intuition
Think of k=5 as asking five of your nearest neighbors to vote. If four say 'spam' and one says 'ham', you go with spam. The noise in any single neighbor averages out over the group.
k=1 is over-committed: one mislabeled outlier anywhere in the training set can flip a prediction. k=20 draws a smoother boundary but misses tight clusters.
Counterexample
Discussion prompt
Think of k=5 as asking five of your nearest neighbors to vote. If four say 'spam' and one says 'ham', you go with spam. The noise in any single neighbor averages out over the group.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Classify query (2.5, 2.5) using k=3 against 5 labeled training points. Predict the answer before computing.
Commit before you compute: what does kNN from scratch — 5-point trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Sort distances: indices [1,2,0,4,3]; k=3 neighbors are classes [0,1,0] -> majority class 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Two class-0 votes beat one class-1 vote: prediction is 0.
Worked example
Classify query (2.5, 2.5) using k=3 against 5 labeled training points. Predict the answer before computing.
import numpy as np
from collections import Counter
X_train = np.array([[1,1],[2,1],[3,4],[5,5],[1,4]], dtype=float)
y_train = np.array([0, 0, 1, 1, 1])
query = np.array([2.5, 2.5])
dists = np.sqrt(np.sum((X_train - query)**2, axis=1))
sorted_idx = np.argsort(dists)
k3 = sorted_idx[:3]
votes = Counter(y_train[k3])
print('dists:', dists.round(4).tolist())
print('sorted_idx:', sorted_idx.tolist())
print('k=3 neighbors:', k3.tolist(), 'classes:', y_train[k3].tolist())
print('prediction:', votes.most_common(1)[0][0])Sort distances: indices [1,2,0,4,3]; k=3 neighbors are classes [0,1,0] -> majority class 0
Why: Two class-0 votes beat one class-1 vote: prediction is 0. Note indices 1 and 2 are equidistant (1.5811) — tie-break by stable argsort order.
| idx | point | class | L2 dist |
|---|---|---|---|
| 0 | (1,1) | 0 | 2.1213 |
| 1 | (2,1) | 0 | 1.5811 |
| 2 | (3,4) | 1 | 1.5811 |
| 3 | (5,5) | 1 | 3.5355 |
| 4 | (1,4) | 1 | 2.1213 |
Comparison
Comparison matrix
From kNN from scratch — 5-point trace: refill the point column from what you know. The rest of the table is as it appeared.
| idx | point | class | L2 dist |
|---|---|---|---|
| 0 | (1,1) | 0 | 2.1213 |
| 1 | (2,1) | 0 | 1.5811 |
| 2 | (3,4) | 1 | 1.5811 |
| 3 | (5,5) | 1 | 3.5355 |
| 4 | (1,4) | 1 | 2.1213 |
Section
Part 2 of 4
Concept
| metric | formula | best for |
|---|---|---|
| L2 (Euclidean) | sqrt(sum (a-b)^2) | continuous, isotropic features |
| L1 (Manhattan) | sum |a-b| | sparse data, outlier-robust |
| Minkowski (p) | (sum |a-b|^p)^(1/p) | generalizes L1/L2 (p=1 or 2) |
| Cosine | 1 - dot(a,b)/(|a||b|) | text/embedding direction |
| Mahalanobis | sqrt((a-b)^T S^-1 (a-b)) | accounts for feature covariance |
For raw pixel or sensor features: L2 is the default. For high-dimensional count vectors (word counts, embeddings): cosine. For correlated features with different scales: Mahalanobis (or just StandardScaler + L2).
Trade off
Comparison matrix
From L2, L1, Minkowski, cosine, Mahalanobis: every row here is a choice with a cost. Fill the formula column, then say which row you would actually pick and what you give up for it.
| metric | formula | best for |
|---|---|---|
| L2 (Euclidean) | sqrt(sum (a-b)^2) | continuous, isotropic features |
| L1 (Manhattan) | sum |a-b| | sparse data, outlier-robust |
| Minkowski (p) | (sum |a-b|^p)^(1/p) | generalizes L1/L2 (p=1 or 2) |
| Cosine | 1 - dot(a,b)/(|a||b|) | text/embedding direction |
| Mahalanobis | sqrt((a-b)^T S^-1 (a-b)) | accounts for feature covariance |
Missing information
Discussion prompt
Compute L2, L1, and cosine distance from A=(1,2) to B=(4,6) and to C=(1,3). Which is closer under each metric?
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
B lies along nearly the same direction as A (both tilted northeast), so the angle is smaller. C is one unit directly above A, so L2/L1 are small but the direction shifts more.
Worked example
Compute L2, L1, and cosine distance from A=(1,2) to B=(4,6) and to C=(1,3). Which is closer under each metric?
import numpy as np
A = np.array([1.0, 2.0])
B = np.array([4.0, 6.0])
C = np.array([1.0, 3.0])
def cos_dist(u, v):
return 1 - np.dot(u, v) / (np.linalg.norm(u) * np.linalg.norm(v))
print('L2 A-B:', np.linalg.norm(A-B)) # 5.0
print('L2 A-C:', np.linalg.norm(A-C)) # 1.0
print('L1 A-B:', np.sum(np.abs(A-B))) # 7.0
print('L1 A-C:', np.sum(np.abs(A-C))) # 1.0
print('cos A-B:', round(cos_dist(A,B),4)) # 0.0077
print('cos A-C:', round(cos_dist(A,C),4)) # 0.0101C is closest under L2 (1.0) and L1 (1.0); B is closer under cosine (0.0077 vs 0.0101)
Why: B lies along nearly the same direction as A (both tilted northeast), so the angle is smaller. C is one unit directly above A, so L2/L1 are small but the direction shifts more.
| metric | dist(A,B) | dist(A,C) | nearest |
|---|---|---|---|
| L2 | 5.0000 | 1.0000 | C |
| L1 | 7.0000 | 1.0000 | C |
| cosine | 0.0077 | 0.0101 | B |
Pattern
Step through it
Step through Distance metric worked example one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Features: age in years (range 0-90) and income_normalized (range 0-1). Feed both directly to L2 kNN.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1.
Apply StandardScaler before kNN whenever features differ in scale.
Why: A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1. kNN effectively ignores income entirely.
Trap
Features: age in years (range 0-90) and income_normalized (range 0-1). Feed both directly to L2 kNN.
L2 distance is dominated by age alone
Why: A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1. kNN effectively ignores income entirely.
Apply StandardScaler before kNN whenever features differ in scale.
After StandardScaler, each feature contributes equally to L2 distance
Why: Z-score normalization gives every feature mean 0 and std 1. This is mandatory for distance-based models including kNN, SVM (L15), and clustering (L52).
Break the constraint
Discussion prompt
The rule this trap just fixed:
Z-score normalization gives every feature mean 0 and std 1. This is mandatory for distance-based models including kNN, SVM (L15), and clustering (L52).
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1. kNN effectively ignores income entirely.
Section
Part 3 of 4
Concept
As dimensionality grows, the ratio of the farthest to the nearest neighbor shrinks toward 1 — every point looks equally far from every other. The notion of 'nearest' becomes meaningless.
\[ \frac{d_{\max}(n,d)}{d_{\min}(n,d)} \xrightarrow{d \to \infty} 1 \]
This is because in high d, distance concentrates around the mean. Adding irrelevant dimensions drowns signal in noise.
Analogy
Discussion prompt
Explain In high d, all points become equidistant by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
As dimensionality grows, the ratio of the farthest to the nearest neighbor shrinks toward 1 — every point looks equally far from every other. The notion of 'nearest' becomes meaningless.
Estimation
Predict first
Draw 1000 standard-normal points in d dimensions, then measure max/min distance to a query. Watch the ratio collapse as d increases.
Commit before you compute: what does Curse of dimensionality — measured come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: d=2 ratio 60.4; d=10 ratio 3.6; d=50 ratio 1.8; d=100 ratio 1.4; d=500 ratio 1.2
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Even at d=500 the ratio is only 1.19 -- the nearest and farthest neighbor are 19% apart in distance.
Worked example
Draw 1000 standard-normal points in d dimensions, then measure max/min distance to a query. Watch the ratio collapse as d increases.
import numpy as np
rng = np.random.default_rng(0)
dims = [2, 10, 50, 100, 500]
print(f'dim | max_dist | min_dist | ratio')
for d in dims:
pts = rng.standard_normal((1000, d))
q = rng.standard_normal((1, d))
dists = np.sqrt(np.sum((pts - q)**2, axis=1))
print(f'{d:>3} | {dists.max():8.4f} | {dists.min():8.4f} | {dists.max()/dists.min():6.4f}')d=2 ratio 60.4; d=10 ratio 3.6; d=50 ratio 1.8; d=100 ratio 1.4; d=500 ratio 1.2
Why: Even at d=500 the ratio is only 1.19 -- the nearest and farthest neighbor are 19% apart in distance. kNN's ranking is still somewhat meaningful here but degrades fast with irrelevant features.
| dim | max dist | min dist | ratio |
|---|---|---|---|
| 2 | 4.4253 | 0.0733 | 60.36 |
| 10 | 7.1824 | 1.9777 | 3.63 |
| 50 | 12.6616 | 7.1912 | 1.76 |
| 100 | 16.5239 | 11.4392 | 1.44 |
| 500 | 34.2870 | 28.9619 | 1.18 |
Pattern
Step through it
Step through Curse of dimensionality — measured one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Raw image pixels (d=784 for MNIST), raw text TF-IDF vectors (d=10k+), and unfiltered sensor arrays all violate the low-d assumption kNN requires.
The KD-tree index is also exact only for d < 20. Above that it degrades to brute force — a second reason high-d kNN is slow.
Sorting
Sort into buckets
These are the pieces of Lesson 69: kNN and the Curse of Dimensionality, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 4
Concept
| method | exact? | complexity (query) | good for |
|---|---|---|---|
| brute force | yes | O(nd) | small n or large d |
| KD-tree | yes | O(log n) avg | d < 20, moderate n |
| ball tree | yes | O(log n) avg | d 20-100, curved manifolds |
| FAISS / HNSW | approx | sub-linear | millions of vectors, d > 100 |
sklearn selects the algorithm automatically (algorithm='auto') based on n and d. For production with 1M+ vectors and d=128 embeddings, FAISS is the standard (not in sklearn).
Discrimination
Sort into buckets
Sort these by exact?, from memory, without looking back at KD-tree, ball tree, and approximate NN. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Missing information
Discussion prompt
Compare predict-time for brute force, KD-tree, and ball tree on n=5000, d=10 classification (make_classification). Same accuracy expected.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
KD-tree is 26x faster than brute force at d=10. All three produce identical predictions — the index structure doesn't change the result, only how fast you find the neighbors.
Worked example
Compare predict-time for brute force, KD-tree, and ball tree on n=5000, d=10 classification (make_classification). Same accuracy expected.
import time
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=5000, n_features=10, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
for algo in ['brute', 'kd_tree', 'ball_tree']:
knn = KNeighborsClassifier(n_neighbors=5, algorithm=algo)
t0 = time.time()
knn.fit(X_tr, y_tr)
acc = accuracy_score(y_te, knn.predict(X_te))
print(f'{algo:12s}: acc={acc:.4f}, time={time.time()-t0:.3f}s')brute=2.4s, kd_tree=0.09s, ball_tree=0.13s; all accuracy 0.9670
Why: KD-tree is 26x faster than brute force at d=10. All three produce identical predictions — the index structure doesn't change the result, only how fast you find the neighbors.
| algorithm | accuracy | time (s) |
|---|---|---|
| brute | 0.9670 | 2.398 |
| kd_tree | 0.9670 | 0.109 |
| ball_tree | 0.9670 | 0.138 |
Invariant
Step through it
Step through KD-tree vs brute force — timing on 5000 points one row at a time. One of these columns never changes — find it, and say why it cannot.
Concept
Standard kNN gives every neighbor one vote. Weighted kNN gives each neighbor a vote proportional to 1/distance — a very close neighbor counts more than a distant one.
\[ \hat{y} = \arg\max_c \sum_{i \in N_k(x)} \frac{\mathbf{1}[y_i = c]}{d(x, x_i) + \epsilon} \]
This is the continuous generalization of majority vote: as k grows large, weighted kNN approaches a kernel density classifier with a box kernel.
Explain it
Discussion prompt
Explain Weighted kNN: 1/distance votes to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Standard kNN gives every neighbor one vote. Weighted kNN gives each neighbor a vote proportional to 1/distance — a very close neighbor counts more than a distant one.
Estimation
Predict first
Re-classify query (2.5, 2.5) with k=3 1/distance weights. The two class-0 neighbors are equidistant from the class-1 neighbor — so weighted vote changes nothing here, but the math is the point.
Commit before you compute: what does Weighted kNN trace — same 5-point example come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Weights [0.6325, 0.6325, 0.4714]; class-0 total 1.1039, class-1 total 0.6325 -> prediction 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Neighbors 1 and 2 are equidistant (1.5811) so equal weight; neighbor 0 is farther (2.1213) so lower weight 0.4714.
Worked example
Re-classify query (2.5, 2.5) with k=3 1/distance weights. The two class-0 neighbors are equidistant from the class-1 neighbor — so weighted vote changes nothing here, but the math is the point.
import numpy as np
from collections import Counter
X_train = np.array([[1,1],[2,1],[3,4],[5,5],[1,4]], dtype=float)
y_train = np.array([0, 0, 1, 1, 1])
query = np.array([2.5, 2.5])
dists = np.sqrt(np.sum((X_train - query)**2, axis=1))
k3 = np.argsort(dists)[:3]
weights = 1.0 / (dists[k3] + 1e-10)
wv = {}
for i, w in zip(k3, weights):
c = y_train[i]; wv[c] = wv.get(c, 0) + w
print('dists k3:', dists[k3].round(4).tolist())
print('weights: ', weights.round(4).tolist())
print('weighted votes:', {int(k): round(v,4) for k,v in wv.items()})
print('prediction:', max(wv, key=wv.get))Weights [0.6325, 0.6325, 0.4714]; class-0 total 1.1039, class-1 total 0.6325 -> prediction 0
Why: Neighbors 1 and 2 are equidistant (1.5811) so equal weight; neighbor 0 is farther (2.1213) so lower weight 0.4714. Class 0 still wins but the margin is softer than raw vote 2-1.
| neighbor | class | dist | 1/dist weight |
|---|---|---|---|
| (2,1) | 0 | 1.5811 | 0.6325 |
| (3,4) | 1 | 1.5811 | 0.6325 |
| (1,1) | 0 | 2.1213 | 0.4714 |
Comparison
Comparison matrix
From Weighted kNN trace — same 5-point example: refill the class column from what you know. The rest of the table is as it appeared.
| neighbor | class | dist | 1/dist weight |
|---|---|---|---|
| (2,1) | 0 | 1.5811 | 0.6325 |
| (3,4) | 1 | 1.5811 | 0.6325 |
| (1,1) | 0 | 2.1213 | 0.4714 |
Concept
On make_moons (n=400, noise=0.3): k=1 memorizes training set (acc 1.00), but test accuracy is lower. Larger k smooths the boundary — test accuracy can improve until k is so large it underfits.
| k | train acc | test acc | regime |
|---|---|---|---|
| 1 | 1.0000 | 0.8900 | overfitting (high variance) |
| 5 | 0.9067 | 0.8900 | balanced |
| 20 | 0.9033 | 0.9000 | slightly smoother boundary |
| 50 | 0.9000 | 0.9000 | approaching underfitting |
Rule of thumb: tune k via cross-validation. Common starting points are k=sqrt(n) or k=5. Always an odd k for binary classification to avoid ties.
Analogy
Discussion prompt
Explain Bias-variance tradeoff over k by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
On make_moons (n=400, noise=0.3): k=1 memorizes training set (acc 1.00), but test accuracy is lower. Larger k smooths the boundary — test accuracy can improve until k is so large it underfits.
Anomaly
Predict first
A student writes this, and it looks reasonable:
k=1 achieves 1.0000 training accuracy on every dataset — it memorizes perfectly. Use k=1 as the default.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: k=1 always classifies every training point correctly (it is its own nearest neighbor).
Evaluate k on a held-out validation set (or cross-validation), never on the training set.
Why: k=1 always classifies every training point correctly (it is its own nearest neighbor). Training accuracy is a useless signal here -- it is always 1.0 regardless of data.
Trap
k=1 achieves 1.0000 training accuracy on every dataset — it memorizes perfectly. Use k=1 as the default.
Deploy k=1 based on training-set performance
Why: k=1 always classifies every training point correctly (it is its own nearest neighbor). Training accuracy is a useless signal here -- it is always 1.0 regardless of data.
Evaluate k on a held-out validation set (or cross-validation), never on the training set.
Sweep k via CV; pick the k with best validation accuracy
Why: On make_moons, k=1 has 89% test accuracy vs k=20 which has 90%. The improvement is modest here but on noisy data the gap is large. k=1 is always the most overfit choice.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
1/distance — a very close neighbor counts more than a distant one.; Rule of thumb: tune k via cross-validation. Common starting points are k=sqrt(n) or k=5. Always an odd k for binary classification to avoid ties.age in years (range 0-90) and income_normalized (range 0-1). Feed both directly to L2 kNN.; k=1 achieves 1.0000 training accuracy on every dataset — it memorizes perfectly. Use k=1 as the default.Ranking
Put in order
These are the steps of kNN decision recipe, scrambled. Put them back in order before the next slide shows you.
StandardScaler before any distance computationkd_tree for d<20; ball_tree for d<100; brute or FAISS for d>100weights='distance') when neighbors vary greatly in proximityWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
StandardScaler before any distance computationkd_tree for d<20; ball_tree for d<100; brute or FAISS for d>100weights='distance') when neighbors vary greatly in proximityEdge cases
Discussion prompt
kNN decision recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
StandardScaler before any distance computationkd_tree for d<20; ball_tree for d<100; brute or FAISS for d>100weights='distance') when neighbors vary greatly in proximityElimination
Eliminate the wrong options
As the number of dimensions d grows, what happens to the ratio max_distance / min_distance among n random points?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Distance concentrates around the mean in high d. In our experiment: d=2 gives ratio 60.4; d=10 gives 3.6; d=500 gives 1.18. The gap between nearest and farthest neighbor vanishes, making kNN's ranking meaningless.
Check
Think before clicking.
Check your understanding
As the number of dimensions d grows, what happens to the ratio max_distance / min_distance among n random points?
Answer: A
Why: Distance concentrates around the mean in high d. In our experiment: d=2 gives ratio 60.4; d=10 gives 3.6; d=500 gives 1.18. The gap between nearest and farthest neighbor vanishes, making kNN's ranking meaningless.
Prediction
Predict first
You have n=1M vectors in d=128 dimensions (e.g. image embeddings). Which nearest-neighbor strategy is best?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: FAISS or HNSW approximate nearest neighbor
Why: At d=128, KD-tree degrades to near-brute-force (it only wins for d<20). Brute-force over 1M vectors is infeasible at inference time. FAISS/HNSW gives sub-linear approximate search that is fast enough and accurate enough for embedding retrieval.
Check
Match the method to the data regime.
Check your understanding
You have n=1M vectors in d=128 dimensions (e.g. image embeddings). Which nearest-neighbor strategy is best?
Answer: A
Why: At d=128, KD-tree degrades to near-brute-force (it only wins for d<20). Brute-force over 1M vectors is infeasible at inference time. FAISS/HNSW gives sub-linear approximate search that is fast enough and accurate enough for embedding retrieval.
Elimination
Eliminate the wrong options
Increasing k in kNN (holding everything else constant) has what effect on model bias and variance?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Large k averages more neighbors, smoothing the decision boundary — this reduces variance (less sensitive to individual points) but introduces bias (the boundary can't capture fine structure). k=1 is maximum variance/minimum bias; k=n is maximum bias/minimum variance.
Check
Predict the direction of the effect.
Check your understanding
Increasing k in kNN (holding everything else constant) has what effect on model bias and variance?
Answer: A
Why: Large k averages more neighbors, smoothing the decision boundary — this reduces variance (less sensitive to individual points) but introduces bias (the boundary can't capture fine structure). k=1 is maximum variance/minimum bias; k=n is maximum bias/minimum variance.
Section
Project
Concept
Implement kNN from scratch supporting both L2 and cosine distance. Benchmark it against sklearn on iris, then demonstrate the curse of dimensionality.
| # | requirement | tool |
|---|---|---|
| 1 | kNN from scratch (L2 + cosine) | numpy only |
| 2 | match sklearn KNeighborsClassifier | k=3 on iris |
| 3 | curse of dimensionality table | max/min ratio vs d |
| 4 | bias-variance sweep over k | make_moons, k=1,5,20 |
Build rules: implement euclidean_dist and cosine_dist as vectorized numpy functions (no loops over training set), then sweep over the query set.
Counterexample
Discussion prompt
Implement kNN from scratch supporting both L2 and cosine distance. Benchmark it against sklearn on iris, then demonstrate the curse of dimensionality.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: implement euclidean_dist and cosine_dist as vectorized numpy functions (no loops over training set), then sweep over the query set.
Worked example
Your turn: implement kNN from scratch with L2 and cosine options. Predict whether L2 or cosine will score higher on iris.
Hint: for L2 use np.sqrt(np.sum((X_train - x)**2, axis=1)); for cosine compute dot product divided by norms.
import numpy as np
from collections import Counter
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
iris = load_iris()
X_tr, X_te, y_tr, y_te = train_test_split(
iris.data, iris.target, test_size=0.2, random_state=42)
def knn_predict(X_train, y_train, X_test, k=3, metric='l2'):
preds = []
for x in X_test:
if metric == 'l2':
dists = np.sqrt(np.sum((X_train - x)**2, axis=1))
else: # cosine
dots = X_train @ x
norms = np.linalg.norm(X_train, axis=1) * np.linalg.norm(x)
dists = 1 - dots / (norms + 1e-10)
idx = np.argsort(dists)[:k]
preds.append(Counter(y_train[idx]).most_common(1)[0][0])
return np.array(preds)
print('L2 acc:', accuracy_score(y_te, knn_predict(X_tr, y_tr, X_te, 3, 'l2')))
print('cos acc:', accuracy_score(y_te, knn_predict(X_tr, y_tr, X_te, 3, 'cosine')))| metric | k=3 accuracy on iris (test) |
|---|---|
| L2 (from scratch) | 1.0000 |
| cosine (from scratch) | 0.9667 |
| sklearn L2 (reference) | 1.0000 |
Worked example
Your turn: measure max/min distance ratio as d increases from 2 to 500. Predict at what dimension the ratio drops below 2.
Hint: rng.standard_normal((1000, d)) for points; a single query row; np.sqrt(np.sum((pts - q)**2, axis=1)) for distances.
import numpy as np
rng = np.random.default_rng(0)
for d in [2, 10, 50, 100, 500]:
pts = rng.standard_normal((1000, d))
q = rng.standard_normal((1, d))
dists = np.sqrt(np.sum((pts - q)**2, axis=1))
print(f'd={d:4d} ratio={dists.max()/dists.min():.4f}')| d | ratio max/min |
|---|---|
| 2 | 60.3558 |
| 10 | 3.6317 |
| 50 | 1.7607 |
| 100 | 1.4445 |
| 500 | 1.1839 |
Trade off
Comparison matrix
From Milestone 2 — curse of dimensionality: every row here is a choice with a cost. Fill the ratio max/min column, then say which row you would actually pick and what you give up for it.
| d | ratio max/min |
|---|---|
| 2 | 60.3558 |
| 10 | 3.6317 |
| 50 | 1.7607 |
| 100 | 1.4445 |
| 500 | 1.1839 |
Concept
import numpy as np
from collections import Counter
from sklearn.datasets import load_iris, make_moons
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
# 1. kNN from scratch on iris
iris = load_iris()
X_tr, X_te, y_tr, y_te = train_test_split(iris.data, iris.target, test_size=0.2, random_state=42)
def knn_l2(X_train, y_train, X_test, k):
preds = []
for x in X_test:
dists = np.sqrt(np.sum((X_train - x)**2, axis=1))
preds.append(Counter(y_train[np.argsort(dists)[:k]]).most_common(1)[0][0])
return np.array(preds)
print('scratch L2 k=3:', accuracy_score(y_te, knn_l2(X_tr, y_tr, X_te, 3)))
# 2. Bias-variance sweep on make_moons
X, y = make_moons(n_samples=400, noise=0.3, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=42)
for k in [1, 5, 20]:
knn = KNeighborsClassifier(n_neighbors=k).fit(Xtr, ytr)
print(f'k={k}: train={accuracy_score(ytr, knn.predict(Xtr)):.4f} test={accuracy_score(yte, knn.predict(Xte)):.4f}')| result | value |
|---|---|
| scratch L2 k=3 (iris) | 1.0000 |
| k=1 train / test (moons) | 1.0000 / 0.8900 |
| k=5 train / test (moons) | 0.9067 / 0.8900 |
| k=20 train / test (moons) | 0.9033 / 0.9000 |
If your scratch kNN matches sklearn and your bias-variance table shows train accuracy falling as k increases while test accuracy stabilizes — you've built kNN and measured its two fundamental failure modes.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| result | value |
|---|---|
| scratch L2 k=3 (iris) | 1.0000 |
| k=1 train / test (moons) | 1.0000 / 0.8900 |
| k=5 train / test (moons) | 0.9067 / 0.8900 |
| k=20 train / test (moons) | 0.9033 / 0.9000 |
Concept
Out loud, slides closed: explain (1) why kNN has no training step and what that means for compute cost, (2) the curse of dimensionality in terms of the max/min ratio, and (3) when to use KD-tree vs FAISS.
Stretch (homework): implement cosine-weighted kNN from scratch and compare to distance-weighted on a text-classification toy set; add FAISS approximate search for a synthetic 1M-point dataset in d=64. Next up: ensemble methods — bagging and random forests.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — kNN: majority-vote classification · Distance metrics · Curse of dimensionality · Efficient search & weighted kNN · Your turn: kNN from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| idea | the one thing to remember |
|---|---|
| kNN prediction cost | O(nd) — no training, all cost at inference |
| curse of dimensionality | max/min ratio -> 1 as d grows; use d < 20 or reduce first |
| index choice | KD-tree d<20, ball tree d<100, FAISS/HNSW for d>=128 or n>=1M |
| k tuning | large k = high bias / low variance; tune via cross-validation |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.