Lesson 69: kNN and the Curse of Dimensionality

USAAIO Lesson 69, from Week 24 of Phase 4. It builds kNN majority-vote classification from scratch, covers the L2, L1, and cosine distance metrics, and works the curse of dimensionality, where the ratio of the maximum to the minimum distance converges to 1. It then covers the KD-tree and ball-tree for efficient lookup, weighted kNN using 1/distance, and the bias-variance trade-off as k changes. Every number was verified with numpy and sklearn in June 2026. The lesson runs to 32 slides.

Subject: Machine Learning · 63 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. kNN & the Curse of Dimensionality

Title

USAAIO · Lesson 69 · Week 24

The simplest possible model — and the hardest possible caveat. Majority vote, distance metrics, KD-trees, and why kNN quietly breaks in high dimensions.

2. By the end of this lesson you can

Objectives

  1. Implement kNN from scratch with L2 and cosine distance and match sklearn
  2. Explain and compute the curse of dimensionality (max/min ratio converging to 1)
  3. Choose between KD-tree, ball tree, and brute force by data shape
  4. Apply 1/distance weighting and explain the bias-variance tradeoff over k
  5. Decide when kNN is and is not the right tool for a given problem

3. What survived from AdaBoost?

Warm-up

Discussion prompt

Before we open Lesson 69: kNN and the Curse of Dimensionality: without looking back, what was the main idea of AdaBoost, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

AdaBoost from scratch — sample-weight updates, the alpha formula, exponential-loss minimization, gradient descent in function space, and convergence theory. Implement decision-stump AdaBoost, verify it matches sklearn, and compare it to a single deep tree.

4. kNN: majority-vote classification

Section

Part 1 of 4

5. The algorithm: no training, O(n) prediction

Concept

kNN stores the full training set. At prediction time it computes the distance from the query to every training point, takes the k closest, and returns the majority class.

  1. No training step — model IS the training data
  2. Prediction cost: O(nd) per query (n points, d dimensions)
  3. Lazy learning: all computation deferred to inference

k=1 is the most sensitive (nearest neighbor wins alone). Larger k smooths the decision boundary and reduces variance at the cost of bias (Lesson 32).

6. Why majority vote?

Intuition

Think of k=5 as asking five of your nearest neighbors to vote. If four say 'spam' and one says 'ham', you go with spam. The noise in any single neighbor averages out over the group.

k=1 is over-committed: one mislabeled outlier anywhere in the training set can flip a prediction. k=20 draws a smoother boundary but misses tight clusters.

7. Break it if you can: Why majority vote?

Counterexample

Discussion prompt

Think of k=5 as asking five of your nearest neighbors to vote. If four say 'spam' and one says 'ham', you go with spam. The noise in any single neighbor averages out over the group.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

8. Guess the shape of the answer: kNN from scratch — 5-point trace

Estimation

Predict first

Classify query (2.5, 2.5) using k=3 against 5 labeled training points. Predict the answer before computing.

Commit before you compute: what does kNN from scratch — 5-point trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Sort distances: indices [1,2,0,4,3]; k=3 neighbors are classes [0,1,0] -> majority class 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Two class-0 votes beat one class-1 vote: prediction is 0.

9. kNN from scratch — 5-point trace

Worked example

Classify query (2.5, 2.5) using k=3 against 5 labeled training points. Predict the answer before computing.

import numpy as np
from collections import Counter

X_train = np.array([[1,1],[2,1],[3,4],[5,5],[1,4]], dtype=float)
y_train = np.array([0, 0, 1, 1, 1])
query   = np.array([2.5, 2.5])

dists = np.sqrt(np.sum((X_train - query)**2, axis=1))
sorted_idx = np.argsort(dists)
k3 = sorted_idx[:3]
votes = Counter(y_train[k3])
print('dists:', dists.round(4).tolist())
print('sorted_idx:', sorted_idx.tolist())
print('k=3 neighbors:', k3.tolist(), 'classes:', y_train[k3].tolist())
print('prediction:', votes.most_common(1)[0][0])

Sort distances: indices [1,2,0,4,3]; k=3 neighbors are classes [0,1,0] -> majority class 0

Why: Two class-0 votes beat one class-1 vote: prediction is 0. Note indices 1 and 2 are equidistant (1.5811) — tie-break by stable argsort order.

idxpointclassL2 dist
0(1,1)02.1213
1(2,1)01.5811
2(3,4)11.5811
3(5,5)13.5355
4(1,4)12.1213

10. Fill in: point for kNN from scratch — 5-point trace

Comparison

Comparison matrix

From kNN from scratch — 5-point trace: refill the point column from what you know. The rest of the table is as it appeared.

idxpointclassL2 dist
0(1,1)02.1213
1(2,1)01.5811
2(3,4)11.5811
3(5,5)13.5355
4(1,4)12.1213

11. Distance metrics

Section

Part 2 of 4

12. L2, L1, Minkowski, cosine, Mahalanobis

Concept

metricformulabest for
L2 (Euclidean)sqrt(sum (a-b)^2)continuous, isotropic features
L1 (Manhattan)sum |a-b|sparse data, outlier-robust
Minkowski (p)(sum |a-b|^p)^(1/p)generalizes L1/L2 (p=1 or 2)
Cosine1 - dot(a,b)/(|a||b|)text/embedding direction
Mahalanobissqrt((a-b)^T S^-1 (a-b))accounts for feature covariance

For raw pixel or sensor features: L2 is the default. For high-dimensional count vectors (word counts, embeddings): cosine. For correlated features with different scales: Mahalanobis (or just StandardScaler + L2).

13. What each one costs: L2, L1, Minkowski, cosine, Mahalanobis

Trade off

Comparison matrix

From L2, L1, Minkowski, cosine, Mahalanobis: every row here is a choice with a cost. Fill the formula column, then say which row you would actually pick and what you give up for it.

metricformulabest for
L2 (Euclidean)sqrt(sum (a-b)^2)continuous, isotropic features
L1 (Manhattan)sum |a-b|sparse data, outlier-robust
Minkowski (p)(sum |a-b|^p)^(1/p)generalizes L1/L2 (p=1 or 2)
Cosine1 - dot(a,b)/(|a||b|)text/embedding direction
Mahalanobissqrt((a-b)^T S^-1 (a-b))accounts for feature covariance

14. What has to be given first: Distance metric worked example

Missing information

Discussion prompt

Compute L2, L1, and cosine distance from A=(1,2) to B=(4,6) and to C=(1,3). Which is closer under each metric?

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

B lies along nearly the same direction as A (both tilted northeast), so the angle is smaller. C is one unit directly above A, so L2/L1 are small but the direction shifts more.

15. Distance metric worked example

Worked example

Compute L2, L1, and cosine distance from A=(1,2) to B=(4,6) and to C=(1,3). Which is closer under each metric?

import numpy as np
A = np.array([1.0, 2.0])
B = np.array([4.0, 6.0])
C = np.array([1.0, 3.0])

def cos_dist(u, v):
    return 1 - np.dot(u, v) / (np.linalg.norm(u) * np.linalg.norm(v))

print('L2  A-B:', np.linalg.norm(A-B))   # 5.0
print('L2  A-C:', np.linalg.norm(A-C))   # 1.0
print('L1  A-B:', np.sum(np.abs(A-B)))   # 7.0
print('L1  A-C:', np.sum(np.abs(A-C)))   # 1.0
print('cos A-B:', round(cos_dist(A,B),4)) # 0.0077
print('cos A-C:', round(cos_dist(A,C),4)) # 0.0101

C is closest under L2 (1.0) and L1 (1.0); B is closer under cosine (0.0077 vs 0.0101)

Why: B lies along nearly the same direction as A (both tilted northeast), so the angle is smaller. C is one unit directly above A, so L2/L1 are small but the direction shifts more.

metricdist(A,B)dist(A,C)nearest
L25.00001.0000C
L17.00001.0000C
cosine0.00770.0101B

16. Watch it run: Distance metric worked example

Pattern

Step through it

Step through Distance metric worked example one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: metric is L2
  2. Step 2: metric is L1
  3. Step 3: metric is cosine

17. Something is wrong here: using raw features without scaling

Anomaly

Predict first

A student writes this, and it looks reasonable:

Features: age in years (range 0-90) and income_normalized (range 0-1). Feed both directly to L2 kNN.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1.

Apply StandardScaler before kNN whenever features differ in scale.

Why: A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1. kNN effectively ignores income entirely.

18. Trap: using raw features without scaling

Trap

The trap

Features: age in years (range 0-90) and income_normalized (range 0-1). Feed both directly to L2 kNN.

L2 distance is dominated by age alone

Why: A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1. kNN effectively ignores income entirely.

The fix

Apply StandardScaler before kNN whenever features differ in scale.

After StandardScaler, each feature contributes equally to L2 distance

Why: Z-score normalization gives every feature mean 0 and std 1. This is mandatory for distance-based models including kNN, SVM (L15), and clustering (L52).

19. Break it on purpose: using raw features without scaling

Break the constraint

Discussion prompt

The rule this trap just fixed:

Z-score normalization gives every feature mean 0 and std 1. This is mandatory for distance-based models including kNN, SVM (L15), and clustering (L52).

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

A difference of 50 in age contributes 2500 to squared L2, while the entire income range contributes at most 1. kNN effectively ignores income entirely.

20. Curse of dimensionality

Section

Part 3 of 4

21. In high d, all points become equidistant

Concept

As dimensionality grows, the ratio of the farthest to the nearest neighbor shrinks toward 1 — every point looks equally far from every other. The notion of 'nearest' becomes meaningless.

\[ \frac{d_{\max}(n,d)}{d_{\min}(n,d)} \xrightarrow{d \to \infty} 1 \]

This is because in high d, distance concentrates around the mean. Adding irrelevant dimensions drowns signal in noise.

22. By analogy: In high d, all points become equidistant

Analogy

Discussion prompt

Explain In high d, all points become equidistant by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

As dimensionality grows, the ratio of the farthest to the nearest neighbor shrinks toward 1 — every point looks equally far from every other. The notion of 'nearest' becomes meaningless.

23. Guess the shape of the answer: Curse of dimensionality — measured

Estimation

Predict first

Draw 1000 standard-normal points in d dimensions, then measure max/min distance to a query. Watch the ratio collapse as d increases.

Commit before you compute: what does Curse of dimensionality — measured come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: d=2 ratio 60.4; d=10 ratio 3.6; d=50 ratio 1.8; d=100 ratio 1.4; d=500 ratio 1.2

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Even at d=500 the ratio is only 1.19 -- the nearest and farthest neighbor are 19% apart in distance.

24. Curse of dimensionality — measured

Worked example

Draw 1000 standard-normal points in d dimensions, then measure max/min distance to a query. Watch the ratio collapse as d increases.

import numpy as np
rng = np.random.default_rng(0)
dims = [2, 10, 50, 100, 500]
print(f'dim | max_dist | min_dist | ratio')
for d in dims:
    pts = rng.standard_normal((1000, d))
    q   = rng.standard_normal((1, d))
    dists = np.sqrt(np.sum((pts - q)**2, axis=1))
    print(f'{d:>3}  | {dists.max():8.4f} | {dists.min():8.4f} | {dists.max()/dists.min():6.4f}')

d=2 ratio 60.4; d=10 ratio 3.6; d=50 ratio 1.8; d=100 ratio 1.4; d=500 ratio 1.2

Why: Even at d=500 the ratio is only 1.19 -- the nearest and farthest neighbor are 19% apart in distance. kNN's ranking is still somewhat meaningful here but degrades fast with irrelevant features.

dimmax distmin distratio
24.42530.073360.36
107.18241.97773.63
5012.66167.19121.76
10016.523911.43921.44
50034.287028.96191.18

25. Watch it run: Curse of dimensionality — measured

Pattern

Step through it

Step through Curse of dimensionality — measured one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: dim is 2
  2. Step 2: dim is 10
  3. Step 3: dim is 50
  4. Step 4: dim is 100
  5. Step 5: dim is 500

26. Practical consequence: when to avoid kNN

Concept

Raw image pixels (d=784 for MNIST), raw text TF-IDF vectors (d=10k+), and unfiltered sensor arrays all violate the low-d assumption kNN requires.

The KD-tree index is also exact only for d < 20. Above that it degrades to brute force — a second reason high-d kNN is slow.

27. Where does each piece belong: Lesson 69: kNN and the Curse of…

Sorting

Sort into buckets

These are the pieces of Lesson 69: kNN and the Curse of Dimensionality, out of order. Put each one back under the part of the lesson it belongs to.

kNN: majority-vote classification
The algorithm: no training, O(n) prediction; Why majority vote?; kNN from scratch — 5-point trace
Distance metrics
L2, L1, Minkowski, cosine, Mahalanobis; Distance metric worked example
Curse of dimensionality
In high d, all points become equidistant; Curse of dimensionality — measured; Practical consequence: when to avoid kNN
s1
kNN: majority-vote classification is where Lesson 69: kNN and the Curse of Dimensionality puts The algorithm: no training, O(n) prediction, Why majority vote?, kNN from scratch — 5-point trace. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Distance metrics is where Lesson 69: kNN and the Curse of Dimensionality puts L2, L1, Minkowski, cosine, Mahalanobis, Distance metric worked example. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Curse of dimensionality is where Lesson 69: kNN and the Curse of Dimensionality puts In high d, all points become equidistant, Curse of dimensionality — measured, Practical consequence: when to avoid kNN. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

28. Efficient search & weighted kNN

Section

Part 4 of 4

29. KD-tree, ball tree, and approximate NN

Concept

methodexact?complexity (query)good for
brute forceyesO(nd)small n or large d
KD-treeyesO(log n) avgd < 20, moderate n
ball treeyesO(log n) avgd 20-100, curved manifolds
FAISS / HNSWapproxsub-linearmillions of vectors, d > 100

sklearn selects the algorithm automatically (algorithm='auto') based on n and d. For production with 1M+ vectors and d=128 embeddings, FAISS is the standard (not in sklearn).

30. Which is which, by exact?

Discrimination

Sort into buckets

Sort these by exact?, from memory, without looking back at KD-tree, ball tree, and approximate NN. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

yes
brute force; KD-tree; ball tree
approx
FAISS / HNSW
g1
exact? is "yes" for brute force, KD-tree, ball tree — that is what the table on "KD-tree, ball tree, and approximate NN" records, and it is the single property separating this group from the rest.
g2
exact? is "approx" for FAISS / HNSW — that is what the table on "KD-tree, ball tree, and approximate NN" records, and it is the single property separating this group from the rest.

31. What has to be given first: KD-tree vs brute force — timing on 5000…

Missing information

Discussion prompt

Compare predict-time for brute force, KD-tree, and ball tree on n=5000, d=10 classification (make_classification). Same accuracy expected.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

KD-tree is 26x faster than brute force at d=10. All three produce identical predictions — the index structure doesn't change the result, only how fast you find the neighbors.

32. KD-tree vs brute force — timing on 5000 points

Worked example

Compare predict-time for brute force, KD-tree, and ball tree on n=5000, d=10 classification (make_classification). Same accuracy expected.

import time
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score

X, y = make_classification(n_samples=5000, n_features=10, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)

for algo in ['brute', 'kd_tree', 'ball_tree']:
    knn = KNeighborsClassifier(n_neighbors=5, algorithm=algo)
    t0 = time.time()
    knn.fit(X_tr, y_tr)
    acc = accuracy_score(y_te, knn.predict(X_te))
    print(f'{algo:12s}: acc={acc:.4f}, time={time.time()-t0:.3f}s')

brute=2.4s, kd_tree=0.09s, ball_tree=0.13s; all accuracy 0.9670

Why: KD-tree is 26x faster than brute force at d=10. All three produce identical predictions — the index structure doesn't change the result, only how fast you find the neighbors.

algorithmaccuracytime (s)
brute0.96702.398
kd_tree0.96700.109
ball_tree0.96700.138

33. What stays fixed: KD-tree vs brute force — timing on 5000 points

Invariant

Step through it

Step through KD-tree vs brute force — timing on 5000 points one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: algorithm is brute
  2. Step 2: algorithm is kd_tree
  3. Step 3: algorithm is ball_tree

34. Weighted kNN: 1/distance votes

Concept

Standard kNN gives every neighbor one vote. Weighted kNN gives each neighbor a vote proportional to 1/distance — a very close neighbor counts more than a distant one.

\[ \hat{y} = \arg\max_c \sum_{i \in N_k(x)} \frac{\mathbf{1}[y_i = c]}{d(x, x_i) + \epsilon} \]

This is the continuous generalization of majority vote: as k grows large, weighted kNN approaches a kernel density classifier with a box kernel.

35. Teach it back: Weighted kNN: 1/distance votes

Explain it

Discussion prompt

Explain Weighted kNN: 1/distance votes to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Standard kNN gives every neighbor one vote. Weighted kNN gives each neighbor a vote proportional to 1/distance — a very close neighbor counts more than a distant one.

36. Guess the shape of the answer: Weighted kNN trace — same 5-point example

Estimation

Predict first

Re-classify query (2.5, 2.5) with k=3 1/distance weights. The two class-0 neighbors are equidistant from the class-1 neighbor — so weighted vote changes nothing here, but the math is the point.

Commit before you compute: what does Weighted kNN trace — same 5-point example come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Weights [0.6325, 0.6325, 0.4714]; class-0 total 1.1039, class-1 total 0.6325 -> prediction 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Neighbors 1 and 2 are equidistant (1.5811) so equal weight; neighbor 0 is farther (2.1213) so lower weight 0.4714.

37. Weighted kNN trace — same 5-point example

Worked example

Re-classify query (2.5, 2.5) with k=3 1/distance weights. The two class-0 neighbors are equidistant from the class-1 neighbor — so weighted vote changes nothing here, but the math is the point.

import numpy as np
from collections import Counter

X_train = np.array([[1,1],[2,1],[3,4],[5,5],[1,4]], dtype=float)
y_train = np.array([0, 0, 1, 1, 1])
query   = np.array([2.5, 2.5])
dists = np.sqrt(np.sum((X_train - query)**2, axis=1))
k3 = np.argsort(dists)[:3]
weights = 1.0 / (dists[k3] + 1e-10)
wv = {}
for i, w in zip(k3, weights):
    c = y_train[i]; wv[c] = wv.get(c, 0) + w
print('dists k3:', dists[k3].round(4).tolist())
print('weights: ', weights.round(4).tolist())
print('weighted votes:', {int(k): round(v,4) for k,v in wv.items()})
print('prediction:', max(wv, key=wv.get))

Weights [0.6325, 0.6325, 0.4714]; class-0 total 1.1039, class-1 total 0.6325 -> prediction 0

Why: Neighbors 1 and 2 are equidistant (1.5811) so equal weight; neighbor 0 is farther (2.1213) so lower weight 0.4714. Class 0 still wins but the margin is softer than raw vote 2-1.

neighborclassdist1/dist weight
(2,1)01.58110.6325
(3,4)11.58110.6325
(1,1)02.12130.4714

38. Fill in: class for Weighted kNN trace — same 5-point example

Comparison

Comparison matrix

From Weighted kNN trace — same 5-point example: refill the class column from what you know. The rest of the table is as it appeared.

neighborclassdist1/dist weight
(2,1)01.58110.6325
(3,4)11.58110.6325
(1,1)02.12130.4714

39. Bias-variance tradeoff over k

Concept

On make_moons (n=400, noise=0.3): k=1 memorizes training set (acc 1.00), but test accuracy is lower. Larger k smooths the boundary — test accuracy can improve until k is so large it underfits.

ktrain acctest accregime
11.00000.8900overfitting (high variance)
50.90670.8900balanced
200.90330.9000slightly smoother boundary
500.90000.9000approaching underfitting

Rule of thumb: tune k via cross-validation. Common starting points are k=sqrt(n) or k=5. Always an odd k for binary classification to avoid ties.

40. By analogy: Bias-variance tradeoff over k

Analogy

Discussion prompt

Explain Bias-variance tradeoff over k by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

On make_moons (n=400, noise=0.3): k=1 memorizes training set (acc 1.00), but test accuracy is lower. Larger k smooths the boundary — test accuracy can improve until k is so large it underfits.

41. Something is wrong here: choosing k=1 because it gets 100% training accuracy

Anomaly

Predict first

A student writes this, and it looks reasonable:

k=1 achieves 1.0000 training accuracy on every dataset — it memorizes perfectly. Use k=1 as the default.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: k=1 always classifies every training point correctly (it is its own nearest neighbor).

Evaluate k on a held-out validation set (or cross-validation), never on the training set.

Why: k=1 always classifies every training point correctly (it is its own nearest neighbor). Training accuracy is a useless signal here -- it is always 1.0 regardless of data.

42. Trap: choosing k=1 because it gets 100% training accuracy

Trap

The trap

k=1 achieves 1.0000 training accuracy on every dataset — it memorizes perfectly. Use k=1 as the default.

Deploy k=1 based on training-set performance

Why: k=1 always classifies every training point correctly (it is its own nearest neighbor). Training accuracy is a useless signal here -- it is always 1.0 regardless of data.

The fix

Evaluate k on a held-out validation set (or cross-validation), never on the training set.

Sweep k via CV; pick the k with best validation accuracy

Why: On make_moons, k=1 has 89% test accuracy vs k=20 which has 90%. The improvement is modest here but on noisy data the gap is large. k=1 is always the most overfit choice.

43. Which of these survive contact with Lesson 69: kNN and the Curse of…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Raw image pixels (d=784 for MNIST), raw text TF-IDF vectors (d=10k+), and unfiltered sensor arrays all violate the low-d assumption kNN requires.; Standard kNN gives every neighbor one vote. Weighted kNN gives each neighbor a vote proportional to 1/distance — a very close neighbor counts more than a distant one.; Rule of thumb: tune k via cross-validation. Common starting points are k=sqrt(n) or k=5. Always an odd k for binary classification to avoid ties.
Breaks
Features: age in years (range 0-90) and income_normalized (range 0-1). Feed both directly to L2 kNN.; k=1 achieves 1.0000 training accuracy on every dataset — it memorizes perfectly. Use k=1 as the default.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 69: kNN and the Curse of Dimensionality puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

44. Rebuild the recipe: kNN decision recipe

Ranking

Put in order

These are the steps of kNN decision recipe, scrambled. Put them back in order before the next slide shows you.

  1. Scale features with StandardScaler before any distance computation
  2. Check dimensionality: if d > 20, reduce with PCA first
  3. Choose metric: L2 (default), cosine (embeddings/text), L1 (sparse/outlier-robust)
  4. Tune k via cross-validation; start at k=5 or k=sqrt(n), always odd for binary
  5. Choose algorithm: kd_tree for d<20; ball_tree for d<100; brute or FAISS for d>100
  6. Weighted kNN (weights='distance') when neighbors vary greatly in proximity

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

45. kNN decision recipe

Pattern

  1. Scale features with StandardScaler before any distance computation
  2. Check dimensionality: if d > 20, reduce with PCA first
  3. Choose metric: L2 (default), cosine (embeddings/text), L1 (sparse/outlier-robust)
  4. Tune k via cross-validation; start at k=5 or k=sqrt(n), always odd for binary
  5. Choose algorithm: kd_tree for d<20; ball_tree for d<100; brute or FAISS for d>100
  6. Weighted kNN (weights='distance') when neighbors vary greatly in proximity

46. Where does it stop working: kNN decision recipe

Edge cases

Discussion prompt

kNN decision recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Scale features with StandardScaler before any distance computation
  2. Check dimensionality: if d > 20, reduce with PCA first
  3. Choose metric: L2 (default), cosine (embeddings/text), L1 (sparse/outlier-robust)
  4. Tune k via cross-validation; start at k=5 or k=sqrt(n), always odd for binary
  5. Choose algorithm: kd_tree for d<20; ball_tree for d<100; brute or FAISS for d>100
  6. Weighted kNN (weights='distance') when neighbors vary greatly in proximity

47. Rule out three: Check yourself — curse of dimensionality

Elimination

Eliminate the wrong options

As the number of dimensions d grows, what happens to the ratio max_distance / min_distance among n random points?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. It approaches 1 — all points become roughly equidistant
  • B. It grows without bound — the farthest point gets much farther
  • C. It stays constant because distances all scale with sqrt(d)
  • D. It approaches 0 — all points collapse to the origin

Survives elimination: A

Why: Distance concentrates around the mean in high d. In our experiment: d=2 gives ratio 60.4; d=10 gives 3.6; d=500 gives 1.18. The gap between nearest and farthest neighbor vanishes, making kNN's ranking meaningless.

48. Check yourself — curse of dimensionality

Check

Think before clicking.

Check your understanding

As the number of dimensions d grows, what happens to the ratio max_distance / min_distance among n random points?

  • A. It approaches 1 — all points become roughly equidistant (correct)
  • B. It grows without bound — the farthest point gets much farther
  • C. It stays constant because distances all scale with sqrt(d)
  • D. It approaches 0 — all points collapse to the origin

Answer: A

Why: Distance concentrates around the mean in high d. In our experiment: d=2 gives ratio 60.4; d=10 gives 3.6; d=500 gives 1.18. The gap between nearest and farthest neighbor vanishes, making kNN's ranking meaningless.

Why B tempts people
Both max and min distance scale with sqrt(d) — they grow, but at the same rate, so their ratio shrinks rather than grows.
Why C tempts people
The mean distance does scale with sqrt(d), but max and min concentrate tightly around that mean; the ratio (max/min) narrows, not stays flat.
Why D tempts people
Points do not collapse to the origin in L2 norm — they spread outward with sqrt(d). The origin has no special role.

49. Answer it before you see the options: Check yourself — index selection

Prediction

Predict first

You have n=1M vectors in d=128 dimensions (e.g. image embeddings). Which nearest-neighbor strategy is best?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: FAISS or HNSW approximate nearest neighbor

Why: At d=128, KD-tree degrades to near-brute-force (it only wins for d<20). Brute-force over 1M vectors is infeasible at inference time. FAISS/HNSW gives sub-linear approximate search that is fast enough and accurate enough for embedding retrieval.

50. Check yourself — index selection

Check

Match the method to the data regime.

Check your understanding

You have n=1M vectors in d=128 dimensions (e.g. image embeddings). Which nearest-neighbor strategy is best?

  • A. FAISS or HNSW approximate nearest neighbor (correct)
  • B. KD-tree exact search
  • C. Brute-force O(nd) scan
  • D. Ball tree exact search

Answer: A

Why: At d=128, KD-tree degrades to near-brute-force (it only wins for d<20). Brute-force over 1M vectors is infeasible at inference time. FAISS/HNSW gives sub-linear approximate search that is fast enough and accurate enough for embedding retrieval.

Why B tempts people
KD-tree is efficient only for d < 20; at d=128 it reverts to O(nd) without the overhead of tree construction. It is not recommended above d=20.
Why C tempts people
Brute-force is O(nd) per query: 1M x 128 = 128M multiplications every inference call. This is feasible for batch offline jobs but not real-time retrieval.
Why D tempts people
Ball tree handles curved manifolds better than KD-tree and works to d~100, but at d=128 and n=1M it is still too slow for real-time use; FAISS is the correct tool.

51. Rule out three: Check yourself — bias-variance over k

Elimination

Eliminate the wrong options

Increasing k in kNN (holding everything else constant) has what effect on model bias and variance?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Increases bias, decreases variance
  • B. Decreases bias, increases variance
  • C. Increases both bias and variance
  • D. Decreases both bias and variance

Survives elimination: A

Why: Large k averages more neighbors, smoothing the decision boundary — this reduces variance (less sensitive to individual points) but introduces bias (the boundary can't capture fine structure). k=1 is maximum variance/minimum bias; k=n is maximum bias/minimum variance.

52. Check yourself — bias-variance over k

Check

Predict the direction of the effect.

Check your understanding

Increasing k in kNN (holding everything else constant) has what effect on model bias and variance?

  • A. Increases bias, decreases variance (correct)
  • B. Decreases bias, increases variance
  • C. Increases both bias and variance
  • D. Decreases both bias and variance

Answer: A

Why: Large k averages more neighbors, smoothing the decision boundary — this reduces variance (less sensitive to individual points) but introduces bias (the boundary can't capture fine structure). k=1 is maximum variance/minimum bias; k=n is maximum bias/minimum variance.

Why B tempts people
This is backwards: large k smooths the boundary (lower variance, higher bias). Small k memorizes (high variance, low bias).
Why C tempts people
You cannot simultaneously increase both: the bias-variance tradeoff is a tradeoff. Smoothing always trades variance for bias.
Why D tempts people
Increasing k does reduce variance, but it does so at the cost of increased bias. There is no free lunch — bias must rise when variance falls.

53. Your turn: kNN from scratch

Section

Project

54. Project: kNN classifier with two metrics

Concept

Implement kNN from scratch supporting both L2 and cosine distance. Benchmark it against sklearn on iris, then demonstrate the curse of dimensionality.

#requirementtool
1kNN from scratch (L2 + cosine)numpy only
2match sklearn KNeighborsClassifierk=3 on iris
3curse of dimensionality tablemax/min ratio vs d
4bias-variance sweep over kmake_moons, k=1,5,20

Build rules: implement euclidean_dist and cosine_dist as vectorized numpy functions (no loops over training set), then sweep over the query set.

55. Break it if you can: Project: kNN classifier with two metrics

Counterexample

Discussion prompt

Implement kNN from scratch supporting both L2 and cosine distance. Benchmark it against sklearn on iris, then demonstrate the curse of dimensionality.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: implement euclidean_dist and cosine_dist as vectorized numpy functions (no loops over training set), then sweep over the query set.

56. Milestone 1 — kNN from scratch

Worked example

Your turn: implement kNN from scratch with L2 and cosine options. Predict whether L2 or cosine will score higher on iris.

Hint: for L2 use np.sqrt(np.sum((X_train - x)**2, axis=1)); for cosine compute dot product divided by norms.

import numpy as np
from collections import Counter
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

iris = load_iris()
X_tr, X_te, y_tr, y_te = train_test_split(
    iris.data, iris.target, test_size=0.2, random_state=42)

def knn_predict(X_train, y_train, X_test, k=3, metric='l2'):
    preds = []
    for x in X_test:
        if metric == 'l2':
            dists = np.sqrt(np.sum((X_train - x)**2, axis=1))
        else:  # cosine
            dots  = X_train @ x
            norms = np.linalg.norm(X_train, axis=1) * np.linalg.norm(x)
            dists = 1 - dots / (norms + 1e-10)
        idx = np.argsort(dists)[:k]
        preds.append(Counter(y_train[idx]).most_common(1)[0][0])
    return np.array(preds)

print('L2  acc:', accuracy_score(y_te, knn_predict(X_tr, y_tr, X_te, 3, 'l2')))
print('cos acc:', accuracy_score(y_te, knn_predict(X_tr, y_tr, X_te, 3, 'cosine')))
metrick=3 accuracy on iris (test)
L2 (from scratch)1.0000
cosine (from scratch)0.9667
sklearn L2 (reference)1.0000

57. Milestone 2 — curse of dimensionality

Worked example

Your turn: measure max/min distance ratio as d increases from 2 to 500. Predict at what dimension the ratio drops below 2.

Hint: rng.standard_normal((1000, d)) for points; a single query row; np.sqrt(np.sum((pts - q)**2, axis=1)) for distances.

import numpy as np
rng = np.random.default_rng(0)
for d in [2, 10, 50, 100, 500]:
    pts = rng.standard_normal((1000, d))
    q   = rng.standard_normal((1, d))
    dists = np.sqrt(np.sum((pts - q)**2, axis=1))
    print(f'd={d:4d}  ratio={dists.max()/dists.min():.4f}')
dratio max/min
260.3558
103.6317
501.7607
1001.4445
5001.1839

58. What each one costs: Milestone 2 — curse of dimensionality

Trade off

Comparison matrix

From Milestone 2 — curse of dimensionality: every row here is a choice with a cost. Fill the ratio max/min column, then say which row you would actually pick and what you give up for it.

dratio max/min
260.3558
103.6317
501.7607
1001.4445
5001.1839

59. The full program

Concept

import numpy as np
from collections import Counter
from sklearn.datasets import load_iris, make_moons
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score

# 1. kNN from scratch on iris
iris = load_iris()
X_tr, X_te, y_tr, y_te = train_test_split(iris.data, iris.target, test_size=0.2, random_state=42)
def knn_l2(X_train, y_train, X_test, k):
    preds = []
    for x in X_test:
        dists = np.sqrt(np.sum((X_train - x)**2, axis=1))
        preds.append(Counter(y_train[np.argsort(dists)[:k]]).most_common(1)[0][0])
    return np.array(preds)
print('scratch L2 k=3:', accuracy_score(y_te, knn_l2(X_tr, y_tr, X_te, 3)))

# 2. Bias-variance sweep on make_moons
X, y = make_moons(n_samples=400, noise=0.3, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=42)
for k in [1, 5, 20]:
    knn = KNeighborsClassifier(n_neighbors=k).fit(Xtr, ytr)
    print(f'k={k}: train={accuracy_score(ytr, knn.predict(Xtr)):.4f} test={accuracy_score(yte, knn.predict(Xte)):.4f}')
resultvalue
scratch L2 k=3 (iris)1.0000
k=1 train / test (moons)1.0000 / 0.8900
k=5 train / test (moons)0.9067 / 0.8900
k=20 train / test (moons)0.9033 / 0.9000

If your scratch kNN matches sklearn and your bias-variance table shows train accuracy falling as k increases while test accuracy stabilizes — you've built kNN and measured its two fundamental failure modes.

60. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

resultvalue
scratch L2 k=3 (iris)1.0000
k=1 train / test (moons)1.0000 / 0.8900
k=5 train / test (moons)0.9067 / 0.8900
k=20 train / test (moons)0.9033 / 0.9000

61. Show it off

Concept

Out loud, slides closed: explain (1) why kNN has no training step and what that means for compute cost, (2) the curse of dimensionality in terms of the max/min ratio, and (3) when to use KD-tree vs FAISS.

Stretch (homework): implement cosine-weighted kNN from scratch and compare to distance-weighted on a text-classification toy set; add FAISS approximate search for a synthetic 1M-point dataset in d=64. Next up: ensemble methods — bagging and random forests.

62. Connect it up: Lesson 69: kNN and the Curse of Dimensionality

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — kNN: majority-vote classification · Distance metrics · Curse of dimensionality · Efficient search & weighted kNN · Your turn: kNN from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

63. What you can do now

Recap

ideathe one thing to remember
kNN prediction costO(nd) — no training, all cost at inference
curse of dimensionalitymax/min ratio -> 1 as d grows; use d < 20 or reduce first
index choiceKD-tree d<20, ball tree d<100, FAISS/HNSW for d>=128 or n>=1M
k tuninglarge k = high bias / low variance; tune via cross-validation

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 69 (Week 24 — Advanced PyTorch + Full Pipeline Projects) — Barron · USAAIO Round 2 Preparation, 2026
  2. kNN accuracy, curse-of-dimensionality ratios, KD-tree vs ball-tree timings verified — numpy 2.2.6 + scikit-learn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108