Lesson 52: k-Means Clustering

USAAIO Lesson 52, from Phase 3 on unsupervised learning. It covers k-Means as assign-then-update, k-Means++ initialization, the elbow method and WCSS, the silhouette score, the EM interpretation in which the E step assigns and the M step updates the centroids, and soft k-Means as a Gaussian mixture model. It then shows k-Means failing on non-spherical clusters and DBSCAN fixing it. You build k-Means from scratch with both random and k++ initialization, implement the elbow method and the silhouette score from scratch, and verify convergence on make_blobs. The lesson runs to 29 slides.

Subject: Machine Learning · 54 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. k-Means Clustering

Title

USAAIO · Lesson 52 · Phase 3 — Unsupervised Learning

Assign to the nearest centroid; update centroids as means; repeat. Today: the full algorithm, k-Means++ init, elbow & silhouette diagnostics, and the EM view that unifies soft k-Means with GMM.

2. By the end of this lesson you can

Objectives

  1. Implement k-Means assign-then-update from scratch and trace convergence on a toy dataset
  2. Explain k-Means++ initialization and prove why it selects spread-out seeds
  3. Choose k with the elbow method (WCSS vs k) and the silhouette score formula
  4. Recast k-Means as EM (E=assignment, M=centroid update) and state what soft k-Means adds
  5. Diagnose k-Means failure on non-spherical data and fix it with DBSCAN

3. k-Means: assign & update

Section

Part 1 of 4

4. The algorithm in two moves

Concept

k-Means alternates two cheap steps until the centroids stop moving — it always converges, but only to a local minimum of WCSS.

stepactioncost
E (assign)label each point with its nearest centroid (Euclidean distance)O(nkd)
M (update)move each centroid to the mean of its assigned pointsO(nd)
checkif no label changed, stopO(n)

WCSS = within-cluster sum of squares — the objective k-Means greedily decreases each iteration. It cannot increase.

5. What each one costs: The algorithm in two moves

Trade off

Comparison matrix

From The algorithm in two moves: every row here is a choice with a cost. Fill the cost column, then say which row you would actually pick and what you give up for it.

stepactioncost
E (assign)label each point with its nearest centroid (Euclidean distance)O(nkd)
M (update)move each centroid to the mean of its assigned pointsO(nd)
checkif no label changed, stopO(n)

6. WCSS — the objective

Concept

WCSS measures total intra-cluster spread. k-Means is a coordinate-descent optimizer on this objective.

\[ \text{WCSS} = \sum_{k=1}^{K} \sum_{x_i \in C_k} \|x_i - \mu_k\|^2 \]

WCSS decreases monotonically — the E-step cannot increase it (you relabeled to the nearest center) and the M-step cannot increase it (the mean minimizes squared distance). Convergence is guaranteed in finite steps because there are finitely many label assignments.

7. Break it if you can: WCSS — the objective

Counterexample

Discussion prompt

WCSS measures total intra-cluster spread. k-Means is a coordinate-descent optimizer on this objective.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

8. Guess the shape of the answer: Trace: 6 points, 2 centroids, 2 iterations

Estimation

Predict first

Six 2D points; initial centroids at [1,1] and [5,7]. Run two iterations by hand.

Commit before you compute: what does Trace: 6 points, 2 centroids, 2 iterations come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Iter 1: labels = [0,0,0,1,1,1], c0=[1.833,2.333], c1=[4.333,5.667], WCSS=10.6667

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Points closer to [1,1] join cluster 0; those closer to [5,7] join cluster 1.

9. Trace: 6 points, 2 centroids, 2 iterations

Worked example

Six 2D points; initial centroids at [1,1] and [5,7]. Run two iterations by hand.

import numpy as np
X = np.array([[1.,1.],[1.5,2.],[3.,4.],[5.,7.],[3.5,5.],[4.5,5.]])
centers = np.array([[1.,1.],[5.,7.]])
for it in range(2):
    dists = np.array([[np.sum((x-c)**2) for c in centers] for x in X])
    labels = dists.argmin(axis=1)
    centers = np.array([X[labels==k].mean(0) for k in range(2)])
    wcss = sum(np.sum((X[labels==k]-centers[k])**2) for k in range(2))
    print(f'iter {it+1}: labels={labels}, c0={centers[0].round(3)}, c1={centers[1].round(3)}, WCSS={wcss:.4f}')

Iter 1: labels = [0,0,0,1,1,1], c0=[1.833,2.333], c1=[4.333,5.667], WCSS=10.6667

Why: Points closer to [1,1] join cluster 0; those closer to [5,7] join cluster 1. Centroids shift to their group means.

iterlabelsc0c1WCSS
0 (init)—[1.0, 1.0][5.0, 7.0]—
10,0,0,1,1,1[1.833, 2.333][4.333, 5.667]10.6667
2 (converged)0,0,0,1,1,1[1.833, 2.333][4.333, 5.667]10.6667

10. Fill in: c0 for Trace: 6 points, 2 centroids, 2 iterations

Comparison

Comparison matrix

From Trace: 6 points, 2 centroids, 2 iterations: refill the c0 column from what you know. The rest of the table is as it appeared.

iterlabelsc0c1WCSS
0 (init)—[1.0, 1.0][5.0, 7.0]—
10,0,0,1,1,1[1.833, 2.333][4.333, 5.667]10.6667
2 (converged)0,0,0,1,1,1[1.833, 2.333][4.333, 5.667]10.6667

11. k-Means++ initialization

Section

Part 2 of 4

12. Why random init fails

Concept

With random initialization, two initial centroids may land in the same true cluster — the algorithm then wastes iterations and often converges to a bad local minimum.

initWCSS (make_blobs, 3 clusters)converges at
random (seed=0)5693.43iter 10
k-Means++ (seed=0)566.86iter 2
sklearn (optimal)566.86—

Random init produced WCSS 10× worse — it stuck in a local minimum. k-Means++ picks seeds that are spread out by design.

13. By analogy: Why random init fails

Analogy

Discussion prompt

Explain Why random init fails by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

With random initialization, two initial centroids may land in the same true cluster — the algorithm then wastes iterations and often converges to a bad local minimum.

14. k-Means++ seeding rule

Concept

k-Means++ samples each successive seed proportional to d(x, nearest chosen centroid)² — far points get exponentially more probability.

\[ P(x_i \text{ chosen next}) = \frac{d(x_i,\, C)^2}{\sum_j d(x_j,\, C)^2} \]

After seed 1 = [0,0], the 1D example below shows: x=[11,0] has probability 0.4276, while x=[1,0] has only 0.0035. The farthest point is ~120× more likely to be chosen.

15. Guess the shape of the answer: k-Means++ selection probabilities

Estimation

Predict first

Six collinear points at 0,1,5,6,10,11. First seed = [0,0]. Compute the d² probabilities for the second seed.

Commit before you compute: what does k-Means++ selection probabilities come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: x=11 has P=0.4276; x=10 has P=0.3534; x=1 has P=0.0035

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. d² grows quadratically with distance — the seeding rule over-represents outlying points, spreading initial seeds across the data.

16. k-Means++ selection probabilities

Worked example

Six collinear points at 0,1,5,6,10,11. First seed = [0,0]. Compute the d² probabilities for the second seed.

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.],[10.,0.],[11.,0.]])
c1 = X[0]  # first seed chosen uniformly
d2 = np.array([np.sum((x-c1)**2) for x in X])
probs = d2 / d2.sum()
for xi, di, pi in zip(X[:,0], d2, probs):
    print(f'x={xi:4.0f}  d^2={di:5.0f}  P={pi:.4f}')

x=11 has P=0.4276; x=10 has P=0.3534; x=1 has P=0.0035

Why: d² grows quadratically with distance — the seeding rule over-represents outlying points, spreading initial seeds across the data.

xd²(from c1=0)P(chosen as c2)
000.0000
110.0035
5250.0883
6360.1272
101000.3534
111210.4276

17. Watch it run: k-Means++ selection probabilities

Pattern

Step through it

Step through k-Means++ selection probabilities one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: x is 0
  2. Step 2: x is 1
  3. Step 3: x is 5
  4. Step 4: x is 6
  5. Step 5: x is 10
  6. Step 6: x is 11

18. Something is wrong here: using Euclidean distance in d², not squared distance

Anomaly

Predict first

A student writes this, and it looks reasonable:

Implement k-Means++ by weighting each candidate proportional to ‖x − nearest_c‖ (the plain Euclidean distance).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that…

Weight proportional to d² — the squared distance to the nearest already-chosen centroid.

Why: The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that k-Means++ provides.

19. Trap: using Euclidean distance in d², not squared distance

Trap

The trap

Implement k-Means++ by weighting each candidate proportional to ‖x − nearest_c‖ (the plain Euclidean distance).

P(x) ∝ d instead of d²

Why: The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that k-Means++ provides.

The fix

Weight proportional to d² — the squared distance to the nearest already-chosen centroid.

P(x) ∝ ‖x − nearest_c‖²

Why: The squared-distance weighting is the definition of k-Means++; it provides a O(log k) expected-WCSS approximation guarantee. Plain distance loses this bound.

20. Break it on purpose: using Euclidean distance in d², not squared…

Break the constraint

Discussion prompt

The rule this trap just fixed:

The squared-distance weighting is the definition of k-Means++; it provides a O(log k) expected-WCSS approximation guarantee. Plain distance loses this bound.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that k-Means++ provides.

21. Elbow & silhouette

Section

Part 3 of 4

22. Elbow method: choosing k

Concept

Plot WCSS vs k. WCSS always decreases as k increases (more clusters → tighter fit) — pick k at the elbow where the marginal gain flattens.

kWCSS (make_blobs 3-cluster)
120402.34
25763.46
3566.86
4496.43
5427.13
6375.03
7308.20

The sharp drop from k=2 (5763) to k=3 (567) — a 10× decrease — is the elbow. k=3→4 drops only 70 (12%). Elbow correctly identifies the true 3 clusters.

23. Watch it run: Elbow method: choosing k

Pattern

Step through it

Step through Elbow method: choosing k one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: k is 1
  2. Step 2: k is 2
  3. Step 3: k is 3
  4. Step 4: k is 4
  5. Step 5: k is 5
  6. Step 6: k is 6
  7. Step 7: k is 7

24. Silhouette score

Concept

Silhouette measures how tightly a point sits within its cluster vs how far it is from the nearest other cluster. Range: [-1, 1]; higher is better.

\[ s(i) = \frac{b(i) - a(i)}{\max(a(i),\, b(i))} \]

a(i) = mean distance to all other points in the same cluster (intra-cluster cohesion). b(i) = mean distance to all points in the nearest different cluster (separation).

25. By analogy: Silhouette score

Analogy

Discussion prompt

Explain Silhouette score by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Silhouette measures how tightly a point sits within its cluster vs how far it is from the nearest other cluster. Range: [-1, 1]; higher is better.

26. Guess the shape of the answer: Silhouette: compute s(i) for one point

Estimation

Predict first

Point 0 in the make_blobs dataset is assigned to cluster 1 after k-Means (seed 42). Compute its silhouette score from scratch.

Commit before you compute: what does Silhouette: compute s(i) for one point come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: a=1.4486 (tight cluster), b=15.5553 (far from next cluster), s=0.9069

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. b >> a means the point is deep inside its cluster and far from any other — near-perfect silhouette of 0.91.

27. Silhouette: compute s(i) for one point

Worked example

Point 0 in the make_blobs dataset is assigned to cluster 1 after k-Means (seed 42). Compute its silhouette score from scratch.

import numpy as np
from sklearn.datasets import make_blobs
from sklearn.cluster import KMeans
X, _ = make_blobs(n_samples=300, centers=3, cluster_std=1.0, random_state=42)
km = KMeans(n_clusters=3, random_state=42, n_init=10).fit(X)
labels = km.labels_
x0, c0 = X[0], labels[0]
a = np.mean(np.sqrt(np.sum((X[labels==c0] - x0)**2, axis=1)))
b = min(np.mean(np.sqrt(np.sum((X[labels==k] - x0)**2, axis=1)))
        for k in range(3) if k != c0)
print(f'a={a:.4f}, b={b:.4f}, s={( b-a)/max(a,b):.4f}')

a=1.4486 (tight cluster), b=15.5553 (far from next cluster), s=0.9069

Why: b >> a means the point is deep inside its cluster and far from any other — near-perfect silhouette of 0.91.

quantityvaluemeaning
a (intra-cluster mean dist)1.4486tight — close to cluster peers
b (nearest-cluster mean dist)15.5553well-separated from next cluster
s(i) = (b−a)/max(a,b)0.9069near-perfect assignment

28. Work backwards from the answer: Silhouette: compute s(i) for one point

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

a=1.4486 (tight cluster), b=15.5553 (far from next cluster), s=0.9069

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Point 0 in the make_blobs dataset is assigned to cluster 1 after k-Means (seed 42). Compute its silhouette score from scratch.

29. k-Means as EM

Concept

k-Means is a hard-assignment special case of the Expectation-Maximization algorithm (Lesson 47).

EM stepk-Means actionGMM action
E-stephard-assign each point to nearest centroid (0/1)soft-assign: compute responsibility r_ik ∈ [0,1]
M-stepupdate centroid = mean of assigned pointsupdate mean, covariance, mixture weight
objectiveminimize WCSSmaximize log-likelihood

Soft k-Means replaces the hard 0/1 assignment with fractional responsibilities — that is exactly a GMM with spherical, equal-variance Gaussians.

30. Fill in: k-Means action for k-Means as EM

Comparison

Comparison matrix

From k-Means as EM: refill the k-Means action column from what you know. The rest of the table is as it appeared.

EM stepk-Means actionGMM action
E-stephard-assign each point to nearest centroid (0/1)soft-assign: compute responsibility r_ik ∈ [0,1]
M-stepupdate centroid = mean of assigned pointsupdate mean, covariance, mixture weight
objectiveminimize WCSSmaximize log-likelihood

31. Something is wrong here: k-Means on non-spherical clusters

Anomaly

Predict first

A student writes this, and it looks reasonable:

Apply k-Means (k=2) to a two-moon dataset — it will find 2 clusters.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: k-Means uses Euclidean distance to a centroid — it can only produce convex, roughly spherical clusters.

Use DBSCAN (density-based; no shape assumption) with eps=0.25, min_samples=5.

Why: k-Means uses Euclidean distance to a centroid — it can only produce convex, roughly spherical clusters. The two moons are non-convex and interlocked; k-Means cuts them down the middle and gets 74.5% accuracy.

32. Trap: k-Means on non-spherical clusters

Trap

The trap

Apply k-Means (k=2) to a two-moon dataset — it will find 2 clusters.

k-Means reports 2 clusters; silhouette = 0.4815

Why: k-Means uses Euclidean distance to a centroid — it can only produce convex, roughly spherical clusters. The two moons are non-convex and interlocked; k-Means cuts them down the middle and gets 74.5% accuracy.

The fix

Use DBSCAN (density-based; no shape assumption) with eps=0.25, min_samples=5.

DBSCAN reports 2 clusters, 0 noise; silhouette = 0.3187; accuracy = 100%

Why: DBSCAN connects dense neighborhoods regardless of shape — it perfectly recovers the two interlocked moons. The silhouette is lower because the metric itself assumes convex clusters, but the true assignment is correct.

33. Which of these survive contact with Lesson 52: k-Means Clustering?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
WCSS = within-cluster sum of squares — the objective k-Means greedily decreases each iteration. It cannot increase.; WCSS measures total intra-cluster spread. k-Means is a coordinate-descent optimizer on this objective.; With random initialization, two initial centroids may land in the same true cluster — the algorithm then wastes iterations and often converges to a bad local minimum.
Breaks
Implement k-Means++ by weighting each candidate proportional to ‖x − nearest_c‖ (the plain Euclidean distance).; Apply k-Means (k=2) to a two-moon dataset — it will find 2 clusters.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 52: k-Means Clustering puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

34. Rebuild the recipe: k-Means recipe

Ranking

Put in order

These are the steps of k-Means recipe, scrambled. Put them back in order before the next slide shows you.

  1. Init: pick k seeds via k-Means++ (d² weighting)
  2. E-step: assign each point to nearest centroid — labels = argmin_k ‖x − μk‖²
  3. M-step: update each centroid — μk = mean({xi : label=k})
  4. Repeat until labels don't change (WCSS only decreases; local minimum guaranteed)
  5. Choose k: elbow on WCSS vs k; confirm with silhouette score
  6. Sanity-check shape assumption: if clusters are non-spherical, use DBSCAN or GMM

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

35. k-Means recipe

Pattern

  1. Init: pick k seeds via k-Means++ (d² weighting)
  2. E-step: assign each point to nearest centroid — labels = argmin_k ‖x − μk‖²
  3. M-step: update each centroid — μk = mean({xi : label=k})
  4. Repeat until labels don't change (WCSS only decreases; local minimum guaranteed)
  5. Choose k: elbow on WCSS vs k; confirm with silhouette score
  6. Sanity-check shape assumption: if clusters are non-spherical, use DBSCAN or GMM

36. Where does it stop working: k-Means recipe

Edge cases

Discussion prompt

k-Means recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Init: pick k seeds via k-Means++ (d² weighting)
  2. E-step: assign each point to nearest centroid — labels = argmin_k ‖x − μk‖²
  3. M-step: update each centroid — μk = mean({xi : label=k})
  4. Repeat until labels don't change (WCSS only decreases; local minimum guaranteed)
  5. Choose k: elbow on WCSS vs k; confirm with silhouette score
  6. Sanity-check shape assumption: if clusters are non-spherical, use DBSCAN or GMM

37. Rule out three: Check yourself — convergence

Elimination

Eliminate the wrong options

k-Means is guaranteed to converge. The solution it converges to is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. a local minimum of WCSS — not necessarily the global minimum
  • B. the global minimum of WCSS — guaranteed by the E/M alternation
  • C. a saddle point of the log-likelihood
  • D. the global minimum only when initialized with k-Means++

Survives elimination: A

Why: Each E-step and M-step can only decrease or maintain WCSS, so the sequence of WCSS values converges. But the algorithm is coordinate descent on a non-convex objective — the limit depends on initialization and is in general a local minimum. Our random-init run hit WCSS 5693 vs the global 567.

38. Check yourself — convergence

Check

k-Means always converges — but to what?

Check your understanding

k-Means is guaranteed to converge. The solution it converges to is:

  • A. a local minimum of WCSS — not necessarily the global minimum (correct)
  • B. the global minimum of WCSS — guaranteed by the E/M alternation
  • C. a saddle point of the log-likelihood
  • D. the global minimum only when initialized with k-Means++

Answer: A

Why: Each E-step and M-step can only decrease or maintain WCSS, so the sequence of WCSS values converges. But the algorithm is coordinate descent on a non-convex objective — the limit depends on initialization and is in general a local minimum. Our random-init run hit WCSS 5693 vs the global 567.

Why B tempts people
The WCSS surface is non-convex for k>1; coordinate descent on a non-convex function cannot guarantee the global minimum.
Why C tempts people
k-Means minimizes WCSS, not the log-likelihood (that's EM/GMM). k-Means has no log-likelihood objective.
Why D tempts people
k-Means++ gives a better expected initialization (O(log k) bound) but does not turn the algorithm into a global optimizer — it still converges to a local minimum.

39. Answer it before you see the options: Check yourself — silhouette formula

Prediction

Predict first

Point i has a(i)=1.45 (mean intra-cluster distance) and b(i)=15.56 (nearest-cluster mean distance). Its silhouette score s(i) is closest to:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0.91

Why: s(i) = (b−a)/max(a,b) = (15.56−1.45)/max(1.45,15.56) = 14.11/15.56 ≈ 0.907 ≈ 0.91. The point is well inside its cluster and far from the next — high positive score.

40. Check yourself — silhouette formula

Check

A point has a=1.45 and b=15.56. What is its silhouette score?

Check your understanding

Point i has a(i)=1.45 (mean intra-cluster distance) and b(i)=15.56 (nearest-cluster mean distance). Its silhouette score s(i) is closest to:

  • A. 0.91 (correct)
  • B. 0.09
  • C. −0.91
  • D. 14.11

Answer: A

Why: s(i) = (b−a)/max(a,b) = (15.56−1.45)/max(1.45,15.56) = 14.11/15.56 ≈ 0.907 ≈ 0.91. The point is well inside its cluster and far from the next — high positive score.

Why B tempts people
0.09 = a/b, not (b−a)/max(a,b). Confusing the ratio direction.
Why C tempts people
Negative silhouette means a > b (point is closer to a different cluster than its own). Here b >> a, so the score is strongly positive.
Why D tempts people
14.11 = b−a is the numerator, not the normalized score. Forgetting to divide by max(a,b).

41. Rule out three: Check yourself — EM connection

Elimination

Eliminate the wrong options

Soft k-Means differs from hard k-Means in that it:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. assigns fractional cluster memberships (responsibilities) instead of 0/1 hard labels
  • B. uses Manhattan distance instead of Euclidean distance
  • C. allows k to change during training
  • D. guarantees convergence to the global optimum

Survives elimination: A

Why: In the EM view of k-Means, the E-step assigns a responsibility r_ik ∈ {0,1}. Soft k-Means relaxes this to r_ik ∈ [0,1] via a softmax — recovering the GMM with equal spherical covariances. Both still alternate E and M steps.

42. Check yourself — EM connection

Check

What distinguishes soft k-Means from hard k-Means?

Check your understanding

Soft k-Means differs from hard k-Means in that it:

  • A. assigns fractional cluster memberships (responsibilities) instead of 0/1 hard labels (correct)
  • B. uses Manhattan distance instead of Euclidean distance
  • C. allows k to change during training
  • D. guarantees convergence to the global optimum

Answer: A

Why: In the EM view of k-Means, the E-step assigns a responsibility r_ik ∈ {0,1}. Soft k-Means relaxes this to r_ik ∈ [0,1] via a softmax — recovering the GMM with equal spherical covariances. Both still alternate E and M steps.

Why B tempts people
Both hard and soft k-Means use squared Euclidean distance in the standard formulation. Manhattan distance is a separate variant.
Why C tempts people
k is fixed in both formulations. Algorithms that adapt k (like X-means or DBSCAN) are separate.
Why D tempts people
Neither hard nor soft k-Means guarantees the global optimum — both are coordinate-descent methods on a non-convex objective.

43. Your turn: build k-Means from scratch

Section

Project

44. Project: k-Means from scratch

Concept

Build k-Means with both random and k++ initialization on make_blobs(centers=3). Implement the elbow method and silhouette from scratch, then demonstrate k-Means failing on moons.

#requirementtarget value
1k++ init, k=3, make_blobsWCSS=566.86, silhouette=0.848
2elbow WCSS for k=1..7big drop at k=2→3 (5763→567)
3silhouette from scratch, point 0s=0.907
4k-Means on moons (k=2)accuracy=74.5%; silhouette=0.482
5DBSCAN on moons (eps=0.25)accuracy=100%; 2 clusters

Build rules: no sklearn.cluster.KMeans in milestones 1-3 (except to check). For DBSCAN you may use sklearn.cluster.DBSCAN.

45. Break it if you can: Project: k-Means from scratch

Counterexample

Discussion prompt

Build k-Means with both random and k++ initialization on make_blobs(centers=3). Implement the elbow method and silhouette from scratch, then demonstrate k-Means failing on moons.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: no sklearn.cluster.KMeans in milestones 1-3 (except to check). For DBSCAN you may use sklearn.cluster.DBSCAN.

46. Milestone 1 — k-Means++ from scratch

Worked example

Your turn: implement kmeans_pp(X, k) — k++ seeding then assign-update loop. Predict: does k++ converge faster than random init?

Hint: seed loop — d2 = min distance² to any chosen center; rng.choice(n, p=d2/d2.sum()).

import numpy as np
from sklearn.datasets import make_blobs
X, _ = make_blobs(n_samples=300, centers=3, cluster_std=1.0, random_state=42)

def kmeans_pp(X, k, seed=42):
    rng = np.random.default_rng(seed)
    c = [X[rng.integers(len(X))]]
    for _ in range(k-1):
        d2 = np.array([min(np.sum((x-ci)**2) for ci in c) for x in X])
        c.append(X[rng.choice(len(X), p=d2/d2.sum())])
    centers = np.array(c)
    for it in range(100):
        labels = np.array([np.argmin([np.sum((x-ci)**2) for ci in centers]) for x in X])
        new_c = np.array([X[labels==j].mean(0) for j in range(k)])
        if np.allclose(new_c, centers):
            print(f'converged iter {it+1}'); break
        centers = new_c
    return labels, centers

labels, centers = kmeans_pp(X, 3)
wcss = sum(np.sum((X[labels==k]-centers[k])**2) for k in range(3))
print(f'WCSS={wcss:.4f}')
runWCSSconverged at iter
random init (seed=0)5693.4310
k-Means++ (seed=42)566.862
sklearn reference566.86—

47. Milestone 2 — elbow method from scratch

Worked example

Your turn: run k-Means++ for k=1 to 7 and compute WCSS each time. Predict where the elbow is.

Hint: call kmeans_pp(X, k) for each k; compute sum((X[labels==j] - centers[j])**2 for j in range(k)).

from sklearn.cluster import KMeans  # check with sklearn
wcss_k = []
for k in range(1, 8):
    km = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
    wcss_k.append(km.inertia_)
    print(f'k={k}: WCSS={km.inertia_:.2f}')
kWCSSdrop from k-1
120402.34—
25763.4614638.88
3566.865196.60 ← elbow
4496.4370.43
5427.1369.30
6375.0352.10
7308.2066.83

48. Milestone 3 — silhouette from scratch

Worked example

Your turn: for the k=3 labels from Milestone 1, compute s(i) for every i without sklearn's silhouette_score. Check your mean vs 0.848.

Hint: a[i] = mean(dist(X[i], X[same cluster])); b[i] = min over other clusters of mean(dist(X[i], X[cluster k])).

from sklearn.metrics import silhouette_score
# scratch implementation
def sil_scratch(X, labels):
    k = len(set(labels))
    scores = []
    for i, x in enumerate(X):
        c = labels[i]
        a = np.mean(np.sqrt(np.sum((X[labels==c]-x)**2, axis=1)))
        bs = [np.mean(np.sqrt(np.sum((X[labels==j]-x)**2, axis=1)))
              for j in range(k) if j != c]
        b = min(bs)
        scores.append((b-a)/max(a,b))
    return np.mean(scores)

print(f'scratch silhouette: {sil_scratch(X, labels):.4f}')
print(f'sklearn silhouette: {silhouette_score(X, labels):.4f}')
methodmean silhouette
scratch0.8480
sklearn0.8480
match✓

49. What each one costs: Milestone 3 — silhouette from scratch

Trade off

Comparison matrix

From Milestone 3 — silhouette from scratch: every row here is a choice with a cost. Fill the mean silhouette column, then say which row you would actually pick and what you give up for it.

methodmean silhouette
scratch0.8480
sklearn0.8480
match✓

50. The full program

Concept

import numpy as np
from sklearn.datasets import make_blobs, make_moons
from sklearn.cluster import DBSCAN
from sklearn.metrics import silhouette_score

X, _ = make_blobs(n_samples=300, centers=3, cluster_std=1.0, random_state=42)

def kmeans_pp(X, k, seed=42):
    rng = np.random.default_rng(seed)
    c = [X[rng.integers(len(X))]]
    for _ in range(k-1):
        d2 = np.array([min(np.sum((x-ci)**2) for ci in c) for x in X])
        c.append(X[rng.choice(len(X), p=d2/d2.sum())])
    centers = np.array(c)
    for _ in range(100):
        labels = np.array([np.argmin([np.sum((x-ci)**2) for ci in centers]) for x in X])
        new_c = np.array([X[labels==j].mean(0) for j in range(k)])
        if np.allclose(new_c, centers): break
        centers = new_c
    return labels, centers

labels, centers = kmeans_pp(X, 3)
wcss = sum(np.sum((X[labels==k]-centers[k])**2) for k in range(3))
print(f'WCSS={wcss:.4f}, sil={silhouette_score(X, labels):.4f}')

X_m, y_m = make_moons(n_samples=200, noise=0.07, random_state=42)
lb, _ = kmeans_pp(X_m, 2)
db = DBSCAN(eps=0.25, min_samples=5).fit_predict(X_m)
print(f'moons k-Means sil={silhouette_score(X_m, lb):.4f}')
print(f'moons DBSCAN sil={silhouette_score(X_m, db):.4f}, clusters={len(set(db))}')
outputvalue
WCSS (blobs, k=3)566.8596
silhouette (blobs)0.8480
k-Means sil (moons)0.4815
DBSCAN sil (moons)0.3187
DBSCAN clusters2 (correct)

If your WCSS and silhouette match — you have the complete k-Means toolkit. The lower DBSCAN silhouette does not mean it's worse; it means the silhouette metric itself favors convex clusters.

51. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
WCSS (blobs, k=3)566.8596
silhouette (blobs)0.8480
k-Means sil (moons)0.4815
DBSCAN sil (moons)0.3187
DBSCAN clusters2 (correct)

52. Show it off

Concept

Out loud, slides closed: (1) walk through one full E + M step without looking at code, (2) explain why k-Means++ picks seeds that are far apart, (3) state the silhouette formula and explain what a score near 1 vs near -1 means.

Stretch (homework from lesson plan): plot the k-Means convergence curve (WCSS vs iteration) for random vs k++ init on the same axes; implement the elbow method as a function; prove algebraically that k-Means M-step minimizes WCSS given fixed labels (i.e., the mean minimizes sum of squared deviations). Next up: Gaussian Mixture Models — the soft-assignment generalization of k-Means.

53. Connect it up: Lesson 52: k-Means Clustering

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — k-Means: assign & update · k-Means++ initialization · Elbow & silhouette · Your turn: build k-Means from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

54. What you can do now

Recap

ideathe one thing to remember
k-Means objectiveminimize WCSS — local minimum, not global
k-Means++ initP(next seed) ∝ d² to nearest chosen center
elbow methodpick k at the sharpest WCSS drop
silhouette score(b−a)/max(a,b); range [−1, 1], higher = better
EM viewE=hard assign, M=centroid mean; soft→GMM

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 52 (Phase 3 — Unsupervised Learning: k-Means) — Barron · USAAIO Round 2 Preparation, 2026
  2. k-Means scratch, k-Means++, elbow, silhouette, DBSCAN on make_blobs/make_moons verified — numpy 2.2.6 + sklearn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108