USAAIO Lesson 52, from Phase 3 on unsupervised learning. It covers k-Means as assign-then-update, k-Means++ initialization, the elbow method and WCSS, the silhouette score, the EM interpretation in which the E step assigns and the M step updates the centroids, and soft k-Means as a Gaussian mixture model. It then shows k-Means failing on non-spherical clusters and DBSCAN fixing it. You build k-Means from scratch with both random and k++ initialization, implement the elbow method and the silhouette score from scratch, and verify convergence on make_blobs. The lesson runs to 29 slides.
Subject: Machine Learning · 54 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 52 · Phase 3 — Unsupervised Learning
Assign to the nearest centroid; update centroids as means; repeat. Today: the full algorithm, k-Means++ init, elbow & silhouette diagnostics, and the EM view that unifies soft k-Means with GMM.
Objectives
Section
Part 1 of 4
Concept
k-Means alternates two cheap steps until the centroids stop moving — it always converges, but only to a local minimum of WCSS.
| step | action | cost |
|---|---|---|
| E (assign) | label each point with its nearest centroid (Euclidean distance) | O(nkd) |
| M (update) | move each centroid to the mean of its assigned points | O(nd) |
| check | if no label changed, stop | O(n) |
WCSS = within-cluster sum of squares — the objective k-Means greedily decreases each iteration. It cannot increase.
Trade off
Comparison matrix
From The algorithm in two moves: every row here is a choice with a cost. Fill the cost column, then say which row you would actually pick and what you give up for it.
| step | action | cost |
|---|---|---|
| E (assign) | label each point with its nearest centroid (Euclidean distance) | O(nkd) |
| M (update) | move each centroid to the mean of its assigned points | O(nd) |
| check | if no label changed, stop | O(n) |
Concept
WCSS measures total intra-cluster spread. k-Means is a coordinate-descent optimizer on this objective.
\[ \text{WCSS} = \sum_{k=1}^{K} \sum_{x_i \in C_k} \|x_i - \mu_k\|^2 \]
WCSS decreases monotonically — the E-step cannot increase it (you relabeled to the nearest center) and the M-step cannot increase it (the mean minimizes squared distance). Convergence is guaranteed in finite steps because there are finitely many label assignments.
Counterexample
Discussion prompt
WCSS measures total intra-cluster spread. k-Means is a coordinate-descent optimizer on this objective.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Six 2D points; initial centroids at [1,1] and [5,7]. Run two iterations by hand.
Commit before you compute: what does Trace: 6 points, 2 centroids, 2 iterations come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Iter 1: labels = [0,0,0,1,1,1], c0=[1.833,2.333], c1=[4.333,5.667], WCSS=10.6667
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Points closer to [1,1] join cluster 0; those closer to [5,7] join cluster 1.
Worked example
Six 2D points; initial centroids at [1,1] and [5,7]. Run two iterations by hand.
import numpy as np
X = np.array([[1.,1.],[1.5,2.],[3.,4.],[5.,7.],[3.5,5.],[4.5,5.]])
centers = np.array([[1.,1.],[5.,7.]])
for it in range(2):
dists = np.array([[np.sum((x-c)**2) for c in centers] for x in X])
labels = dists.argmin(axis=1)
centers = np.array([X[labels==k].mean(0) for k in range(2)])
wcss = sum(np.sum((X[labels==k]-centers[k])**2) for k in range(2))
print(f'iter {it+1}: labels={labels}, c0={centers[0].round(3)}, c1={centers[1].round(3)}, WCSS={wcss:.4f}')Iter 1: labels = [0,0,0,1,1,1], c0=[1.833,2.333], c1=[4.333,5.667], WCSS=10.6667
Why: Points closer to [1,1] join cluster 0; those closer to [5,7] join cluster 1. Centroids shift to their group means.
| iter | labels | c0 | c1 | WCSS |
|---|---|---|---|---|
| 0 (init) | — | [1.0, 1.0] | [5.0, 7.0] | — |
| 1 | 0,0,0,1,1,1 | [1.833, 2.333] | [4.333, 5.667] | 10.6667 |
| 2 (converged) | 0,0,0,1,1,1 | [1.833, 2.333] | [4.333, 5.667] | 10.6667 |
Comparison
Comparison matrix
From Trace: 6 points, 2 centroids, 2 iterations: refill the c0 column from what you know. The rest of the table is as it appeared.
| iter | labels | c0 | c1 | WCSS |
|---|---|---|---|---|
| 0 (init) | — | [1.0, 1.0] | [5.0, 7.0] | — |
| 1 | 0,0,0,1,1,1 | [1.833, 2.333] | [4.333, 5.667] | 10.6667 |
| 2 (converged) | 0,0,0,1,1,1 | [1.833, 2.333] | [4.333, 5.667] | 10.6667 |
Section
Part 2 of 4
Concept
With random initialization, two initial centroids may land in the same true cluster — the algorithm then wastes iterations and often converges to a bad local minimum.
| init | WCSS (make_blobs, 3 clusters) | converges at |
|---|---|---|
| random (seed=0) | 5693.43 | iter 10 |
| k-Means++ (seed=0) | 566.86 | iter 2 |
| sklearn (optimal) | 566.86 | — |
Random init produced WCSS 10× worse — it stuck in a local minimum. k-Means++ picks seeds that are spread out by design.
Analogy
Discussion prompt
Explain Why random init fails by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
With random initialization, two initial centroids may land in the same true cluster — the algorithm then wastes iterations and often converges to a bad local minimum.
Concept
k-Means++ samples each successive seed proportional to d(x, nearest chosen centroid)² — far points get exponentially more probability.
\[ P(x_i \text{ chosen next}) = \frac{d(x_i,\, C)^2}{\sum_j d(x_j,\, C)^2} \]
After seed 1 = [0,0], the 1D example below shows: x=[11,0] has probability 0.4276, while x=[1,0] has only 0.0035. The farthest point is ~120× more likely to be chosen.
Estimation
Predict first
Six collinear points at 0,1,5,6,10,11. First seed = [0,0]. Compute the d² probabilities for the second seed.
Commit before you compute: what does k-Means++ selection probabilities come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: x=11 has P=0.4276; x=10 has P=0.3534; x=1 has P=0.0035
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. d² grows quadratically with distance — the seeding rule over-represents outlying points, spreading initial seeds across the data.
Worked example
Six collinear points at 0,1,5,6,10,11. First seed = [0,0]. Compute the d² probabilities for the second seed.
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.],[10.,0.],[11.,0.]])
c1 = X[0] # first seed chosen uniformly
d2 = np.array([np.sum((x-c1)**2) for x in X])
probs = d2 / d2.sum()
for xi, di, pi in zip(X[:,0], d2, probs):
print(f'x={xi:4.0f} d^2={di:5.0f} P={pi:.4f}')x=11 has P=0.4276; x=10 has P=0.3534; x=1 has P=0.0035
Why: d² grows quadratically with distance — the seeding rule over-represents outlying points, spreading initial seeds across the data.
| x | d²(from c1=0) | P(chosen as c2) |
|---|---|---|
| 0 | 0 | 0.0000 |
| 1 | 1 | 0.0035 |
| 5 | 25 | 0.0883 |
| 6 | 36 | 0.1272 |
| 10 | 100 | 0.3534 |
| 11 | 121 | 0.4276 |
Pattern
Step through it
Step through k-Means++ selection probabilities one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Implement k-Means++ by weighting each candidate proportional to ‖x − nearest_c‖ (the plain Euclidean distance).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that…
Weight proportional to d² — the squared distance to the nearest already-chosen centroid.
Why: The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that k-Means++ provides.
Trap
Implement k-Means++ by weighting each candidate proportional to ‖x − nearest_c‖ (the plain Euclidean distance).
P(x) ∝ d instead of d²
Why: The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that k-Means++ provides.
Weight proportional to d² — the squared distance to the nearest already-chosen centroid.
P(x) ∝ ‖x − nearest_c‖²
Why: The squared-distance weighting is the definition of k-Means++; it provides a O(log k) expected-WCSS approximation guarantee. Plain distance loses this bound.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The squared-distance weighting is the definition of k-Means++; it provides a O(log k) expected-WCSS approximation guarantee. Plain distance loses this bound.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
The plain distance still favors far points, so it 'looks right' locally — but it underweights extreme outliers relative to the theoretically optimal schedule, degrading the O(log k) approximation guarantee that k-Means++ provides.
Section
Part 3 of 4
Concept
Plot WCSS vs k. WCSS always decreases as k increases (more clusters → tighter fit) — pick k at the elbow where the marginal gain flattens.
| k | WCSS (make_blobs 3-cluster) |
|---|---|
| 1 | 20402.34 |
| 2 | 5763.46 |
| 3 | 566.86 |
| 4 | 496.43 |
| 5 | 427.13 |
| 6 | 375.03 |
| 7 | 308.20 |
The sharp drop from k=2 (5763) to k=3 (567) — a 10× decrease — is the elbow. k=3→4 drops only 70 (12%). Elbow correctly identifies the true 3 clusters.
Pattern
Step through it
Step through Elbow method: choosing k one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Silhouette measures how tightly a point sits within its cluster vs how far it is from the nearest other cluster. Range: [-1, 1]; higher is better.
\[ s(i) = \frac{b(i) - a(i)}{\max(a(i),\, b(i))} \]
a(i) = mean distance to all other points in the same cluster (intra-cluster cohesion). b(i) = mean distance to all points in the nearest different cluster (separation).
Analogy
Discussion prompt
Explain Silhouette score by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Silhouette measures how tightly a point sits within its cluster vs how far it is from the nearest other cluster. Range: [-1, 1]; higher is better.
Estimation
Predict first
Point 0 in the make_blobs dataset is assigned to cluster 1 after k-Means (seed 42). Compute its silhouette score from scratch.
Commit before you compute: what does Silhouette: compute s(i) for one point come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: a=1.4486 (tight cluster), b=15.5553 (far from next cluster), s=0.9069
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. b >> a means the point is deep inside its cluster and far from any other — near-perfect silhouette of 0.91.
Worked example
Point 0 in the make_blobs dataset is assigned to cluster 1 after k-Means (seed 42). Compute its silhouette score from scratch.
import numpy as np
from sklearn.datasets import make_blobs
from sklearn.cluster import KMeans
X, _ = make_blobs(n_samples=300, centers=3, cluster_std=1.0, random_state=42)
km = KMeans(n_clusters=3, random_state=42, n_init=10).fit(X)
labels = km.labels_
x0, c0 = X[0], labels[0]
a = np.mean(np.sqrt(np.sum((X[labels==c0] - x0)**2, axis=1)))
b = min(np.mean(np.sqrt(np.sum((X[labels==k] - x0)**2, axis=1)))
for k in range(3) if k != c0)
print(f'a={a:.4f}, b={b:.4f}, s={( b-a)/max(a,b):.4f}')a=1.4486 (tight cluster), b=15.5553 (far from next cluster), s=0.9069
Why: b >> a means the point is deep inside its cluster and far from any other — near-perfect silhouette of 0.91.
| quantity | value | meaning |
|---|---|---|
| a (intra-cluster mean dist) | 1.4486 | tight — close to cluster peers |
| b (nearest-cluster mean dist) | 15.5553 | well-separated from next cluster |
| s(i) = (b−a)/max(a,b) | 0.9069 | near-perfect assignment |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
a=1.4486 (tight cluster), b=15.5553 (far from next cluster), s=0.9069
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Point 0 in the make_blobs dataset is assigned to cluster 1 after k-Means (seed 42). Compute its silhouette score from scratch.
Concept
k-Means is a hard-assignment special case of the Expectation-Maximization algorithm (Lesson 47).
| EM step | k-Means action | GMM action |
|---|---|---|
| E-step | hard-assign each point to nearest centroid (0/1) | soft-assign: compute responsibility r_ik ∈ [0,1] |
| M-step | update centroid = mean of assigned points | update mean, covariance, mixture weight |
| objective | minimize WCSS | maximize log-likelihood |
Soft k-Means replaces the hard 0/1 assignment with fractional responsibilities — that is exactly a GMM with spherical, equal-variance Gaussians.
Comparison
Comparison matrix
From k-Means as EM: refill the k-Means action column from what you know. The rest of the table is as it appeared.
| EM step | k-Means action | GMM action |
|---|---|---|
| E-step | hard-assign each point to nearest centroid (0/1) | soft-assign: compute responsibility r_ik ∈ [0,1] |
| M-step | update centroid = mean of assigned points | update mean, covariance, mixture weight |
| objective | minimize WCSS | maximize log-likelihood |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Apply k-Means (k=2) to a two-moon dataset — it will find 2 clusters.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: k-Means uses Euclidean distance to a centroid — it can only produce convex, roughly spherical clusters.
Use DBSCAN (density-based; no shape assumption) with eps=0.25, min_samples=5.
Why: k-Means uses Euclidean distance to a centroid — it can only produce convex, roughly spherical clusters. The two moons are non-convex and interlocked; k-Means cuts them down the middle and gets 74.5% accuracy.
Trap
Apply k-Means (k=2) to a two-moon dataset — it will find 2 clusters.
k-Means reports 2 clusters; silhouette = 0.4815
Why: k-Means uses Euclidean distance to a centroid — it can only produce convex, roughly spherical clusters. The two moons are non-convex and interlocked; k-Means cuts them down the middle and gets 74.5% accuracy.
Use DBSCAN (density-based; no shape assumption) with eps=0.25, min_samples=5.
DBSCAN reports 2 clusters, 0 noise; silhouette = 0.3187; accuracy = 100%
Why: DBSCAN connects dense neighborhoods regardless of shape — it perfectly recovers the two interlocked moons. The silhouette is lower because the metric itself assumes convex clusters, but the true assignment is correct.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
‖x − nearest_c‖ (the plain Euclidean distance).; Apply k-Means (k=2) to a two-moon dataset — it will find 2 clusters.Ranking
Put in order
These are the steps of k-Means recipe, scrambled. Put them back in order before the next slide shows you.
labels = argmin_k ‖x − μk‖²μk = mean({xi : label=k})Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
labels = argmin_k ‖x − μk‖²μk = mean({xi : label=k})Edge cases
Discussion prompt
k-Means recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
labels = argmin_k ‖x − μk‖²μk = mean({xi : label=k})Elimination
Eliminate the wrong options
k-Means is guaranteed to converge. The solution it converges to is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Each E-step and M-step can only decrease or maintain WCSS, so the sequence of WCSS values converges. But the algorithm is coordinate descent on a non-convex objective — the limit depends on initialization and is in general a local minimum. Our random-init run hit WCSS 5693 vs the global 567.
Check
k-Means always converges — but to what?
Check your understanding
k-Means is guaranteed to converge. The solution it converges to is:
Answer: A
Why: Each E-step and M-step can only decrease or maintain WCSS, so the sequence of WCSS values converges. But the algorithm is coordinate descent on a non-convex objective — the limit depends on initialization and is in general a local minimum. Our random-init run hit WCSS 5693 vs the global 567.
Prediction
Predict first
Point i has a(i)=1.45 (mean intra-cluster distance) and b(i)=15.56 (nearest-cluster mean distance). Its silhouette score s(i) is closest to:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 0.91
Why: s(i) = (b−a)/max(a,b) = (15.56−1.45)/max(1.45,15.56) = 14.11/15.56 ≈ 0.907 ≈ 0.91. The point is well inside its cluster and far from the next — high positive score.
Check
A point has a=1.45 and b=15.56. What is its silhouette score?
Check your understanding
Point i has a(i)=1.45 (mean intra-cluster distance) and b(i)=15.56 (nearest-cluster mean distance). Its silhouette score s(i) is closest to:
Answer: A
Why: s(i) = (b−a)/max(a,b) = (15.56−1.45)/max(1.45,15.56) = 14.11/15.56 ≈ 0.907 ≈ 0.91. The point is well inside its cluster and far from the next — high positive score.
Elimination
Eliminate the wrong options
Soft k-Means differs from hard k-Means in that it:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: In the EM view of k-Means, the E-step assigns a responsibility r_ik ∈ {0,1}. Soft k-Means relaxes this to r_ik ∈ [0,1] via a softmax — recovering the GMM with equal spherical covariances. Both still alternate E and M steps.
Check
What distinguishes soft k-Means from hard k-Means?
Check your understanding
Soft k-Means differs from hard k-Means in that it:
Answer: A
Why: In the EM view of k-Means, the E-step assigns a responsibility r_ik ∈ {0,1}. Soft k-Means relaxes this to r_ik ∈ [0,1] via a softmax — recovering the GMM with equal spherical covariances. Both still alternate E and M steps.
Section
Project
Concept
Build k-Means with both random and k++ initialization on make_blobs(centers=3). Implement the elbow method and silhouette from scratch, then demonstrate k-Means failing on moons.
| # | requirement | target value |
|---|---|---|
| 1 | k++ init, k=3, make_blobs | WCSS=566.86, silhouette=0.848 |
| 2 | elbow WCSS for k=1..7 | big drop at k=2→3 (5763→567) |
| 3 | silhouette from scratch, point 0 | s=0.907 |
| 4 | k-Means on moons (k=2) | accuracy=74.5%; silhouette=0.482 |
| 5 | DBSCAN on moons (eps=0.25) | accuracy=100%; 2 clusters |
Build rules: no sklearn.cluster.KMeans in milestones 1-3 (except to check). For DBSCAN you may use sklearn.cluster.DBSCAN.
Counterexample
Discussion prompt
Build k-Means with both random and k++ initialization on make_blobs(centers=3). Implement the elbow method and silhouette from scratch, then demonstrate k-Means failing on moons.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: no sklearn.cluster.KMeans in milestones 1-3 (except to check). For DBSCAN you may use sklearn.cluster.DBSCAN.
Worked example
Your turn: implement kmeans_pp(X, k) — k++ seeding then assign-update loop. Predict: does k++ converge faster than random init?
Hint: seed loop — d2 = min distance² to any chosen center; rng.choice(n, p=d2/d2.sum()).
import numpy as np
from sklearn.datasets import make_blobs
X, _ = make_blobs(n_samples=300, centers=3, cluster_std=1.0, random_state=42)
def kmeans_pp(X, k, seed=42):
rng = np.random.default_rng(seed)
c = [X[rng.integers(len(X))]]
for _ in range(k-1):
d2 = np.array([min(np.sum((x-ci)**2) for ci in c) for x in X])
c.append(X[rng.choice(len(X), p=d2/d2.sum())])
centers = np.array(c)
for it in range(100):
labels = np.array([np.argmin([np.sum((x-ci)**2) for ci in centers]) for x in X])
new_c = np.array([X[labels==j].mean(0) for j in range(k)])
if np.allclose(new_c, centers):
print(f'converged iter {it+1}'); break
centers = new_c
return labels, centers
labels, centers = kmeans_pp(X, 3)
wcss = sum(np.sum((X[labels==k]-centers[k])**2) for k in range(3))
print(f'WCSS={wcss:.4f}')| run | WCSS | converged at iter |
|---|---|---|
| random init (seed=0) | 5693.43 | 10 |
| k-Means++ (seed=42) | 566.86 | 2 |
| sklearn reference | 566.86 | — |
Worked example
Your turn: run k-Means++ for k=1 to 7 and compute WCSS each time. Predict where the elbow is.
Hint: call kmeans_pp(X, k) for each k; compute sum((X[labels==j] - centers[j])**2 for j in range(k)).
from sklearn.cluster import KMeans # check with sklearn
wcss_k = []
for k in range(1, 8):
km = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
wcss_k.append(km.inertia_)
print(f'k={k}: WCSS={km.inertia_:.2f}')| k | WCSS | drop from k-1 |
|---|---|---|
| 1 | 20402.34 | — |
| 2 | 5763.46 | 14638.88 |
| 3 | 566.86 | 5196.60 ← elbow |
| 4 | 496.43 | 70.43 |
| 5 | 427.13 | 69.30 |
| 6 | 375.03 | 52.10 |
| 7 | 308.20 | 66.83 |
Worked example
Your turn: for the k=3 labels from Milestone 1, compute s(i) for every i without sklearn's silhouette_score. Check your mean vs 0.848.
Hint: a[i] = mean(dist(X[i], X[same cluster])); b[i] = min over other clusters of mean(dist(X[i], X[cluster k])).
from sklearn.metrics import silhouette_score
# scratch implementation
def sil_scratch(X, labels):
k = len(set(labels))
scores = []
for i, x in enumerate(X):
c = labels[i]
a = np.mean(np.sqrt(np.sum((X[labels==c]-x)**2, axis=1)))
bs = [np.mean(np.sqrt(np.sum((X[labels==j]-x)**2, axis=1)))
for j in range(k) if j != c]
b = min(bs)
scores.append((b-a)/max(a,b))
return np.mean(scores)
print(f'scratch silhouette: {sil_scratch(X, labels):.4f}')
print(f'sklearn silhouette: {silhouette_score(X, labels):.4f}')| method | mean silhouette |
|---|---|
| scratch | 0.8480 |
| sklearn | 0.8480 |
| match | ✓ |
Trade off
Comparison matrix
From Milestone 3 — silhouette from scratch: every row here is a choice with a cost. Fill the mean silhouette column, then say which row you would actually pick and what you give up for it.
| method | mean silhouette |
|---|---|
| scratch | 0.8480 |
| sklearn | 0.8480 |
| match | ✓ |
Concept
import numpy as np
from sklearn.datasets import make_blobs, make_moons
from sklearn.cluster import DBSCAN
from sklearn.metrics import silhouette_score
X, _ = make_blobs(n_samples=300, centers=3, cluster_std=1.0, random_state=42)
def kmeans_pp(X, k, seed=42):
rng = np.random.default_rng(seed)
c = [X[rng.integers(len(X))]]
for _ in range(k-1):
d2 = np.array([min(np.sum((x-ci)**2) for ci in c) for x in X])
c.append(X[rng.choice(len(X), p=d2/d2.sum())])
centers = np.array(c)
for _ in range(100):
labels = np.array([np.argmin([np.sum((x-ci)**2) for ci in centers]) for x in X])
new_c = np.array([X[labels==j].mean(0) for j in range(k)])
if np.allclose(new_c, centers): break
centers = new_c
return labels, centers
labels, centers = kmeans_pp(X, 3)
wcss = sum(np.sum((X[labels==k]-centers[k])**2) for k in range(3))
print(f'WCSS={wcss:.4f}, sil={silhouette_score(X, labels):.4f}')
X_m, y_m = make_moons(n_samples=200, noise=0.07, random_state=42)
lb, _ = kmeans_pp(X_m, 2)
db = DBSCAN(eps=0.25, min_samples=5).fit_predict(X_m)
print(f'moons k-Means sil={silhouette_score(X_m, lb):.4f}')
print(f'moons DBSCAN sil={silhouette_score(X_m, db):.4f}, clusters={len(set(db))}')| output | value |
|---|---|
| WCSS (blobs, k=3) | 566.8596 |
| silhouette (blobs) | 0.8480 |
| k-Means sil (moons) | 0.4815 |
| DBSCAN sil (moons) | 0.3187 |
| DBSCAN clusters | 2 (correct) |
If your WCSS and silhouette match — you have the complete k-Means toolkit. The lower DBSCAN silhouette does not mean it's worse; it means the silhouette metric itself favors convex clusters.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| WCSS (blobs, k=3) | 566.8596 |
| silhouette (blobs) | 0.8480 |
| k-Means sil (moons) | 0.4815 |
| DBSCAN sil (moons) | 0.3187 |
| DBSCAN clusters | 2 (correct) |
Concept
Out loud, slides closed: (1) walk through one full E + M step without looking at code, (2) explain why k-Means++ picks seeds that are far apart, (3) state the silhouette formula and explain what a score near 1 vs near -1 means.
Stretch (homework from lesson plan): plot the k-Means convergence curve (WCSS vs iteration) for random vs k++ init on the same axes; implement the elbow method as a function; prove algebraically that k-Means M-step minimizes WCSS given fixed labels (i.e., the mean minimizes sum of squared deviations). Next up: Gaussian Mixture Models — the soft-assignment generalization of k-Means.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — k-Means: assign & update · k-Means++ initialization · Elbow & silhouette · Your turn: build k-Means from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
(b−a)/max(a,b)| idea | the one thing to remember |
|---|---|
| k-Means objective | minimize WCSS — local minimum, not global |
| k-Means++ init | P(next seed) ∝ d² to nearest chosen center |
| elbow method | pick k at the sharpest WCSS drop |
| silhouette score | (b−a)/max(a,b); range [−1, 1], higher = better |
| EM view | E=hard assign, M=centroid mean; soft→GMM |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.