USAAIO Lesson 54, wrapping up Week 19. It covers DBSCAN, with its core, border, and noise points and its epsilon and min_samples parameters, then agglomerative hierarchical clustering with dendrograms and the single, complete, average, and Ward linkage methods. It covers validating clusters without labels, by silhouette score, and with labels, by ARI and NMI, and ends with a practical guide to matching the geometry of a dataset to the right algorithm. It was verified with sklearn and scipy in June 2026. The lesson runs to 28 slides.
Subject: Machine Learning · 58 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 54 · Week 19
Two algorithms that don't need k upfront — plus the validation metrics that tell you whether any clustering is actually good.
Objectives
DBSCAN and explain why it handles arbitrary-shape clustersscipy.linkage, cut it at different heightsWarm-up
Discussion prompt
Before we open Lesson 54: DBSCAN & Hierarchical Clustering: without looking back, what was the main idea of Gaussian Mixture Models & the EM Algorithm, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
GMM density model (K weighted Gaussians), E-step soft assignments and M-step parameter updates derived from expected complete-data log-likelihood, EM convergence guarantees, covariance types and BIC model selection, GMM vs k-Means on non-spherical data, and GMM-based anomaly detection. Build GMM from scratch with NumPy/SciPy and validate against sklearn.
Section
Part 1 of 3
Concept
DBSCAN takes epsilon (neighborhood radius) and min_samples (density threshold) — no k needed.
| parameter | meaning | effect if too large |
|---|---|---|
| epsilon | radius of a neighborhood ball | clusters merge together |
| min_samples | min pts to be a core point | more points become noise |
Typical starting point: epsilon from a k-NN distance plot (elbow), min_samples = 2 * n_features.
Comparison
Comparison matrix
From The two hyperparameters: refill the meaning column from what you know. The rest of the table is as it appeared.
| parameter | meaning | effect if too large |
|---|---|---|
| epsilon | radius of a neighborhood ball | clusters merge together |
| min_samples | min pts to be a core point | more points become noise |
Concept
Every point gets exactly one classification based on its epsilon-neighborhood.
| type | condition | role |
|---|---|---|
| core | >= min_samples within epsilon | seeds a cluster |
| border | within epsilon of a core, but < min_samples own nbrs | extends cluster edge |
| noise | not within epsilon of any core | label = -1 |
A cluster = a core point + every point density-reachable from it (transitively via other cores).
Trade off
Comparison matrix
From Core, border, and noise points: every row here is a choice with a cost. Fill the condition column, then say which row you would actually pick and what you give up for it.
| type | condition | role |
|---|---|---|
| core | >= min_samples within epsilon | seeds a cluster |
| border | within epsilon of a core, but < min_samples own nbrs | extends cluster edge |
| noise | not within epsilon of any core | label = -1 |
Estimation
Predict first
Apply DBSCAN (eps=0.2, min_samples=5) to the two-moon dataset — a geometry k-Means cannot separate.
Commit before you compute: what does DBSCAN on make_moons come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: 2 clusters, 0 noise, 196 core pts, 4 border pts; ARI = 1.0000
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. DBSCAN follows the curved density: the two half-moons are separated exactly.
Worked example
Apply DBSCAN (eps=0.2, min_samples=5) to the two-moon dataset — a geometry k-Means cannot separate.
from sklearn.datasets import make_moons
from sklearn.cluster import DBSCAN
from sklearn.metrics import adjusted_rand_score
X, y = make_moons(n_samples=200, noise=0.05, random_state=42)
db = DBSCAN(eps=0.2, min_samples=5).fit(X)
lbls = db.labels_
n_clusters = len(set(lbls)) - (1 if -1 in lbls else 0)
n_noise = (lbls == -1).sum()
n_core = len(db.core_sample_indices_)
n_border = len(lbls) - n_core - n_noise
print(n_clusters, n_noise, n_core, n_border)
print(round(adjusted_rand_score(y, lbls), 4))2 clusters, 0 noise, 196 core pts, 4 border pts; ARI = 1.0000
Why: DBSCAN follows the curved density: the two half-moons are separated exactly. ARI = 1 means perfect agreement with ground truth.
| metric | DBSCAN (eps=0.2) | KMeans (k=2) |
|---|---|---|
| clusters found | 2 | 2 |
| ARI vs true | 1.000 | 0.217 |
| captures crescent shape | yes | no |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
2 clusters, 0 noise, 196 core pts, 4 border pts; ARI = 1.0000
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Apply DBSCAN (eps=0.2, min_samples=5) to the two-moon dataset — a geometry k-Means cannot separate.
Missing information
Discussion prompt
Nine points arranged so we get all three types at once — trace through each.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
pts 0-3 each have >= 3 neighbors within 0.5 => core; pts 4-6 have exactly 2 neighbors among themselves so they are border points in cluster B; pts 7 and 8 have 0 neighbors => noise.
Worked example
Nine points arranged so we get all three types at once — trace through each.
import numpy as np
from sklearn.cluster import DBSCAN
X = np.array([
[0.0, 0.0], [0.3, 0.0], [0.0, 0.3], [0.15, 0.15], # dense cluster A
[3.0, 3.0], [3.3, 3.0], [3.0, 3.3], # cluster B (3 pts)
[1.65, 0.0], # too far from both
[8.0, 8.0], # isolated
])
db = DBSCAN(eps=0.5, min_samples=3).fit(X)
print(db.labels_) # [0 0 0 0 1 1 1 -1 -1]pts 0-3 label=0 (cluster A), pts 4-6 label=1 (cluster B), pts 7-8 label=-1 (noise)
Why: pts 0-3 each have >= 3 neighbors within 0.5 => core; pts 4-6 have exactly 2 neighbors among themselves so they are border points in cluster B; pts 7 and 8 have 0 neighbors => noise.
| point | neighbors within 0.5 | type | label |
|---|---|---|---|
| pt0 (0,0) | 3 (pts 1,2,3) | core | 0 |
| pt4 (3,3) | 2 (pts 5,6) | border | 1 |
| pt7 (1.65,0) | 0 | noise | -1 |
| pt8 (8,8) | 0 | noise | -1 |
Discrimination
Sort into buckets
Sort these by neighbors within 0.5, from memory, without looking back at Point type classification (eps=0.5, min_samples=3). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use a generous epsilon so no points are labeled noise — every point deserves a cluster.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A large epsilon connects core points across both moons.
Choose epsilon from a k-NN distance plot: sort distances to the k-th nearest neighbor and find the elbow.
Why: A large epsilon connects core points across both moons. DBSCAN returns 1 cluster instead of 2 — the neighborhood balls bridge the gap between the arcs.
Trap
Use a generous epsilon so no points are labeled noise — every point deserves a cluster.
Set eps=2.0 on a two-moon dataset with well-separated arcs
Why: A large epsilon connects core points across both moons. DBSCAN returns 1 cluster instead of 2 — the neighborhood balls bridge the gap between the arcs.
Choose epsilon from a k-NN distance plot: sort distances to the k-th nearest neighbor and find the elbow.
Pick epsilon at the elbow of the sorted k-NN distance curve
Why: Below the elbow, points are tightly connected within clusters; above it, epsilon bridges inter-cluster gaps. The elbow keeps clusters separate while accepting a few border points.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Choose epsilon from a k-NN distance plot: sort distances to the k-th nearest neighbor and find the elbow.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
A large epsilon connects core points across both moons. DBSCAN returns 1 cluster instead of 2 — the neighborhood balls bridge the gap between the arcs.
Section
Part 2 of 3
Concept
Hierarchical clustering builds a tree of merges (or splits) — no k required at fit time; choose k by cutting the tree.
| approach | direction | typical use |
|---|---|---|
| agglomerative | bottom-up: each point is its own cluster; merge greedily | more common; scipy/sklearn |
| divisive | top-down: start with all points; split recursively | DIANA; rarely in practice |
The merge history is stored in a dendrogram — a tree you can cut at any height to get k clusters.
Counterexample
Discussion prompt
Hierarchical clustering builds a tree of merges (or splits) — no k required at fit time; choose k by cutting the tree.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The merge history is stored in a dendrogram — a tree you can cut at any height to get k clusters.
Concept
The linkage criterion defines the distance between two clusters (not two points).
| method | dist(A, B) = | weakness |
|---|---|---|
| single | min dist between any pt in A and B | chaining: elongated chains form |
| complete | max dist between any pt in A and B | sensitive to outliers |
| average | mean of all pairwise A-B distances | compromise; slower |
| Ward | increase in total within-cluster variance | best for compact, equal-size clusters |
Ward is the default for most ML problems — it minimizes within-cluster variance at every merge, analogous to the k-Means objective.
Analogy
Discussion prompt
Explain Linkage methods by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The linkage criterion defines the distance between two clusters (not two points).
Estimation
Predict first
Fit Ward linkage on 150 three-blob points, then cut at k=2, 3, 4 and compare silhouette + ARI.
Commit before you compute: what does Dendrogram on make_blobs — cut at different heights come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: k=2: sil=0.7042, ARI=0.5681 | k=3: sil=0.8448, ARI=1.0000 | k=4: sil=0.6584, ARI=0.8695
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Silhouette peaks at k=3 and ARI hits 1.0 — both metrics agree the true number of clusters is 3.
Worked example
Fit Ward linkage on 150 three-blob points, then cut at k=2, 3, 4 and compare silhouette + ARI.
from sklearn.datasets import make_blobs
from sklearn.metrics import silhouette_score, adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster
X, y = make_blobs(n_samples=150, centers=3, cluster_std=1.0, random_state=42)
Z = linkage(X, method='ward')
for k in [2, 3, 4]:
lbl = fcluster(Z, k, criterion='maxclust')
sil = silhouette_score(X, lbl)
ari = adjusted_rand_score(y, lbl)
print(f'k={k}: sil={sil:.4f} ARI={ari:.4f}')k=2: sil=0.7042, ARI=0.5681 | k=3: sil=0.8448, ARI=1.0000 | k=4: sil=0.6584, ARI=0.8695
Why: Silhouette peaks at k=3 and ARI hits 1.0 — both metrics agree the true number of clusters is 3. Cutting higher (k=4) creates artificial splits.
| k (cut) | silhouette | ARI vs true | verdict |
|---|---|---|---|
| 2 | 0.7042 | 0.5681 | under-clustered |
| 3 | 0.8448 | 1.0000 | correct cut |
| 4 | 0.6584 | 0.8695 | over-split |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
k=2: sil=0.7042, ARI=0.5681 | k=3: sil=0.8448, ARI=1.0000 | k=4: sil=0.6584, ARI=0.8695
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Fit Ward linkage on 150 three-blob points, then cut at k=2, 3, 4 and compare silhouette + ARI.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Single linkage is simple — just use it for everything. It finds the minimum distance, so it can't go wrong.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Single linkage follows the chain: as long as there is one pair of close points connecting two groups, they merge.
Match linkage to data geometry: Ward for compact equal-sized groups; single linkage only for well-separated clusters with no bridges.
Why: Single linkage follows the chain: as long as there is one pair of close points connecting two groups, they merge. A single outlier lying between two real clusters causes them to collapse into one.
Trap
Single linkage is simple — just use it for everything. It finds the minimum distance, so it can't go wrong.
Apply single linkage to a dataset with noise points forming a bridge between two clusters
Why: Single linkage follows the chain: as long as there is one pair of close points connecting two groups, they merge. A single outlier lying between two real clusters causes them to collapse into one.
Match linkage to data geometry: Ward for compact equal-sized groups; single linkage only for well-separated clusters with no bridges.
Use Ward (or complete) when data has noise or unequal densities
Why: Ward minimizes within-cluster variance at every merge — it is robust to the chaining artifact because it prefers tight, low-variance merges over distant minimum-distance connections.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Section
Part 3 of 3
Concept
Silhouette measures how similar a point is to its own cluster vs the nearest other cluster. Range: -1 to +1; higher is better.
\[ s(i) = \frac{b(i) - a(i)}{\max(a(i),\, b(i))} \]
a(i) = mean intra-cluster distance; b(i) = mean distance to nearest other cluster. On blobs: k=3 gives silhouette 0.8448 vs k=4 giving 0.6584 — the true k wins.
Explain it
Discussion prompt
Explain Silhouette score — no labels needed to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Silhouette measures how similar a point is to its own cluster vs the nearest other cluster. Range: -1 to +1; higher is better.
Concept
ARI (Adjusted Rand Index): measures agreement between predicted and true cluster assignments, corrected for chance. Range: -1 to 1; random ≈ 0, perfect = 1.
| prediction quality | ARI | NMI |
|---|---|---|
| perfect match | 1.0000 | 1.0000 |
| random assignment | -0.1111 | 0.0817 |
| partial agreement | 0.46–0.87 | varies |
NMI (Normalized Mutual Information) captures information overlap, range 0-1. Use ARI when cluster sizes matter; NMI when you want a symmetric information measure.
Comparison
Comparison matrix
From ARI and NMI — when you have ground-truth labels: refill the NMI column from what you know. The rest of the table is as it appeared.
| prediction quality | ARI | NMI |
|---|---|---|
| perfect match | 1.0000 | 1.0000 |
| random assignment | -0.1111 | 0.0817 |
| partial agreement | 0.46–0.87 | varies |
Estimation
Predict first
Run KMeans, GMM, DBSCAN, and Ward on blobs / moons / circles — one table shows the decisive pattern.
Commit before you compute: what does Validation: 3 datasets × 4 methods come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: blobs: all methods ARI~1.0 | moons: KMeans 0.217, DBSCAN 1.000 | circles: KMeans -0.005, DBSCAN 1.000
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On convex well-separated blobs any method works.
Worked example
Run KMeans, GMM, DBSCAN, and Ward on blobs / moons / circles — one table shows the decisive pattern.
from sklearn.datasets import make_blobs, make_moons, make_circles
from sklearn.cluster import DBSCAN, KMeans, AgglomerativeClustering
from sklearn.mixture import GaussianMixture
from sklearn.metrics import adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster
import warnings; warnings.filterwarnings('ignore')
for dname, (X, y), eps in [
('blobs', make_blobs(200, centers=3, random_state=42), 1.5),
('moons', make_moons(200, noise=0.05, random_state=42), 0.2),
('circles', make_circles(200, noise=0.05, factor=0.5, random_state=42), 0.2),
]:
k = len(set(y))
km = KMeans(k, n_init=10, random_state=42).fit_predict(X)
db = DBSCAN(eps=eps, min_samples=5).fit_predict(X)
Z = linkage(X, 'ward'); wd = fcluster(Z, k, criterion='maxclust')
for mn, lb in [('KM', km), ('DB', db), ('Ward', wd)]:
mask = lb != -1
ari = adjusted_rand_score(y[mask], lb[mask])
print(f'{dname:8s} {mn:6s} ARI={ari:.3f}')blobs: all methods ARI~1.0 | moons: KMeans 0.217, DBSCAN 1.000 | circles: KMeans -0.005, DBSCAN 1.000
Why: On convex well-separated blobs any method works. On non-convex shapes (moons, circles) only density-based DBSCAN recovers the true clusters; centroid and linkage methods fail completely.
| dataset | KMeans ARI | DBSCAN ARI | Ward ARI |
|---|---|---|---|
| blobs | 1.000 | 1.000 | 1.000 |
| moons | 0.217 | 1.000 | 0.446 |
| circles | -0.005 | 1.000 | -0.005 |
Invariant
Step through it
Step through Validation: 3 datasets × 4 methods one row at a time. One of these columns never changes — find it, and say why it cannot.
Ranking
Put in order
These are the steps of Algorithm selection guide, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Edge cases
Discussion prompt
Algorithm selection guide works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
With eps=0.5 and min_samples=4, a point has exactly 3 other points within distance 0.5 and is itself within distance 0.5 of a core point. What type is it?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A core point needs >= min_samples (4) neighbors within eps. This point has only 3 — one short — so it is NOT a core point. But it is within eps of a core point, making it a border point (label = cluster of that core, not -1).
Check
Reason through the DBSCAN definitions.
Check your understanding
With eps=0.5 and min_samples=4, a point has exactly 3 other points within distance 0.5 and is itself within distance 0.5 of a core point. What type is it?
Answer: A
Why: A core point needs >= min_samples (4) neighbors within eps. This point has only 3 — one short — so it is NOT a core point. But it is within eps of a core point, making it a border point (label = cluster of that core, not -1).
Prediction
Predict first
Running Ward linkage on make_blobs(centers=3) and cutting at k=2, 3, 4, 5 yields silhouette scores 0.704, 0.845, 0.658, 0.500. Which k do you report as optimal?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: k = 3, highest silhouette
Why: Silhouette peaks at k=3 (0.845). This data was generated with 3 centers, so ARI also reaches 1.0 at k=3 — the two metrics agree. Always report the k that maximizes silhouette; k=2 under-clusters and k=4/5 over-splits.
Check
Use the silhouette formula.
Check your understanding
Running Ward linkage on make_blobs(centers=3) and cutting at k=2, 3, 4, 5 yields silhouette scores 0.704, 0.845, 0.658, 0.500. Which k do you report as optimal?
Answer: A
Why: Silhouette peaks at k=3 (0.845). This data was generated with 3 centers, so ARI also reaches 1.0 at k=3 — the two metrics agree. Always report the k that maximizes silhouette; k=2 under-clusters and k=4/5 over-splits.
Elimination
Eliminate the wrong options
Which linkage method is most susceptible to the 'chaining' problem, where a single outlier bridges two true clusters?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Single linkage uses the MINIMUM distance between clusters. A single outlier lying between two groups is close enough to both to trigger a merge — the chain collapses real clusters into one. Complete and Ward use maximum / variance-increase criteria that resist this.
Check
Recall the linkage trade-offs.
Check your understanding
Which linkage method is most susceptible to the 'chaining' problem, where a single outlier bridges two true clusters?
Answer: A
Why: Single linkage uses the MINIMUM distance between clusters. A single outlier lying between two groups is close enough to both to trigger a merge — the chain collapses real clusters into one. Complete and Ward use maximum / variance-increase criteria that resist this.
Section
Project
Concept
Three milestones: DBSCAN on moons/circles, a dendrogram cut comparison, and the full 3×4 validation table.
| # | task | key tool |
|---|---|---|
| 1 | DBSCAN on make_moons and make_circles | DBSCAN, adjusted_rand_score |
| 2 | Ward dendrogram, cut at k=2/3/4 | linkage, fcluster, silhouette_score |
| 3 | Compare all methods on blobs/moons/circles | KMeans, GMM, DBSCAN, AgglomerativeClustering |
Rules: print every ARI and silhouette before looking at the table on the next slide. Predict which method will fail on moons before running.
Analogy
Discussion prompt
Explain Project: clustering lab by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Three milestones: DBSCAN on moons/circles, a dendrogram cut comparison, and the full 3×4 validation table.
Worked example
Your turn: run DBSCAN (eps=0.2, min_samples=5) on make_moons. Predict: how many clusters, how many noise points?
Hint: import make_moons and DBSCAN; print len(set(labels)) - (1 if -1 in labels else 0) and (labels==-1).sum().
from sklearn.datasets import make_moons
from sklearn.cluster import DBSCAN
from sklearn.metrics import adjusted_rand_score
X, y = make_moons(n_samples=200, noise=0.05, random_state=42)
lbls = DBSCAN(eps=0.2, min_samples=5).fit_predict(X)
n_clust = len(set(lbls)) - (1 if -1 in lbls else 0)
print(n_clust, (lbls==-1).sum(), round(adjusted_rand_score(y, lbls), 4))| output | value |
|---|---|
| n_clusters | 2 |
| n_noise | 0 |
| ARI | 1.0000 |
Pattern
Step through it
Step through Milestone 1 — DBSCAN on moons one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: fit Ward linkage on make_blobs(centers=3), cut at k=2,3,4. Predict which k has the highest silhouette.
Hint: linkage(X, method='ward') then fcluster(Z, k, criterion='maxclust') and silhouette_score(X, lbl).
from sklearn.datasets import make_blobs
from sklearn.metrics import silhouette_score, adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster
X, y = make_blobs(n_samples=150, centers=3, cluster_std=1.0, random_state=42)
Z = linkage(X, method='ward')
for k in [2, 3, 4]:
lbl = fcluster(Z, k, criterion='maxclust')
print(k, round(silhouette_score(X,lbl),4), round(adjusted_rand_score(y,lbl),4))| k | silhouette | ARI |
|---|---|---|
| 2 | 0.7042 | 0.5681 |
| 3 | 0.8448 | 1.0000 |
| 4 | 0.6584 | 0.8695 |
Trade off
Comparison matrix
From Milestone 2 — Ward dendrogram cut: every row here is a choice with a cost. Fill the silhouette column, then say which row you would actually pick and what you give up for it.
| k | silhouette | ARI |
|---|---|---|
| 2 | 0.7042 | 0.5681 |
| 3 | 0.8448 | 1.0000 |
| 4 | 0.6584 | 0.8695 |
Concept
from sklearn.datasets import make_blobs, make_moons, make_circles
from sklearn.cluster import DBSCAN, KMeans
from sklearn.mixture import GaussianMixture
from sklearn.metrics import adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster
import warnings; warnings.filterwarnings('ignore')
for dname, (X, y), eps in [
('blobs', make_blobs(200, centers=3, random_state=42), 1.5),
('moons', make_moons(200, noise=0.05, random_state=42), 0.2),
('circles', make_circles(200, noise=0.05, factor=0.5, random_state=42), 0.2),
]:
k = len(set(y))
km = KMeans(k, n_init=10, random_state=42).fit_predict(X)
db = DBSCAN(eps=eps, min_samples=5).fit_predict(X)
Z = linkage(X, 'ward'); wd = fcluster(Z, k, criterion='maxclust')
for mn, lb in [('KMeans',km),('DBSCAN',db),('Ward',wd)]:
mask = lb != -1
ari = adjusted_rand_score(y[mask], lb[mask])
print(f'{dname:8} {mn:8} ARI={ari:.3f}')| dataset | KMeans ARI | DBSCAN ARI | Ward ARI |
|---|---|---|---|
| blobs | 1.000 | 1.000 | 1.000 |
| moons | 0.217 | 1.000 | 0.446 |
| circles | -0.005 | 1.000 | -0.005 |
If you see DBSCAN at ARI=1.000 on moons and circles while KMeans sits near 0 — you have empirically proven why density-based methods exist.
Comparison
Comparison matrix
From The full comparison program: refill the KMeans ARI column from what you know. The rest of the table is as it appeared.
| dataset | KMeans ARI | DBSCAN ARI | Ward ARI |
|---|---|---|---|
| blobs | 1.000 | 1.000 | 1.000 |
| moons | 0.217 | 1.000 | 0.446 |
| circles | -0.005 | 1.000 | -0.005 |
Concept
Slides closed: explain (1) the three DBSCAN point types and what epsilon controls, (2) why Ward linkage is preferred over single for most problems, (3) when to use silhouette vs ARI.
Stretch (homework): add make_circles to Milestone 1; try different eps values to break DBSCAN on moons; cluster the load_digits PCA embedding with Ward and measure NMI against digit labels.
Counterexample
Discussion prompt
Slides closed: explain (1) the three DBSCAN point types and what epsilon controls, (2) why Ward linkage is preferred over single for most problems, (3) when to use silhouette vs ARI.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Stretch (homework): add make_circles to Milestone 1; try different eps values to break DBSCAN on moons; cluster the load_digits PCA embedding with Ward and measure NMI against digit labels.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — DBSCAN: density-based clustering · Hierarchical clustering & dendrograms · Cluster validation metrics · Your turn: implement it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| idea | the one thing to remember |
|---|---|
| DBSCAN eps too large | clusters merge — use k-NN elbow to set eps |
| single linkage | chaining risk — one bridge collapses two clusters |
| Ward linkage | minimizes variance increase = best default |
| silhouette | no labels needed; peak = right k |
| ARI | corrected for chance; random baseline = 0 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.