Lesson 54: DBSCAN & Hierarchical Clustering

USAAIO Lesson 54, wrapping up Week 19. It covers DBSCAN, with its core, border, and noise points and its epsilon and min_samples parameters, then agglomerative hierarchical clustering with dendrograms and the single, complete, average, and Ward linkage methods. It covers validating clusters without labels, by silhouette score, and with labels, by ARI and NMI, and ends with a practical guide to matching the geometry of a dataset to the right algorithm. It was verified with sklearn and scipy in June 2026. The lesson runs to 28 slides.

Subject: Machine Learning · 58 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. DBSCAN & Hierarchical Clustering

Title

USAAIO · Lesson 54 · Week 19

Two algorithms that don't need k upfront — plus the validation metrics that tell you whether any clustering is actually good.

2. By the end of this lesson you can

Objectives

  1. Classify every point as core, border, or noise given epsilon and min_samples
  2. Run DBSCAN and explain why it handles arbitrary-shape clusters
  3. Build a dendrogram with scipy.linkage, cut it at different heights
  4. Compare single / complete / average / Ward linkage trade-offs
  5. Evaluate any clustering with silhouette (unsupervised) and ARI / NMI (supervised)

3. What survived from Gaussian Mixture Models & the EM Algorithm?

Warm-up

Discussion prompt

Before we open Lesson 54: DBSCAN & Hierarchical Clustering: without looking back, what was the main idea of Gaussian Mixture Models & the EM Algorithm, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

GMM density model (K weighted Gaussians), E-step soft assignments and M-step parameter updates derived from expected complete-data log-likelihood, EM convergence guarantees, covariance types and BIC model selection, GMM vs k-Means on non-spherical data, and GMM-based anomaly detection. Build GMM from scratch with NumPy/SciPy and validate against sklearn.

4. DBSCAN: density-based clustering

Section

Part 1 of 3

5. The two hyperparameters

Concept

DBSCAN takes epsilon (neighborhood radius) and min_samples (density threshold) — no k needed.

parametermeaningeffect if too large
epsilonradius of a neighborhood ballclusters merge together
min_samplesmin pts to be a core pointmore points become noise

Typical starting point: epsilon from a k-NN distance plot (elbow), min_samples = 2 * n_features.

6. Fill in: meaning for The two hyperparameters

Comparison

Comparison matrix

From The two hyperparameters: refill the meaning column from what you know. The rest of the table is as it appeared.

parametermeaningeffect if too large
epsilonradius of a neighborhood ballclusters merge together
min_samplesmin pts to be a core pointmore points become noise

7. Core, border, and noise points

Concept

Every point gets exactly one classification based on its epsilon-neighborhood.

typeconditionrole
core>= min_samples within epsilonseeds a cluster
borderwithin epsilon of a core, but < min_samples own nbrsextends cluster edge
noisenot within epsilon of any corelabel = -1

A cluster = a core point + every point density-reachable from it (transitively via other cores).

8. What each one costs: Core, border, and noise points

Trade off

Comparison matrix

From Core, border, and noise points: every row here is a choice with a cost. Fill the condition column, then say which row you would actually pick and what you give up for it.

typeconditionrole
core>= min_samples within epsilonseeds a cluster
borderwithin epsilon of a core, but < min_samples own nbrsextends cluster edge
noisenot within epsilon of any corelabel = -1

9. Guess the shape of the answer: DBSCAN on make_moons

Estimation

Predict first

Apply DBSCAN (eps=0.2, min_samples=5) to the two-moon dataset — a geometry k-Means cannot separate.

Commit before you compute: what does DBSCAN on make_moons come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: 2 clusters, 0 noise, 196 core pts, 4 border pts; ARI = 1.0000

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. DBSCAN follows the curved density: the two half-moons are separated exactly.

10. DBSCAN on make_moons

Worked example

Apply DBSCAN (eps=0.2, min_samples=5) to the two-moon dataset — a geometry k-Means cannot separate.

from sklearn.datasets import make_moons
from sklearn.cluster import DBSCAN
from sklearn.metrics import adjusted_rand_score

X, y = make_moons(n_samples=200, noise=0.05, random_state=42)
db  = DBSCAN(eps=0.2, min_samples=5).fit(X)
lbls = db.labels_

n_clusters = len(set(lbls)) - (1 if -1 in lbls else 0)
n_noise    = (lbls == -1).sum()
n_core     = len(db.core_sample_indices_)
n_border   = len(lbls) - n_core - n_noise
print(n_clusters, n_noise, n_core, n_border)
print(round(adjusted_rand_score(y, lbls), 4))

2 clusters, 0 noise, 196 core pts, 4 border pts; ARI = 1.0000

Why: DBSCAN follows the curved density: the two half-moons are separated exactly. ARI = 1 means perfect agreement with ground truth.

metricDBSCAN (eps=0.2)KMeans (k=2)
clusters found22
ARI vs true1.0000.217
captures crescent shapeyesno

11. Work backwards from the answer: DBSCAN on make_moons

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

2 clusters, 0 noise, 196 core pts, 4 border pts; ARI = 1.0000

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Apply DBSCAN (eps=0.2, min_samples=5) to the two-moon dataset — a geometry k-Means cannot separate.

12. What has to be given first: Point type classification (eps=0.5…

Missing information

Discussion prompt

Nine points arranged so we get all three types at once — trace through each.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

pts 0-3 each have >= 3 neighbors within 0.5 => core; pts 4-6 have exactly 2 neighbors among themselves so they are border points in cluster B; pts 7 and 8 have 0 neighbors => noise.

13. Point type classification (eps=0.5, min_samples=3)

Worked example

Nine points arranged so we get all three types at once — trace through each.

import numpy as np
from sklearn.cluster import DBSCAN

X = np.array([
    [0.0, 0.0], [0.3, 0.0], [0.0, 0.3], [0.15, 0.15],  # dense cluster A
    [3.0, 3.0], [3.3, 3.0], [3.0, 3.3],                 # cluster B (3 pts)
    [1.65, 0.0],                                          # too far from both
    [8.0, 8.0],                                           # isolated
])
db = DBSCAN(eps=0.5, min_samples=3).fit(X)
print(db.labels_)   # [0 0 0 0 1 1 1 -1 -1]

pts 0-3 label=0 (cluster A), pts 4-6 label=1 (cluster B), pts 7-8 label=-1 (noise)

Why: pts 0-3 each have >= 3 neighbors within 0.5 => core; pts 4-6 have exactly 2 neighbors among themselves so they are border points in cluster B; pts 7 and 8 have 0 neighbors => noise.

pointneighbors within 0.5typelabel
pt0 (0,0)3 (pts 1,2,3)core0
pt4 (3,3)2 (pts 5,6)border1
pt7 (1.65,0)0noise-1
pt8 (8,8)0noise-1

14. Which is which, by neighbors within 0.5

Discrimination

Sort into buckets

Sort these by neighbors within 0.5, from memory, without looking back at Point type classification (eps=0.5, min_samples=3). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

3 (pts 1,2,3)
pt0 (0,0)
2 (pts 5,6)
pt4 (3,3)
0
pt7 (1.65,0); pt8 (8,8)
g1
neighbors within 0.5 is "3 (pts 1,2,3)" for pt0 (0,0) — that is what the table on "Point type classification (eps=0.5…" records, and it is the single property separating this group from the rest.
g2
neighbors within 0.5 is "2 (pts 5,6)" for pt4 (3,3) — that is what the table on "Point type classification (eps=0.5…" records, and it is the single property separating this group from the rest.
g3
neighbors within 0.5 is "0" for pt7 (1.65,0), pt8 (8,8) — that is what the table on "Point type classification (eps=0.5…" records, and it is the single property separating this group from the rest.

15. Something is wrong here: epsilon too large merges clusters

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use a generous epsilon so no points are labeled noise — every point deserves a cluster.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A large epsilon connects core points across both moons.

Choose epsilon from a k-NN distance plot: sort distances to the k-th nearest neighbor and find the elbow.

Why: A large epsilon connects core points across both moons. DBSCAN returns 1 cluster instead of 2 — the neighborhood balls bridge the gap between the arcs.

16. Trap: epsilon too large merges clusters

Trap

The trap

Use a generous epsilon so no points are labeled noise — every point deserves a cluster.

Set eps=2.0 on a two-moon dataset with well-separated arcs

Why: A large epsilon connects core points across both moons. DBSCAN returns 1 cluster instead of 2 — the neighborhood balls bridge the gap between the arcs.

The fix

Choose epsilon from a k-NN distance plot: sort distances to the k-th nearest neighbor and find the elbow.

Pick epsilon at the elbow of the sorted k-NN distance curve

Why: Below the elbow, points are tightly connected within clusters; above it, epsilon bridges inter-cluster gaps. The elbow keeps clusters separate while accepting a few border points.

17. Break it on purpose: epsilon too large merges clusters

Break the constraint

Discussion prompt

The rule this trap just fixed:

Choose epsilon from a k-NN distance plot: sort distances to the k-th nearest neighbor and find the elbow.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

A large epsilon connects core points across both moons. DBSCAN returns 1 cluster instead of 2 — the neighborhood balls bridge the gap between the arcs.

18. Hierarchical clustering & dendrograms

Section

Part 2 of 3

19. Agglomerative vs divisive

Concept

Hierarchical clustering builds a tree of merges (or splits) — no k required at fit time; choose k by cutting the tree.

approachdirectiontypical use
agglomerativebottom-up: each point is its own cluster; merge greedilymore common; scipy/sklearn
divisivetop-down: start with all points; split recursivelyDIANA; rarely in practice

The merge history is stored in a dendrogram — a tree you can cut at any height to get k clusters.

20. Break it if you can: Agglomerative vs divisive

Counterexample

Discussion prompt

Hierarchical clustering builds a tree of merges (or splits) — no k required at fit time; choose k by cutting the tree.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The merge history is stored in a dendrogram — a tree you can cut at any height to get k clusters.

21. Linkage methods

Concept

The linkage criterion defines the distance between two clusters (not two points).

methoddist(A, B) =weakness
singlemin dist between any pt in A and Bchaining: elongated chains form
completemax dist between any pt in A and Bsensitive to outliers
averagemean of all pairwise A-B distancescompromise; slower
Wardincrease in total within-cluster variancebest for compact, equal-size clusters

Ward is the default for most ML problems — it minimizes within-cluster variance at every merge, analogous to the k-Means objective.

22. By analogy: Linkage methods

Analogy

Discussion prompt

Explain Linkage methods by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The linkage criterion defines the distance between two clusters (not two points).

23. Guess the shape of the answer: Dendrogram on make_blobs — cut at different…

Estimation

Predict first

Fit Ward linkage on 150 three-blob points, then cut at k=2, 3, 4 and compare silhouette + ARI.

Commit before you compute: what does Dendrogram on make_blobs — cut at different heights come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: k=2: sil=0.7042, ARI=0.5681 | k=3: sil=0.8448, ARI=1.0000 | k=4: sil=0.6584, ARI=0.8695

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Silhouette peaks at k=3 and ARI hits 1.0 — both metrics agree the true number of clusters is 3.

24. Dendrogram on make_blobs — cut at different heights

Worked example

Fit Ward linkage on 150 three-blob points, then cut at k=2, 3, 4 and compare silhouette + ARI.

from sklearn.datasets import make_blobs
from sklearn.metrics import silhouette_score, adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster

X, y = make_blobs(n_samples=150, centers=3, cluster_std=1.0, random_state=42)
Z = linkage(X, method='ward')

for k in [2, 3, 4]:
    lbl = fcluster(Z, k, criterion='maxclust')
    sil = silhouette_score(X, lbl)
    ari = adjusted_rand_score(y, lbl)
    print(f'k={k}: sil={sil:.4f}  ARI={ari:.4f}')

k=2: sil=0.7042, ARI=0.5681 | k=3: sil=0.8448, ARI=1.0000 | k=4: sil=0.6584, ARI=0.8695

Why: Silhouette peaks at k=3 and ARI hits 1.0 — both metrics agree the true number of clusters is 3. Cutting higher (k=4) creates artificial splits.

k (cut)silhouetteARI vs trueverdict
20.70420.5681under-clustered
30.84481.0000correct cut
40.65840.8695over-split

25. Work backwards from the answer: Dendrogram on make_blobs — cut at different…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

k=2: sil=0.7042, ARI=0.5681 | k=3: sil=0.8448, ARI=1.0000 | k=4: sil=0.6584, ARI=0.8695

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Fit Ward linkage on 150 three-blob points, then cut at k=2, 3, 4 and compare silhouette + ARI.

26. Something is wrong here: single linkage chaining

Anomaly

Predict first

A student writes this, and it looks reasonable:

Single linkage is simple — just use it for everything. It finds the minimum distance, so it can't go wrong.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Single linkage follows the chain: as long as there is one pair of close points connecting two groups, they merge.

Match linkage to data geometry: Ward for compact equal-sized groups; single linkage only for well-separated clusters with no bridges.

Why: Single linkage follows the chain: as long as there is one pair of close points connecting two groups, they merge. A single outlier lying between two real clusters causes them to collapse into one.

27. Trap: single linkage chaining

Trap

The trap

Single linkage is simple — just use it for everything. It finds the minimum distance, so it can't go wrong.

Apply single linkage to a dataset with noise points forming a bridge between two clusters

Why: Single linkage follows the chain: as long as there is one pair of close points connecting two groups, they merge. A single outlier lying between two real clusters causes them to collapse into one.

The fix

Match linkage to data geometry: Ward for compact equal-sized groups; single linkage only for well-separated clusters with no bridges.

Use Ward (or complete) when data has noise or unequal densities

Why: Ward minimizes within-cluster variance at every merge — it is robust to the chaining artifact because it prefers tight, low-variance merges over distant minimum-distance connections.

28. Which of these survive contact with Lesson 54: DBSCAN & Hierarchical Clustering?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
DBSCAN takes epsilon (neighborhood radius) and min_samples (density threshold) — no k needed.; Every point gets exactly one classification based on its epsilon-neighborhood.; Hierarchical clustering builds a tree of merges (or splits) — no k required at fit time; choose k by cutting the tree.
Breaks
Use a generous epsilon so no points are labeled noise — every point deserves a cluster.; Single linkage is simple — just use it for everything. It finds the minimum distance, so it can't go wrong.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 54: DBSCAN & Hierarchical Clustering puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

29. Cluster validation metrics

Section

Part 3 of 3

30. Silhouette score — no labels needed

Concept

Silhouette measures how similar a point is to its own cluster vs the nearest other cluster. Range: -1 to +1; higher is better.

\[ s(i) = \frac{b(i) - a(i)}{\max(a(i),\, b(i))} \]

a(i) = mean intra-cluster distance; b(i) = mean distance to nearest other cluster. On blobs: k=3 gives silhouette 0.8448 vs k=4 giving 0.6584 — the true k wins.

31. Teach it back: Silhouette score — no labels needed

Explain it

Discussion prompt

Explain Silhouette score — no labels needed to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Silhouette measures how similar a point is to its own cluster vs the nearest other cluster. Range: -1 to +1; higher is better.

32. ARI and NMI — when you have ground-truth labels

Concept

ARI (Adjusted Rand Index): measures agreement between predicted and true cluster assignments, corrected for chance. Range: -1 to 1; random ≈ 0, perfect = 1.

prediction qualityARINMI
perfect match1.00001.0000
random assignment-0.11110.0817
partial agreement0.46–0.87varies

NMI (Normalized Mutual Information) captures information overlap, range 0-1. Use ARI when cluster sizes matter; NMI when you want a symmetric information measure.

33. Fill in: NMI for ARI and NMI — when you have ground-truth…

Comparison

Comparison matrix

From ARI and NMI — when you have ground-truth labels: refill the NMI column from what you know. The rest of the table is as it appeared.

prediction qualityARINMI
perfect match1.00001.0000
random assignment-0.11110.0817
partial agreement0.46–0.87varies

34. Guess the shape of the answer: Validation: 3 datasets × 4 methods

Estimation

Predict first

Run KMeans, GMM, DBSCAN, and Ward on blobs / moons / circles — one table shows the decisive pattern.

Commit before you compute: what does Validation: 3 datasets × 4 methods come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: blobs: all methods ARI~1.0 | moons: KMeans 0.217, DBSCAN 1.000 | circles: KMeans -0.005, DBSCAN 1.000

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On convex well-separated blobs any method works.

35. Validation: 3 datasets × 4 methods

Worked example

Run KMeans, GMM, DBSCAN, and Ward on blobs / moons / circles — one table shows the decisive pattern.

from sklearn.datasets import make_blobs, make_moons, make_circles
from sklearn.cluster import DBSCAN, KMeans, AgglomerativeClustering
from sklearn.mixture import GaussianMixture
from sklearn.metrics import adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster
import warnings; warnings.filterwarnings('ignore')

for dname, (X, y), eps in [
    ('blobs',   make_blobs(200, centers=3, random_state=42),  1.5),
    ('moons',   make_moons(200, noise=0.05, random_state=42), 0.2),
    ('circles', make_circles(200, noise=0.05, factor=0.5, random_state=42), 0.2),
]:
    k = len(set(y))
    km  = KMeans(k, n_init=10, random_state=42).fit_predict(X)
    db  = DBSCAN(eps=eps, min_samples=5).fit_predict(X)
    Z   = linkage(X, 'ward'); wd = fcluster(Z, k, criterion='maxclust')
    for mn, lb in [('KM', km), ('DB', db), ('Ward', wd)]:
        mask = lb != -1
        ari = adjusted_rand_score(y[mask], lb[mask])
        print(f'{dname:8s} {mn:6s} ARI={ari:.3f}')

blobs: all methods ARI~1.0 | moons: KMeans 0.217, DBSCAN 1.000 | circles: KMeans -0.005, DBSCAN 1.000

Why: On convex well-separated blobs any method works. On non-convex shapes (moons, circles) only density-based DBSCAN recovers the true clusters; centroid and linkage methods fail completely.

datasetKMeans ARIDBSCAN ARIWard ARI
blobs1.0001.0001.000
moons0.2171.0000.446
circles-0.0051.000-0.005

36. What stays fixed: Validation: 3 datasets × 4 methods

Invariant

Step through it

Step through Validation: 3 datasets × 4 methods one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: dataset is blobs
  2. Step 2: dataset is moons
  3. Step 3: dataset is circles

37. Rebuild the recipe: Algorithm selection guide

Ranking

Put in order

These are the steps of Algorithm selection guide, scrambled. Put them back in order before the next slide shows you.

  1. Convex, similar-sized clusters, known k: k-Means (fast, interpretable)
  2. Soft / elliptical clusters, known k: GMM (Lesson 53)
  3. Arbitrary shape, unknown k, noisy data: DBSCAN (eps + min_samples from k-NN plot)
  4. Unknown k, need hierarchy or dendrograms: Agglomerative with Ward linkage
  5. Validate (no labels): silhouette score over a range of k
  6. Validate (labels available): ARI (agreement) or NMI (information)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

38. Algorithm selection guide

Pattern

  1. Convex, similar-sized clusters, known k: k-Means (fast, interpretable)
  2. Soft / elliptical clusters, known k: GMM (Lesson 53)
  3. Arbitrary shape, unknown k, noisy data: DBSCAN (eps + min_samples from k-NN plot)
  4. Unknown k, need hierarchy or dendrograms: Agglomerative with Ward linkage
  5. Validate (no labels): silhouette score over a range of k
  6. Validate (labels available): ARI (agreement) or NMI (information)

39. Where does it stop working: Algorithm selection guide

Edge cases

Discussion prompt

Algorithm selection guide works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Convex, similar-sized clusters, known k: k-Means (fast, interpretable)
  2. Soft / elliptical clusters, known k: GMM (Lesson 53)
  3. Arbitrary shape, unknown k, noisy data: DBSCAN (eps + min_samples from k-NN plot)
  4. Unknown k, need hierarchy or dendrograms: Agglomerative with Ward linkage
  5. Validate (no labels): silhouette score over a range of k
  6. Validate (labels available): ARI (agreement) or NMI (information)

40. Rule out three: Check yourself — point types

Elimination

Eliminate the wrong options

With eps=0.5 and min_samples=4, a point has exactly 3 other points within distance 0.5 and is itself within distance 0.5 of a core point. What type is it?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Border point
  • B. Core point
  • C. Noise point
  • D. Cannot be determined without knowing the full dataset

Survives elimination: A

Why: A core point needs >= min_samples (4) neighbors within eps. This point has only 3 — one short — so it is NOT a core point. But it is within eps of a core point, making it a border point (label = cluster of that core, not -1).

41. Check yourself — point types

Check

Reason through the DBSCAN definitions.

Check your understanding

With eps=0.5 and min_samples=4, a point has exactly 3 other points within distance 0.5 and is itself within distance 0.5 of a core point. What type is it?

  • A. Border point (correct)
  • B. Core point
  • C. Noise point
  • D. Cannot be determined without knowing the full dataset

Answer: A

Why: A core point needs >= min_samples (4) neighbors within eps. This point has only 3 — one short — so it is NOT a core point. But it is within eps of a core point, making it a border point (label = cluster of that core, not -1).

Why B tempts people
Core requires >= min_samples neighbors within eps. 3 < 4, so the density threshold is not met.
Why C tempts people
Noise means not reachable from any core point. Since this point is within eps of a core, it is density-reachable — it gets a cluster label, not -1.
Why D tempts people
The classification is fully determined by: (1) does the point itself have >= min_samples neighbors? No. (2) Is it within eps of a core point? Yes. That uniquely identifies a border point.

42. Answer it before you see the options: Check yourself — silhouette

Prediction

Predict first

Running Ward linkage on make_blobs(centers=3) and cutting at k=2, 3, 4, 5 yields silhouette scores 0.704, 0.845, 0.658, 0.500. Which k do you report as optimal?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: k = 3, highest silhouette

Why: Silhouette peaks at k=3 (0.845). This data was generated with 3 centers, so ARI also reaches 1.0 at k=3 — the two metrics agree. Always report the k that maximizes silhouette; k=2 under-clusters and k=4/5 over-splits.

43. Check yourself — silhouette

Check

Use the silhouette formula.

Check your understanding

Running Ward linkage on make_blobs(centers=3) and cutting at k=2, 3, 4, 5 yields silhouette scores 0.704, 0.845, 0.658, 0.500. Which k do you report as optimal?

  • A. k = 3, highest silhouette (correct)
  • B. k = 2, simplest model
  • C. k = 4, close to maximum but more granular
  • D. k = 5, most detail

Answer: A

Why: Silhouette peaks at k=3 (0.845). This data was generated with 3 centers, so ARI also reaches 1.0 at k=3 — the two metrics agree. Always report the k that maximizes silhouette; k=2 under-clusters and k=4/5 over-splits.

Why B tempts people
Parsimony matters, but not at the cost of fit quality. k=2 gets silhouette 0.704 vs 0.845 — the lower score means points are less coherent within clusters. Simplicity loses here.
Why C tempts people
k=4 silhouette is 0.658 < 0.845. A lower silhouette means more cross-contamination between adjacent clusters — the extra split is artificial.
Why D tempts people
k=5 silhouette is 0.500 — near the midpoint, indicating roughly as much inter-cluster similarity as intra-cluster. The extra clusters are not real structure.

44. Rule out three: Check yourself — linkage methods

Elimination

Eliminate the wrong options

Which linkage method is most susceptible to the 'chaining' problem, where a single outlier bridges two true clusters?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Single linkage
  • B. Complete linkage
  • C. Ward linkage
  • D. Average linkage

Survives elimination: A

Why: Single linkage uses the MINIMUM distance between clusters. A single outlier lying between two groups is close enough to both to trigger a merge — the chain collapses real clusters into one. Complete and Ward use maximum / variance-increase criteria that resist this.

45. Check yourself — linkage methods

Check

Recall the linkage trade-offs.

Check your understanding

Which linkage method is most susceptible to the 'chaining' problem, where a single outlier bridges two true clusters?

  • A. Single linkage (correct)
  • B. Complete linkage
  • C. Ward linkage
  • D. Average linkage

Answer: A

Why: Single linkage uses the MINIMUM distance between clusters. A single outlier lying between two groups is close enough to both to trigger a merge — the chain collapses real clusters into one. Complete and Ward use maximum / variance-increase criteria that resist this.

Why B tempts people
Complete linkage uses the MAX inter-cluster distance. An outlier doesn't produce a small max distance, so complete is resistant to chaining (though sensitive to outliers in a different way).
Why C tempts people
Ward minimizes the increase in total within-cluster variance. Bridging two real clusters always increases variance substantially — the merge cost is high, so Ward resists chaining.
Why D tempts people
Average linkage averages all pairwise distances. A single outlier can lower the average slightly but rarely enough to prematurely merge well-separated clusters.

46. Your turn: implement it

Section

Project

47. Project: clustering lab

Concept

Three milestones: DBSCAN on moons/circles, a dendrogram cut comparison, and the full 3×4 validation table.

#taskkey tool
1DBSCAN on make_moons and make_circlesDBSCAN, adjusted_rand_score
2Ward dendrogram, cut at k=2/3/4linkage, fcluster, silhouette_score
3Compare all methods on blobs/moons/circlesKMeans, GMM, DBSCAN, AgglomerativeClustering

Rules: print every ARI and silhouette before looking at the table on the next slide. Predict which method will fail on moons before running.

48. By analogy: Project: clustering lab

Analogy

Discussion prompt

Explain Project: clustering lab by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Three milestones: DBSCAN on moons/circles, a dendrogram cut comparison, and the full 3×4 validation table.

49. Milestone 1 — DBSCAN on moons

Worked example

Your turn: run DBSCAN (eps=0.2, min_samples=5) on make_moons. Predict: how many clusters, how many noise points?

Hint: import make_moons and DBSCAN; print len(set(labels)) - (1 if -1 in labels else 0) and (labels==-1).sum().

from sklearn.datasets import make_moons
from sklearn.cluster import DBSCAN
from sklearn.metrics import adjusted_rand_score

X, y = make_moons(n_samples=200, noise=0.05, random_state=42)
lbls = DBSCAN(eps=0.2, min_samples=5).fit_predict(X)
n_clust = len(set(lbls)) - (1 if -1 in lbls else 0)
print(n_clust, (lbls==-1).sum(), round(adjusted_rand_score(y, lbls), 4))
outputvalue
n_clusters2
n_noise0
ARI1.0000

50. Watch it run: Milestone 1 — DBSCAN on moons

Pattern

Step through it

Step through Milestone 1 — DBSCAN on moons one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: output is n_clusters
  2. Step 2: output is n_noise
  3. Step 3: output is ARI

51. Milestone 2 — Ward dendrogram cut

Worked example

Your turn: fit Ward linkage on make_blobs(centers=3), cut at k=2,3,4. Predict which k has the highest silhouette.

Hint: linkage(X, method='ward') then fcluster(Z, k, criterion='maxclust') and silhouette_score(X, lbl).

from sklearn.datasets import make_blobs
from sklearn.metrics import silhouette_score, adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster

X, y = make_blobs(n_samples=150, centers=3, cluster_std=1.0, random_state=42)
Z = linkage(X, method='ward')
for k in [2, 3, 4]:
    lbl = fcluster(Z, k, criterion='maxclust')
    print(k, round(silhouette_score(X,lbl),4), round(adjusted_rand_score(y,lbl),4))
ksilhouetteARI
20.70420.5681
30.84481.0000
40.65840.8695

52. What each one costs: Milestone 2 — Ward dendrogram cut

Trade off

Comparison matrix

From Milestone 2 — Ward dendrogram cut: every row here is a choice with a cost. Fill the silhouette column, then say which row you would actually pick and what you give up for it.

ksilhouetteARI
20.70420.5681
30.84481.0000
40.65840.8695

53. The full comparison program

Concept

from sklearn.datasets import make_blobs, make_moons, make_circles
from sklearn.cluster import DBSCAN, KMeans
from sklearn.mixture import GaussianMixture
from sklearn.metrics import adjusted_rand_score
from scipy.cluster.hierarchy import linkage, fcluster
import warnings; warnings.filterwarnings('ignore')

for dname, (X, y), eps in [
    ('blobs',   make_blobs(200, centers=3, random_state=42),  1.5),
    ('moons',   make_moons(200, noise=0.05, random_state=42), 0.2),
    ('circles', make_circles(200, noise=0.05, factor=0.5, random_state=42), 0.2),
]:
    k  = len(set(y))
    km = KMeans(k, n_init=10, random_state=42).fit_predict(X)
    db = DBSCAN(eps=eps, min_samples=5).fit_predict(X)
    Z  = linkage(X, 'ward'); wd = fcluster(Z, k, criterion='maxclust')
    for mn, lb in [('KMeans',km),('DBSCAN',db),('Ward',wd)]:
        mask = lb != -1
        ari  = adjusted_rand_score(y[mask], lb[mask])
        print(f'{dname:8} {mn:8} ARI={ari:.3f}')
datasetKMeans ARIDBSCAN ARIWard ARI
blobs1.0001.0001.000
moons0.2171.0000.446
circles-0.0051.000-0.005

If you see DBSCAN at ARI=1.000 on moons and circles while KMeans sits near 0 — you have empirically proven why density-based methods exist.

54. Fill in: KMeans ARI for The full comparison program

Comparison

Comparison matrix

From The full comparison program: refill the KMeans ARI column from what you know. The rest of the table is as it appeared.

datasetKMeans ARIDBSCAN ARIWard ARI
blobs1.0001.0001.000
moons0.2171.0000.446
circles-0.0051.000-0.005

55. Show it off

Concept

Slides closed: explain (1) the three DBSCAN point types and what epsilon controls, (2) why Ward linkage is preferred over single for most problems, (3) when to use silhouette vs ARI.

Stretch (homework): add make_circles to Milestone 1; try different eps values to break DBSCAN on moons; cluster the load_digits PCA embedding with Ward and measure NMI against digit labels.

56. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed: explain (1) the three DBSCAN point types and what epsilon controls, (2) why Ward linkage is preferred over single for most problems, (3) when to use silhouette vs ARI.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Stretch (homework): add make_circles to Milestone 1; try different eps values to break DBSCAN on moons; cluster the load_digits PCA embedding with Ward and measure NMI against digit labels.

57. Connect it up: Lesson 54: DBSCAN & Hierarchical Clustering

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — DBSCAN: density-based clustering · Hierarchical clustering & dendrograms · Cluster validation metrics · Your turn: implement it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

58. What you can do now

Recap

ideathe one thing to remember
DBSCAN eps too largeclusters merge — use k-NN elbow to set eps
single linkagechaining risk — one bridge collapses two clusters
Ward linkageminimizes variance increase = best default
silhouetteno labels needed; peak = right k
ARIcorrected for chance; random baseline = 0

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 54 (Week 19 — DBSCAN & Hierarchical Clustering) — Barron · USAAIO Round 2 Preparation, 2026
  2. DBSCAN on make_moons/make_circles, Ward linkage on make_blobs, silhouette/ARI/NMI all verified — sklearn 1.x + scipy, numpy 2.2.6, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108