A PhD-level deck of 52 slides, about two hours long, on why SIFT and RANSAC landmark matching breaks down across before-and-after vitiligo skin photos, and how to fix it. It sets up a diagnosis-first framework - detect, describe, match, estimate - and then applies three principled upgrades to a real OpenCV pipeline: RootSIFT, built on the Hellinger kernel and the L1-normalize-then-square-root identity; a mutual nearest-neighbor cross-check layered on top of Lowe's ratio test, traded off as precision against recall for a robust estimator; and MAGSAC++ in place of vanilla RANSAC, for its sigma-consensus and threshold insensitivity. The deck includes five traps, five checks, a fully scaffolded your-turn implementation, and an end-to-end ablation. Every number and code path was executed against OpenCV 4.12 before the deck was written.
Subject: Computer Vision · 84 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
Computer Vision · Image Registration
Why landmark matching breaks down between before and after treatment photos — and three principled fixes you can defend at a whiteboard.
Objectives
Your pipeline already does the right things — adaptive SIFT, FLANN + ratio test, spatial-prior gating, homography/affine RANSAC. Today we make it accurate and consistent. By the end you can:
Concept
Figure (svg): Two skin photos with three landmarks each; two correct dashed correspondences and one wrong red one.
We want a stable correspondence set between the same skin region photographed weeks apart, then a geometric model that maps Visit 1 to Visit 2 so lesions can be measured.
Between visits the appearance changes (lighting, white-balance, skin texture, mild non-rigid deformation) while the geometry is nearly a similarity. That split — appearance drifts, geometry is stable — is the whole key.
When it 'breaks', the symptom is vague: inconsistent matches. Our job is to make that symptom measurable and attack the right stage.
Counterexample
Discussion prompt
We want a stable correspondence set between the same skin region photographed weeks apart, then a geometric model that maps Visit 1 to Visit 2 so lesions can be measured.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
When it 'breaks', the symptom is vague: inconsistent matches. Our job is to make that symptom measurable and attack the right stage.
Concept
Four stops. The first is a habit; the last three are the fixes we will implement in your stable_landmark_labeler_v2.py.
Matching
Match the pairs
From Today's four levers — match each one to what it actually does. The descriptions have been shuffled.
Why: Diagnose, RootSIFT, Cross-check, MAGSAC++ are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Section
Part 1
Concept
Figure (svg): Vertical flow: Detect, Describe, Match, Estimate, each feeding the next.
Every feature-based registration is the same four stages. A failure in any one looks identical downstream — bad final matches — so you cannot fix it by staring at the output.
Localization principle — Each stage emits a count you can read. Find the first stage whose number is wrong; tune only that stage. Everything downstream is a symptom, not the cause.
Analogy
Discussion prompt
Explain The pipeline is four stages by analogy to something with no Computer Vision in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Every feature-based registration is the same four stages. A failure in any one looks identical downstream — bad final matches — so you cannot fix it by staring at the output.
Intuition
Phrase the complaint 'SIFT and RANSAC don't match consistently' as four testable hypotheses:
RootSIFT attacks #2, cross-check attacks #3, MAGSAC attacks #4. But you must confirm which one before you reach for a fix.
Explain it
Discussion prompt
Explain Four distinct failure modes to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Phrase the complaint 'SIFT and RANSAC don't match consistently' as four testable hypotheses:
Concept
Your MatchQuality already records the right numbers. Read them as a funnel — each row should fall by a sane factor, not collapse.
| Metric | Stage | Healthy reading |
|---|---|---|
| keypoints / image | Detect | hundreds–thousands on both |
| raw_matches | Match (kNN) | ≈ min(#kp) — every query gets a neighbor |
| good_matches_after_lowe | Match (ratio) | a sizeable fraction survives |
| inlier_ratio | Estimate | well above the RANSAC failure floor |
| average_reprojection_error | Estimate | ≈ pixel noise, not tens of px |
Comparison
Comparison matrix
From Instrument every stage: refill the Stage column from what you know. The rest of the table is as it appeared.
| Metric | Stage | Healthy reading |
|---|---|---|
| keypoints / image | Detect | hundreds–thousands on both |
| raw_matches | Match (kNN) | ≈ min(#kp) — every query gets a neighbor |
| good_matches_after_lowe | Match (ratio) | a sizeable fraction survives |
| inlier_ratio | Estimate | well above the RANSAC failure floor |
| average_reprojection_error | Estimate | ≈ pixel noise, not tens of px |
Worked example
A failing run on a self-similar (pore-textured) pair prints something like this. Plenty of keypoints, plenty of matches — yet the geometry is garbage.
keypoints image1 / image2 : 2861 / 2774
raw_matches : 2861
good_matches_after_lowe : 86
ransac_inliers : 11
inlier_ratio : 0.128
avg_reprojection_error : 27.4 pxDetection is fine (thousands). kNN is fine. The ratio test already cut 2861 → 86. The killer is the last two rows: a 13% inlier ratio with 27 px error means the 86 'good' matches are mostly wrong.
| Row | Verdict | Implication |
|---|---|---|
| keypoints high | Detect OK | do NOT lower contrastThreshold |
| good=86, inliers=11 | Match contaminated | fix the matcher (purity) |
| error 27 px | Estimate starved | estimator can't save a bad set |
Diagnosis: stage #3. The fix is descriptor quality + match purity — not a looser RANSAC.
Trade off
Comparison matrix
From Reading a real diagnostic run: every row here is a choice with a cost. Fill the Verdict column, then say which row you would actually pick and what you give up for it.
| Row | Verdict | Implication |
|---|---|---|
| keypoints high | Detect OK | do NOT lower contrastThreshold |
| good=86, inliers=11 | Match contaminated | fix the matcher (purity) |
| error 27 px | Estimate starved | estimator can't save a bad set |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Matches look bad, so loosen everything at once: ratio → 0.9, RANSAC threshold → 10.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: You get more 'matches' and more 'inliers', so it feels better — but you have masked the failing stage and inflated the outlier rate fed to the estimator.
Print the per-stage funnel first: keypoints, raw, ratio survivors, inlier ratio, reprojection error.
Why: You get more 'matches' and more 'inliers', so it feels better — but you have masked the failing stage and inflated the outlier rate fed to the estimator.
Trap
Matches look bad, so loosen everything at once: ratio → 0.9, RANSAC threshold → 10.
Turn every knob outward simultaneously
Why: You get more 'matches' and more 'inliers', so it feels better — but you have masked the failing stage and inflated the outlier rate fed to the estimator.
Result: a confident-looking homography fit to contamination. Worse, and now untraceable.
Print the per-stage funnel first: keypoints, raw, ratio survivors, inlier ratio, reprojection error.
Find the first wrong number, tune only that stage
Why: Thousands of keypoints but a 13% inlier ratio points at the matcher. Loosening RANSAC there only fits the noise faster.
One change, one measured effect. That is how you keep the result reproducible.
Ranking
Put in order
These are the steps of Recipe: diagnosis-first, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Before any parameter change, run this loop. It is the discipline the rest of the deck depends on.
Section
Part 2
Concept
A SIFT descriptor is a 128-D histogram: a 4×4 grid of cells, each holding an 8-bin histogram of gradient orientations, weighted by gradient magnitude.
\[ d \in \mathbb{R}^{128}_{\ge 0},\qquad d = (d_1, d_2, \dots, d_{128}) \]
Two things matter: the entries are non-negative counts, and the vector is a histogram, not an arbitrary point in space. The default matcher ignores both facts.
Intuition
OpenCV's default compares descriptors with L2 (Euclidean) distance. For histograms that has two pathologies.
Between before/after photos, both pathologies fire at once. We want a metric that compresses large bins and behaves like a proper histogram comparison.
Concept
For histograms, the natural similarity is the Bhattacharyya coefficient, and the matching distance is Hellinger:
\[ \mathrm{BC}(p,q)=\sum_i \sqrt{p_i\,q_i}, \qquad H(p,q)=\tfrac{1}{\sqrt{2}}\,\bigl\lVert \sqrt{p}-\sqrt{q}\bigr\rVert_2 \]
The square root is the whole point: it down-weights large bins and turns a histogram comparison into a well-behaved chi-square-family metric. But our matcher only speaks Euclidean. The trick is to make Euclidean compute Hellinger.
Worked example
L1-normalize each descriptor (so it sums to 1, a probability vector p), then take the element-wise square root. Now the ordinary Euclidean distance between two transformed vectors equals the Hellinger distance between the originals:
\[ \bigl\lVert \sqrt{p}-\sqrt{q}\bigr\rVert_2^2 \;=\; \sum_i p_i - 2\sum_i\sqrt{p_i q_i} + \sum_i q_i \;=\; 2\Bigl(1-\sum_i \sqrt{p_i q_i}\Bigr) \]
Because p and q each sum to 1, the cross term is the Bhattacharyya coefficient. Verified numerically on two random 128-bin histograms:
| Quantity | Value |
|---|---|
| sum(p), sum(q) after L1-norm | 1.000000, 1.000000 |
| ‖√p − √q‖² (Euclidean, squared) | 0.220367 |
| 2·(1 − BC), BC = Σ√(pᵢqᵢ) | 0.220367 (BC = 0.889817) |
Identical to six decimals. So the matcher stays Euclidean — zero extra cost at match time — yet it now measures the right thing. That is why RootSIFT is free.
Worked example
The transform is four lines. The eps guards an all-zero descriptor; the astype keeps FLANN happy with float32.
def to_rootsift(descriptors, eps=1e-7):
if descriptors is None or len(descriptors) == 0:
return descriptors
descriptors = descriptors.astype(np.float32)
l1 = np.sum(np.abs(descriptors), axis=1, keepdims=True)
descriptors = descriptors / (l1 + eps) # L1-normalize -> p
descriptors = np.sqrt(descriptors) # element-wise sqrt
return descriptorsA correctness invariant worth asserting in a test: after the transform every row is an L2 unit vector, because Σ(√pᵢ)² = Σpᵢ = 1.
| Check on 5 random rows | Before | After to_rootsift |
|---|---|---|
| row L1 norm Σ|dᵢ| | arbitrary | 1.0 (then sqrt'd) |
| row L2 norm² Σ(d̃ᵢ)² | arbitrary | 1.000000 |
| dtype | float32 | float32 |
Call it right where descriptors are born — inside detect_sift_features, immediately after sift.detectAndCompute.
Concept
RootSIFT changes the coordinate system of the descriptor. FLANN compares whatever rows you stack — it does no normalization. So every descriptor source must be transformed, or you are comparing apples to a different fruit.
Your pipeline has a second source: augment_with_stable_spot_keypoints computes extra SIFT descriptors for mole-like spots and np.vstacks them onto the main set. Those must be RootSIFT'd too.
| Descriptor source | Transform applied? | Consequence if not |
|---|---|---|
| main detectAndCompute | yes | — |
| augmented stable spots | must match main | mole landmarks silently stop matching |
Anomaly
Predict first
A student writes this, and it looks reasonable:
RootSIFT the main descriptors, but leave the augmented stable-spot descriptors as raw SIFT, then vstack and match in one FLANN index.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Their scales differ by orders of magnitude (unit-norm vs raw counts).
Apply to_rootsift to the spot descriptors before the vstack.
Why: Their scales differ by orders of magnitude (unit-norm vs raw counts). The L2 distance is meaningless, so the nearest neighbor is essentially random.
Trap
RootSIFT the main descriptors, but leave the augmented stable-spot descriptors as raw SIFT, then vstack and match in one FLANN index.
Match a RootSIFT query against a raw-SIFT train vector
Why: Their scales differ by orders of magnitude (unit-norm vs raw counts). The L2 distance is meaningless, so the nearest neighbor is essentially random.
On the test pair the mole keypoints produced 0 correct matches — and no error was raised. A silent failure that costs you exactly the most stable landmarks.
Apply to_rootsift to the spot descriptors before the vstack.
Transform every source into the same Hellinger space
Why: Now all rows are unit-norm RootSIFT; distances are comparable across the whole stacked matrix.
Same pair, fixed: 10 matches at 90% geometric correctness. The landmarks come back.
Elimination
Eliminate the wrong options
You apply RootSIFT to the main detector but forget it inside augment_with_stable_spot_keypoints, so the extra mole keypoints keep raw SIFT descriptors. The two sets are vstacked and matched with one FLANN index. What happens?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: RootSIFT is an element-wise change of coordinates (L1-normalize then sqrt), leaving the descriptor 128-D but unit-L2-norm. Raw SIFT rows have arbitrary, much larger magnitudes. FLANN's kd-tree searches the raw stacked vectors with no normalization, so a unit-norm query and a large-magnitude train vector produce a meaningless distance — the mole keypoints match essentially at random and drop out as outliers, with no exception raised.
Check
Reason about the mechanism before answering.
Check your understanding
You apply RootSIFT to the main detector but forget it inside augment_with_stable_spot_keypoints, so the extra mole keypoints keep raw SIFT descriptors. The two sets are vstacked and matched with one FLANN index. What happens?
Answer: A
Why: RootSIFT is an element-wise change of coordinates (L1-normalize then sqrt), leaving the descriptor 128-D but unit-L2-norm. Raw SIFT rows have arbitrary, much larger magnitudes. FLANN's kd-tree searches the raw stacked vectors with no normalization, so a unit-norm query and a large-magnitude train vector produce a meaningless distance — the mole keypoints match essentially at random and drop out as outliers, with no exception raised.
Section
Part 3
Concept
For each query descriptor we take its two nearest neighbors and keep the match only if the best is decisively closer than the runner-up — Lowe's ratio test.
raw = flann.knnMatch(des1, des2, k=2)
good = []
for m, n in raw:
if m.distance < lowe_ratio * n.distance: # e.g. 0.7
good.append(m)The logic: if the 1st and 2nd neighbors are nearly tied, the match is ambiguous and probably wrong. The ratio is a distinctiveness gate.
| lowe_ratio | Behavior | On repetitive skin |
|---|---|---|
| 0.6–0.7 | strict, distinctive only | few but cleaner matches |
| 0.8 | permissive | more matches, more ambiguity |
| 0.9 | very loose | floods in near-duplicate (wrong) matches |
Socratic
Discussion prompt
For each query descriptor we take its two nearest neighbors and keep the match only if the best is decisively closer than the runner-up — Lowe's ratio test.
Suppose that were not true. What is the first thing in Robust SIFT + RANSAC Landmark Matching (Before/After Registration) that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Answer:
The logic: if the 1st and 2nd neighbors are nearly tied, the match is ambiguous and probably wrong. The ratio is a distinctiveness gate.
Intuition
Figure (svg): Point a's nearest in image 2 is b, but b's nearest back in image 1 is a different point a-prime, so the match is not mutual.
The ratio test is one-directional. It asks: among image-2 points, is a's best neighbor distinctive? It never asks whether that image-2 point agrees that a is its best.
On self-similar skin, a's nearest neighbor b can pass the ratio test while b's own nearest neighbor back is a different point a'. That asymmetric, one-sided match is wrong — and it survives.
This is the dominant failure mode behind 'good=86 but only ~12% correct'.
Concept
Mutual (symmetric) NN — Keep a match a→b only if, matching in the reverse direction, b's nearest neighbor is a. The agreement must be two-sided.
Layer it after the ratio test, not instead of it. Ratio removes ambiguous matches; mutual-NN removes asymmetric ones. Different failure modes, both real.
Definition probe
Sort into buckets
Every line below is part of the definition of Localization principle or of Mutual (symmetric) NN — one or the other, never both. Put each where it belongs.
Worked example
A DMatch from knnMatch(desA, desB) has queryIdx into desA and trainIdx into desB. Get this backwards and the filter silently deletes everything.
reverse = flann.knnMatch(des2, des1, k=1) # note: des2 first
best_back = {}
for pair in reverse:
if pair:
rm = pair[0]
best_back[rm.queryIdx] = rm.trainIdx # des2 idx -> des1 idx
mutual = [m for m in good if best_back.get(m.trainIdx) == m.queryIdx]In the reverse pass rm.queryIdx indexes des2; we map it to its best des1 index. A forward match m survives only if that reverse map sends m.trainIdx back to m.queryIdx.
| Symbol | Indexes into | Meaning |
|---|---|---|
| m.queryIdx | des1 (before) | the image-1 keypoint |
| m.trainIdx | des2 (after) | its claimed image-2 match |
| best_back[trainIdx] | des1 | who image-2 point picks in return |
| survive iff == | — | two-sided agreement |
Worked example
Same before/after pair, RootSIFT descriptors, lowe_ratio = 0.7. We score a match as correct if it lands within 5 px of the known ground-truth transform.
# ratio test only
_, good = match_descriptors(des1, des2, 0.7, cross_check=False)
# ratio test + mutual NN
_, good = match_descriptors(des1, des2, 0.7, cross_check=True)| Config | good matches | correct | correct % |
|---|---|---|---|
| ratio only | 86 | ≈10 | 11.6% |
| ratio + cross-check | 10 | 9 | 90.0% |
Recall fell (86 → 10) but precision went 11.6% → 90%. For a robust estimator that is an excellent trade — and the next two slides say exactly why.
Concept
Crossing RootSIFT with cross-check on the same pair (matcher run through your adaptive path). Read the correct % column — that is match purity.
| RootSIFT | cross-check | good | correct % |
|---|---|---|---|
| off | off | 87 | 13.8% |
| off | on | 39 | 59.0% |
| on | off | 85 | 11.8% |
| on | on | 37 | 70.3% |
Cross-check is the big mover (13.8 → 59). RootSIFT compounds it (59 → 70) and — from the last slide — rescues the mole-spot landmarks the mixed-space bug was killing. Together: ~14% → ~70%, a 5× purer set.
Comparison
Comparison matrix
From The full ablation: refill the correct % column from what you know. The rest of the table is as it appeared.
| RootSIFT | cross-check | good | correct % |
|---|---|---|---|
| off | off | 87 | 13.8% |
| off | on | 39 | 59.0% |
| on | off | 85 | 11.8% |
| on | on | 37 | 70.3% |
Intuition
RANSAC's iteration count to hit a good sample with confidence p, given inlier ratio w and sample size s, is:
\[ k \;=\; \frac{\log(1-p)}{\log\!\bigl(1-w^{\,s}\bigr)} \]
For a homography s = 4. At w = 0.14, w⁴ ≈ 0.0004, so you need thousands of samples to even see a clean one. At w = 0.70, w⁴ ≈ 0.24 — a handful of samples suffice.
| inlier ratio w | w⁴ | samples for 99% conf |
|---|---|---|
| 0.14 | 0.00038 | ≈ 12000 |
| 0.59 | 0.121 | ≈ 36 |
| 0.70 | 0.240 | ≈ 17 |
So a smaller, purer set isn't just cleaner — it makes robust estimation tractable. Throwing away 80% of matches to triple w is a bargain.
Pattern
Step through it
Step through Why precision beats recall for RANSAC one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Build the reverse map, then keep m where best_back[m.queryIdx] == m.trainIdx.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: queryIdx indexes des1, but the reverse search keys are des2 indices.
Keep m where best_back[m.trainIdx] == m.queryIdx.
Why: queryIdx indexes des1, but the reverse search keys are des2 indices. You look up the wrong table entry for almost every match.
Trap
Build the reverse map, then keep m where best_back[m.queryIdx] == m.trainIdx.
Index the reverse map by queryIdx
Why: queryIdx indexes des1, but the reverse search keys are des2 indices. You look up the wrong table entry for almost every match.
Symptom: cross-check 'works' but deletes nearly all matches, or keeps a random few. Easy to mistake for 'cross-check is too strict'.
Keep m where best_back[m.trainIdx] == m.queryIdx.
Index the reverse map by trainIdx (the des2 side)
Why: The reverse match's queryIdx is the des2 point; key the map by m.trainIdx and require it points back to m.queryIdx. Symmetry is enforced on the correct indices.
Now the survivor count is sane and the precision jump (e.g. 13.8 → 59%) appears.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The reverse match's queryIdx is the des2 point; key the map by m.trainIdx and require it points back to m.queryIdx. Symmetry is enforced on the correct indices.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
queryIdx indexes des1, but the reverse search keys are des2 indices. You look up the wrong table entry for almost every match.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Too few matches survive, so push lowe_ratio to 0.9.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: On repetitive skin the 1st and 2nd neighbors are near-duplicates; 0.9 admits exactly the ambiguous matches that are most often geometrically wrong.
Keep the ratio tight; recover recall with cross-check, more keypoints, and the spatial prior.
Why: On repetitive skin the 1st and 2nd neighbors are near-duplicates; 0.9 admits exactly the ambiguous matches that are most often geometrically wrong.
Trap
Too few matches survive, so push lowe_ratio to 0.9.
Accept matches where the best barely beats the 2nd-best
Why: On repetitive skin the 1st and 2nd neighbors are near-duplicates; 0.9 admits exactly the ambiguous matches that are most often geometrically wrong.
More matches, lower inlier ratio, harder RANSAC. You bought recall and paid in purity — the wrong direction.
Keep the ratio tight; recover recall with cross-check, more keypoints, and the spatial prior.
Tighten ratio, then add mutual-NN
Why: Purity first. A smaller, cleaner set raises w and stabilizes the estimator; you add back true matches by detecting more candidates, not by loosening the gate.
If you still lack matches, the cause is upstream (detection/description), not the ratio.
Prediction
Predict first
On self-similar skin the ratio test alone returns 86 'good' matches but only ~12% are geometrically correct. Adding a mutual nearest-neighbor cross-check drops it to 10 matches at ~90% correct. Why is this usually the right trade for a RANSAC-based estimator?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: RANSAC's required iterations grow steeply as the inlier ratio falls, so trading recall for a far higher inlier ratio makes robust estimation faster and more reliable.
Why: The sample count k = log(1−p)/log(1−w^s) blows up as the inlier ratio w drops; for a homography s=4, so w from 0.12 to 0.90 changes w⁴ by ~600×. A smaller but far purer set makes a clean minimal sample likely within a handful of iterations and yields a more accurate fit. Trading recall for precision is therefore exactly what a robust estimator wants.
Check
Tie it back to the estimator.
Check your understanding
On self-similar skin the ratio test alone returns 86 'good' matches but only ~12% are geometrically correct. Adding a mutual nearest-neighbor cross-check drops it to 10 matches at ~90% correct. Why is this usually the right trade for a RANSAC-based estimator?
Answer: A
Why: The sample count k = log(1−p)/log(1−w^s) blows up as the inlier ratio w drops; for a homography s=4, so w from 0.12 to 0.90 changes w⁴ by ~600×. A smaller but far purer set makes a clean minimal sample likely within a handful of iterations and yields a more accurate fit. Trading recall for precision is therefore exactly what a robust estimator wants.
Commit first
Predict first
Keypoint counts are healthy on both images (thousands each) and raw match count is high, but RANSAC returns an inlier ratio of 0.03 and findHomography's average reprojection error is large. Which stage is the prime suspect?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: Matching: many descriptor matches form but few are mutually consistent, so the correspondence set is contaminated before geometry.
Why: A 3% inlier ratio with large reprojection error, despite healthy keypoint and raw-match counts, is the signature of a contaminated match set: descriptors are being paired but the pairs are mostly geometrically inconsistent. The fix is upstream of the estimator — descriptor quality (RootSIFT) and match purity (cross-check) — not a different RANSAC threshold.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Use the funnel.
Check your understanding
Keypoint counts are healthy on both images (thousands each) and raw match count is high, but RANSAC returns an inlier ratio of 0.03 and findHomography's average reprojection error is large. Which stage is the prime suspect?
Answer: A
Why: A 3% inlier ratio with large reprojection error, despite healthy keypoint and raw-match counts, is the signature of a contaminated match set: descriptors are being paired but the pairs are mostly geometrically inconsistent. The fix is upstream of the estimator — descriptor quality (RootSIFT) and match purity (cross-check) — not a different RANSAC threshold.
Section
Part 4
Concept
Repeat: draw a minimal sample (4 points for a homography), fit a candidate model, count inliers — points whose reprojection error is below a threshold τ — and keep the model with the most. Then refit on all inliers.
\[ \text{inlier}(i) \iff r_i(\theta) < \tau,\qquad \hat\theta = \arg\max_\theta \bigl|\{\,i : r_i(\theta) < \tau\,\}\bigr| \]
Everything hinges on one number, τ — a hard binary cut on the residual. That single hard threshold is the brittleness we are about to remove.
Intuition
τ is really a claim: 'true correspondences have residual below this.' But the residual of a genuine inlier is set by keypoint localization noise σ — which you don't know and which changes per image pair.
That is exactly the 'works on one pair, fails on another' complaint. The cure is to stop committing to a single τ.
Concept
Figure (svg): Weight versus residual: RANSAC is a step function dropping to zero at tau; MAGSAC is a smooth decaying weight.
MAGSAC++ refuses to pick one τ. It treats the noise scale σ as unknown and integrates the model quality over a range of σ — 'σ-consensus' — so no point is hard-classified as in or out.
\[ Q(\theta) \;=\; \int_{0}^{\sigma_{\max}} \; \sum_i \rho_\sigma\!\bigl(r_i(\theta)\bigr)\; p(\sigma)\, d\sigma \]
Each point gets a soft weight that decays with its residual, and the final model is an iteratively-reweighted least squares fit. The threshold you pass becomes a loose upper bound on σ, not a knife-edge.
Worked example
Synthetic benchmark: 200 true correspondences + 120 random outliers (38% contamination), 1 px Gaussian noise, known homography. We sweep the threshold and measure the corner reprojection error of the recovered model against ground truth.
for t in (1.0, 2.0, 3.0, 5.0, 8.0):
Hr, _ = cv2.findHomography(src, dst, cv2.RANSAC, t)
Hm, _ = cv2.findHomography(src, dst, cv2.USAC_MAGSAC, t,
maxIters=10000, confidence=0.9999)
# compare corner error of Hr vs Hm to the true H| threshold | RANSAC err (px) | RANSAC inliers | MAGSAC err (px) | MAGSAC inliers |
|---|---|---|---|---|
| 1.0 | 1.001 | 77 / 200 | 0.178 | 81 / 200 |
| 2.0 | 0.565 | 168 | 0.178 | 169 |
| 3.0 | 0.257 | 198 | 0.181 | 198 |
| 5.0 | 0.175 | 200 | 0.181 | 200 |
| 8.0 | 0.175 | 200 | 0.181 | 200 |
RANSAC's error swings 5.7× (1.001 → 0.175) and at τ=1 it keeps only 77 of 200 true inliers. MAGSAC sits at ~0.18 px from τ=1 to τ=8 — essentially flat. That flatness is the whole point: you stop having to tune τ per pair.
Discrimination
Sort into buckets
Sort these by MAGSAC err (px), from memory, without looking back at Threshold-insensitivity, measured. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
Detect the flag once at import — older OpenCV builds lack USAC and we degrade gracefully. Then pass it where you already call findHomography.
_H_METHOD = getattr(cv2, "USAC_MAGSAC", cv2.RANSAC)
try:
H, mask = cv2.findHomography(src, dst, _H_METHOD, threshold,
maxIters=10000, confidence=0.9999)
except cv2.error:
# some builds advertise the flag but reject the kwargs
H, mask = cv2.findHomography(src, dst, cv2.RANSAC, threshold)maxIters and confidence matter: MAGSAC benefits from a high iteration budget, and the cost is cheap once the match set is small and pure.
| Build | _H_METHOD resolves to | Behavior |
|---|---|---|
| OpenCV ≥ 4.5 (yours: 4.12) | cv2.USAC_MAGSAC | σ-consensus, threshold-stable |
| older / no USAC | cv2.RANSAC | unchanged, still runs |
| USAC flag but kwargs rejected | fallback in except | plain RANSAC, no crash |
Concept
A homography has 8 DOF (full projective); a similarity / partial-affine has 4 (rotation, uniform scale, translation). More DOF fit the inliers better but also fit the noise better — classic overfitting.
Your compute_geometric_model_ransac already fits both and ranks them by (inlier count, −avg error, model bonus). The right instinct: prefer the lower-DOF model unless the extra DOF earn a real inlier gain.
| Model | DOF | Use when |
|---|---|---|
| Partial affine (similarity) | 4 | camera roughly fronto-parallel, same distance — the usual clinical setup |
| Full homography | 8 | real perspective change between visits |
| (neither, globally) | — | strong non-rigid skin deformation — see next slide |
Intuition
A homography assumes one flat, rigid surface viewed under projection. Skin stretches and bends between visits, so a single global model can never make every landmark agree — there is genuine model error, not just noise.
Knowing the model is wrong is itself a diagnosis — it tells you no estimator tweak will close the gap.
Explain it
Discussion prompt
Explain Skin is not a rigid plane to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Knowing the model is wrong is itself a diagnosis — it tells you no estimator tweak will close the gap.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The homography looks loose, so drop --ransac-threshold to 1.0 to 'keep only tight inliers'.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A threshold at the noise scale rejects genuine inliers that happen to sit just outside: 77 of 200 kept, corner error ~1.0 px — looser, not stricter.
Let MAGSAC marginalize over σ; pass the threshold as a loose upper bound.
Why: A threshold at the noise scale rejects genuine inliers that happen to sit just outside: 77 of 200 kept, corner error ~1.0 px — looser, not stricter.
Trap
The homography looks loose, so drop --ransac-threshold to 1.0 to 'keep only tight inliers'.
Use a 1 px hard cut on a 1 px-noise problem
Why: A threshold at the noise scale rejects genuine inliers that happen to sit just outside: 77 of 200 kept, corner error ~1.0 px — looser, not stricter.
You conflated 'strict threshold' with 'accurate model'. They're opposite when τ approaches σ.
Let MAGSAC marginalize over σ; pass the threshold as a loose upper bound.
Switch to cv2.USAC_MAGSAC, keep τ generous
Why: σ-consensus weights points by residual instead of cutting at τ; corner error holds at ~0.18 px from τ=1 to τ=8, and more true inliers are retained.
Accuracy now comes from the weighting scheme, not from a number you guessed.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
MatchQuality already records the right numbers. Read them as a funnel — each row should fall by a sane factor, not collapse.0.9, RANSAC threshold → 10.; RootSIFT the main descriptors, but leave the augmented stable-spot descriptors as raw SIFT, then vstack and match in one FLANN index.Prediction
Predict first
RANSAC is unstable: at threshold 1.0 the corner error is ~1.0 px and only 77/200 true inliers are kept, but at 5.0 px the error drops to ~0.18 px. You switch to USAC_MAGSAC. What is the principled reason MAGSAC's corner error stays ~0.18 px across thresholds 1–8?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: It marginalizes the model over a range of noise scales (σ-consensus) and weights every point by its residual, never committing to a single hard inlier/outlier cut.
Why: MAGSAC++ treats σ as unknown and integrates model quality over a σ range, giving each point a soft, residual-decaying weight and solving an iteratively-reweighted least squares. Because nothing is hard-classified at a single τ, the result barely depends on the threshold you pass (it acts as a loose upper bound) — hence the flat ~0.18 px error.
Check
Mechanism, not folklore.
Check your understanding
RANSAC is unstable: at threshold 1.0 the corner error is ~1.0 px and only 77/200 true inliers are kept, but at 5.0 px the error drops to ~0.18 px. You switch to USAC_MAGSAC. What is the principled reason MAGSAC's corner error stays ~0.18 px across thresholds 1–8?
Answer: A
Why: MAGSAC++ treats σ as unknown and integrates model quality over a σ range, giving each point a soft, residual-decaying weight and solving an iteratively-reweighted least squares. Because nothing is hard-classified at a single τ, the result barely depends on the threshold you pass (it acts as a loose upper bound) — hence the flat ~0.18 px error.
Section
Part 5
Concept
You will edit stable_landmark_labeler_v2.py yourself. Three requirements, each a function you already have a home for.
| # | Requirement | Where it goes |
|---|---|---|
| 1 | RootSIFT every descriptor source | to_rootsift + detect_sift_features + augment_with_stable_spot_keypoints |
| 2 | Mutual NN cross-check after ratio test | match_descriptors |
| 3 | MAGSAC++ with graceful fallback | compute_homography_ransac |
Analogy
Discussion prompt
Explain The brief by analogy to something with no Computer Vision in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
You will edit stable_landmark_labeler_v2.py yourself. Three requirements, each a function you already have a home for.
Worked example
Your turn. Write to_rootsift(descriptors), call it inside detect_sift_features after detectAndCompute, and (crucially) inside augment_with_stable_spot_keypoints before the vstack. Predict out loud: what should each row's L2 norm be afterward?
Hint. Two NumPy ops on axis=1: an L1 sum with keepdims=True, a divide, then np.sqrt. Guard None/empty and add an eps.
def to_rootsift(descriptors, eps=1e-7):
if descriptors is None or len(descriptors) == 0:
return descriptors
descriptors = descriptors.astype(np.float32)
l1 = np.sum(np.abs(descriptors), axis=1, keepdims=True)
return np.sqrt(descriptors / (l1 + eps))| Self-check | Expected |
|---|---|
| row L2 norm² after transform | 1.000000 |
| descriptor shape | unchanged (N × 128) |
| spot descriptors transformed too? | yes — or moles stop matching |
Worked example
Your turn. After the forward ratio test builds good, add an optional cross_check. Predict: on the contaminated pair, will good go up or down, and what should happen to correct %?
Hint. Run flann.knnMatch(des2, des1, k=1), build a dict from each reverse match's queryIdx to its trainIdx, then keep forward m only if best_back[m.trainIdx] == m.queryIdx. Mind which index is which.
if cross_check and good:
reverse = flann.knnMatch(des2, des1, k=1)
best_back = {p[0].queryIdx: p[0].trainIdx for p in reverse if p}
good = [m for m in good if best_back.get(m.trainIdx) == m.queryIdx]| Self-check (your pair) | Expected direction |
|---|---|
| good count | falls (recall ↓) |
| correct % | rises sharply (e.g. ~14% → ~59%) |
| inlier_ratio downstream | rises |
Worked example
Your turn. Resolve the method flag once at module load, then use it in compute_homography_ransac with a try/except cv2.error fallback. Predict: does your build (cv2.__version__) have USAC_MAGSAC?
Hint. getattr(cv2, "USAC_MAGSAC", cv2.RANSAC) gives a safe default. Pass maxIters=10000, confidence=0.9999; catch cv2.error and retry with plain cv2.RANSAC.
_H_METHOD = getattr(cv2, "USAC_MAGSAC", cv2.RANSAC)
try:
H, mask = cv2.findHomography(src, dst, _H_METHOD, threshold,
maxIters=10000, confidence=0.9999)
except cv2.error:
H, mask = cv2.findHomography(src, dst, cv2.RANSAC, threshold)| Self-check | Expected on OpenCV 4.12 |
|---|---|
| hasattr(cv2, 'USAC_MAGSAC') | True |
| corner error vs threshold | ≈ flat (~0.18 px, τ=1→8) |
| fallback path taken? | no (only on old builds) |
Trade off
Comparison matrix
From Milestone 3 — MAGSAC++: every row here is a choice with a cost. Fill the Expected on OpenCV 4.12 column, then say which row you would actually pick and what you give up for it.
| Self-check | Expected on OpenCV 4.12 |
|---|---|
| hasattr(cv2, 'USAC_MAGSAC') | True |
| corner error vs threshold | ≈ flat (~0.18 px, τ=1→8) |
| fallback path taken? | no (only on old builds) |
Worked example
Run all three on your before/after pair and print the funnel plus the ablation. This is the result you can defend.
# flags now wired through the pipeline
# --use-rootsift / --no-use-rootsift
# --cross-check-matches / --no-cross-check-matches
# MAGSAC is automatic (with fallback)
python stable_landmark_labeler_v2.py \
--image1 before.jpg --image2 after.jpg \
--use-rootsift --cross-check-matches --output-dir out/| Metric | Before fixes | After fixes |
|---|---|---|
| correct-match rate | ~13.8% | ~70.3% |
| inlier ratio (w) | ~0.13 | well above floor |
| homography corner error | τ-sensitive, up to ~1.0 px | ~0.18 px, τ-stable |
| mole-spot landmarks | silently dropped | recovered |
Three localized changes — one per failing stage — turned a contaminated, unstable fit into a clean, reproducible one. If your pair shows the same shape, you fixed it.
Comparison
Comparison matrix
From Put it together — the end-to-end win: refill the After fixes column from what you know. The rest of the table is as it appeared.
| Metric | Before fixes | After fixes |
|---|---|---|
| correct-match rate | ~13.8% | ~70.3% |
| inlier ratio (w) | ~0.13 | well above floor |
| homography corner error | τ-sensitive, up to ~1.0 px | ~0.18 px, τ-stable |
| mole-spot landmarks | silently dropped | recovered |
Constraint
Discussion prompt
Run The accuracy recipe, generalized with this step confiscated:
Robust estimator, soft threshold — MAGSAC++ over fixed-τ RANSAC; pass τ as an upper bound.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
This transfers to any feature-based registration, not just skin.
Edge cases
Discussion prompt
The accuracy recipe, generalized works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
This transfers to any feature-based registration, not just skin.
Elimination
Eliminate the wrong options
Your fixes raised geometrically-correct matches from ~14% to ~70% and gave a threshold-stable ~0.18 px homography. A reviewer asks which single change most improves the PURITY of the correspondence set fed to RANSAC. Best-supported answer?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Purity is a property of the correspondence set produced by the matcher. The ablation shows cross-check alone moves correct-match rate from 13.8% to 59% (with RootSIFT compounding it to ~70%). RootSIFT helps description and rescues mole landmarks, but the dominant purity lever is the mutual-NN constraint.
Check
Synthesis. Read carefully.
Check your understanding
Your fixes raised geometrically-correct matches from ~14% to ~70% and gave a threshold-stable ~0.18 px homography. A reviewer asks which single change most improves the PURITY of the correspondence set fed to RANSAC. Best-supported answer?
Answer: A
Why: Purity is a property of the correspondence set produced by the matcher. The ablation shows cross-check alone moves correct-match rate from 13.8% to 59% (with RootSIFT compounding it to ~70%). RootSIFT helps description and rescues mole landmarks, but the dominant purity lever is the mutual-NN constraint.
Concept
You can now stand at a whiteboard and defend the pipeline end to end. Be ready to:
Reproduce the ablation on a fresh pair and you've shown the fix generalizes — the real test of understanding.
Counterexample
Discussion prompt
You can now stand at a whiteboard and defend the pipeline end to end. Be ready to:
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Reproduce the ablation on a fresh pair and you've shown the fix generalizes — the real test of understanding.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Diagnose Before You Tune · The Descriptor: RootSIFT · The Matcher: Ratio + Mutual NN · The Estimator: RANSAC → MAGSAC++ · Your Turn: Implement All Three. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You took a working pipeline and made it accurate and reproducible by fixing the right stage three times.
| Failing stage | Fix | Evidence |
|---|---|---|
| Description drift | RootSIFT (all sources) | 0 → 10 mole matches; 90% correct |
| Match contamination | mutual cross-check | 13.8% → 59–70% correct |
| Estimation brittleness | MAGSAC++ | ~0.18 px, flat over τ=1–8 |
Want this taught 1-on-1? Alexander tutors Computer Vision — $55/session, free consultation.