A working-session deck for a hand-annotated PASTIS parcel instance-segmentation project, built around six findings measured from the student's own files: 83 instances per 128x128 tile, box sides of 4 to 38 px, and annotations byte-identical across all 61 observations of a tile. It covers round-trip QA of written artifacts, why per-observation duplication makes a random split leak an entire validation set, configuring Mask R-CNN anchors and min_size from measured object scale instead of COCO defaults, the dormant coordinate bug created by a duplicated 850 literal, polygons as source versus rasters as build output via COCO, fixed per-band normalisation, and a temporal-consistency filter that replaces a guessed pseudo-label confidence threshold with agreement across the time series.
Subject: Computer Vision · 84 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
Computer Vision · Instance Segmentation
You drew 83 parcels by hand and the polygons are good. The problem is one layer up — in what the pipeline writes to disk, and in what you are about to ask a model to learn from it.
Objectives
Your annotator works. It has undo, per-polygon commit, an escape hatch for cloudy dates, and it rasterises correctly. This deck is not about fixing it — it is about the layer above it, where the expensive mistakes live.
Warm-up
Answer from memory. You wrote this code, so this should be quick — and if it is not quick, that is itself the finding.
Discussion prompt
For one PASTIS tile, annotate_first_apply_to_all finishes and writes files to disk.
Without looking: how many files, and what is in each one? Then the harder half — how many of those files contain different information from one another?
Hint: Count the loop. save_annotations_for_all_observations runs once per observation, and the tile has around 61 of them.
Answer:
Three files per observation — _boxes.npy, _labels.npy, _masks.npy — times ~61 observations, so roughly 183 files per tile.
The second half is the one that matters. All 61 observations receive the same arrays. I hashed the files you sent: obs 000 through 004 are byte-identical. So 183 files carry the information content of three.
Concept
A model is the easy part. Datasets are where projects quietly fail, and they fail in a small number of repeatable ways. Every one of them is a question you can ask before you train.
| Question | What goes wrong when you skip it | Where we answer it |
|---|---|---|
| Are the labels correct? | You optimise toward noise and never find out | Part 1 |
| How many independent labels are there? | Validation measures memorisation | Part 2 |
| What do the objects look like, numerically? | Framework defaults silently mismatch your data | Part 3 |
| Can you regenerate everything from source? | One fix means redoing all the labelling | Part 5 |
Notice that none of these mention architecture. You picked Mask R-CNN, which is a reasonable choice — and it is very nearly the least important decision on this list.
Socratic
Discussion prompt
You have annotated one tile: 83 parcels, drawn by hand, saved to disk.
How do you know they are correct?
Not "I was careful" — that is a claim about your intent, not about the bytes on disk. What procedure would show you a mistake if one existed?
Hint: You looked at an image and clicked. Then a series of transformations happened to those clicks. Which of those transformations have you ever seen the output of?
Answer:
Right now there is no such procedure. Nothing in the pipeline reads a saved .npy back and puts it next to the image it describes. The annotations are write-only.
That is the single most important structural gap in the project, and everything in Part 1 follows from closing it.
Section
Part 1
Concept
Your pipeline is a one-way street. Clicks go in at 850×850, get scaled to 128×128, get rasterised, get saved. At no point does anything load a saved file and ask does this match the image it came from?
Round-trip check — Load the artifact you wrote, render it against the input it describes, and look. If the two disagree, you have found a bug that no unit test was ever going to catch.
This matters more in vision than almost anywhere else, because the failure mode is geometric. A coordinate scaling bug does not raise an exception. It does not produce a NaN. It produces perfectly well-formed arrays of plausible numbers that happen to describe the wrong parts of the image.
The model will then train happily on them, converge, and report a loss curve that looks exactly like a healthy one.
Anomaly
Here is what you sent me, as bytes rather than as filenames:
| File | Shape | dtype | MD5 (first 12) |
|---|---|---|---|
| S2_10000_obs_000_boxes.npy | (83, 4) | float32 | 2eb1625e4fde |
| S2_10000_obs_002_boxes.npy | (83, 4) | float32 | 2eb1625e4fde |
| S2_10000_obs_004_boxes.npy | (83, 4) | float32 | 2eb1625e4fde |
| S2_10000_obs_000_labels.npy | (83,) | int64 | 8fc8dd18697e |
| S2_10000_obs_003_labels.npy | (83,) | int64 | 8fc8dd18697e |
Predict first
Look at the last column before reading on.
Every hash in a given group is identical. Two questions:
1. Is this a bug?
2. Whether or not it is a bug, what does it do to the size of your training set?
Correct: It is not a bug — it is exactly what annotate_first_apply_to_all was written to do, and the reasoning behind it is sound. But it means 61 files per tile carry one tile's worth of information, so your dataset has ~61× fewer independent labels than it has samples.
Hold onto the second answer. Part 2 is entirely about it.
Why: The duplication is deliberate and correct as a modelling decision: parcel geometry does not move across a growing season, so re-drawing it per date would be wasted labour. The trap is that it is also a statement about your dataset's independence structure, and nothing downstream knows that. A random train/val split over observations will put the same geometry in both halves.
Worked example
This is annotation_qa.py run against your files. It loads every saved array, regroups them by tile and observation, and checks each one against the others and against itself.
5 annotation records across 1 tiles
--- duplication ---
1/1 tiles have byte-identical annotations across all their observations.
--- object scale (this drives your anchor config) ---
box side px, native 128: min 4 p25 12 median 16 p75 20 max 38
instances per tile: median 83, max 83
--- hard checks ---
FAIL S2_10000 obs 1: labels/masks present but no boxes
FAIL S2_10000 obs 3: labels/masks present but no boxes| Line | What it is | What you do about it |
|---|---|---|
| 1 tile, 5 records | independence structure | split by tile, not by record |
| byte-identical | expected duplication | store once, serve many |
| side 4–38 px | object scale | set anchors from this |
| 83 instances | density | raise box_detections_per_img |
| 2 FAILs | missing files | in this case an upload artifact |
The two FAIL lines are not a real defect — you simply did not send boxes for observations 1 and 3. But notice what happened: the check found the gap without being told what to look for. That is the whole argument for having one.
Pattern
Five steps. Run them as a script, not as a notebook cell you scroll past, and make the script exit non-zero when something fails so it can move into CI later.
Step 3 is the sharpest one and the cheapest to write. You already have two representations of the same fact — the polygon-derived box and the mask — so make them check each other.
def boxes_from_masks(masks):
"""Recompute boxes with the same +1 exclusive convention as
create_training_targets, so the two are directly comparable."""
out = np.zeros((len(masks), 4), dtype=np.float32)
for i, mask in enumerate(masks):
ys, xs = np.where(mask == 1)
out[i] = [xs.min(), ys.min(), xs.max() + 1, ys.max() + 1]
return out
drift = np.abs(boxes_from_masks(masks) - saved_boxes).max()
assert drift <= 0.5, f"coordinate bug: {drift:.1f}px drift"| Drift | Meaning |
|---|---|
| 0 px | masks and boxes agree — geometry is internally consistent |
| ± a few px | a rounding or convention mismatch between the two paths |
| large / erratic | a scaling bug — the two were computed in different coordinate spaces |
Concept
It is worth being precise about why these problems are invisible, because the reason generalises far past this project.
This is the defining property of data bugs: they are silent by construction. Code bugs announce themselves; data bugs just quietly shift what your model learns. The only defence is to build the observation channel yourself, deliberately, before you need it.
Check
Reason about the mechanism, not the vibe.
Check your understanding
You add a check that recomputes each mask's bounding box and compares it to the saved box. On your current data it reports 0 px drift everywhere. What have you actually proven?
Answer: A
Why: Both artifacts are computed from the same polygon inside create_training_targets, so agreement between them rules out a bug between those two steps and nothing more. If the polygon itself was recorded in the wrong coordinate space, the mask and the box will be wrong together and agree perfectly. That is exactly why the overlay against the image is a separate, non-negotiable check — it is the only one that tests the polygon itself.
points array a few lines apart, so their agreement is a weaker claim than it looks.Explain it to yourself
Discussion prompt
In your own words, and without using the phrase "good practice":
Why is a check that compares two things you derived from the same source still worth writing — and what is it specifically unable to tell you?
Hint: Think about which pairs of pipeline stages the check sits between, and which stages it sits entirely downstream of.
Answer:
A good answer names the boundary: the check covers everything from points onward, and covers nothing before it. Cheap, genuinely useful, and honest about its limit.
Being able to state what a test cannot see is most of what separates a useful test suite from a reassuring one.
Section
Part 2
Concept
Start by giving the design its due, because the reasoning behind it is right:
def save_annotations_for_all_observations(
file_path, number_of_images, boxes, labels, masks,
annotation_path, mask_path,
):
for observation_index in range(number_of_images):
annotation_name = f"{file_path.stem}_obs_{observation_index:03d}"
np.save(annotation_path / f"{annotation_name}_boxes.npy", boxes)
np.save(annotation_path / f"{annotation_name}_labels.npy", labels)
np.save(mask_path / f"{annotation_name}_masks.npy", masks)| Claim | True? | Why |
|---|---|---|
| Parcel geometry is static across a season | Yes | field boundaries do not move between satellite passes |
| So one annotation serves every date | Yes | re-drawing per date would be wasted labour |
| So every date is a training sample | Yes | different reflectance, same geometry — real augmentation |
| So every date is an independent label | No | this is the step that does not follow |
The first three are sound. The fourth is the one nothing in the code disagrees with, because nothing in the code has any opinion about independence at all.
Estimation
Estimate before you calculate. Being roughly right here is the skill; being precise is arithmetic.
Predict first
You finish your plan: 100 tiles, annotated by hand, each with about 61 observations.
1. How many training samples is that?
2. How many independent labels is that?
3. If you use a standard 80/20 random split over samples, roughly what fraction of your validation tiles have their geometry already memorised from training?
Correct: About 6,100 samples; exactly 100 independent labels; and essentially 100% of validation tiles will have appeared in training.
0.2^61 ≈ 10^-43. There is no sampling luck that rescues this — it is structural.
Why: With 61 observations per tile and a random 80/20 split over the ~6,100 samples, the chance that a given tile contributes zero of its 61 observations to the training set is 0.2^61 — indistinguishable from zero. So effectively every tile in your validation set has had its exact parcel geometry seen during training. Your val mAP would then be measuring how well the model recalls 100 specific field layouts, not how well it segments fields.
Concept
This distinction has a name in statistics and a different name in every applied field, which is part of why it keeps being rediscovered the hard way.
Grouped data — Observations that share a latent source. Within a group, samples are correlated; across groups they are not. The group — not the sample — is the unit of independence.
Figure (svg): Two split strategies. In a random split over observations, coloured squares representing tile A and tile B appear in both the train and validation boxes. In a tile-level split, every square from tile A is in train and every square from tile B is in validation.
The same structure, under other names: multiple photographs of one patient; several sentences from one document; repeated measurements on one subject; several frames from one video. In every case, splitting at the sample level and reporting the resulting score is the same mistake.
Worked example
This is what almost everyone writes first. It uses a fixed seed, it shuffles properly, it even stratifies. It is still wrong.
# WRONG for grouped data, however careful it looks.
records = sorted(annotation_dir.glob('*_boxes.npy')) # ~6100 entries
rng = np.random.default_rng(0)
rng.shuffle(records)
cut = int(0.8 * len(records))
train, val = records[:cut], records[cut:]
# S2_10000_obs_007 -> train
# S2_10000_obs_031 -> val <-- same 83 parcels, same geometry| Property of this code | Reads as | Actually |
|---|---|---|
| fixed seed | reproducible | reproducibly wrong |
| shuffled | unbiased | unbiased within the wrong unit |
| 80/20 | standard | standard for independent samples |
| one file = one sample | obvious | the assumption that fails |
Every line of it is defensible in isolation. The error is in a premise that is never written down anywhere: that one file is one independent thing.
Trap
You train, and validation mAP comes back at 0.87. You are pleased. You tune a little, get 0.89, and write it in the README.
Accept the number because it is high
Why: A high validation score feels like evidence that everything upstream is fine, so it quietly certifies the split, the labels, and the config all at once — none of which it actually tested.
Then the model meets a region it has not seen and collapses to 0.3, and you have no idea which of the twenty things you changed is responsible — because your baseline was never real.
Split by tile first, before the first training run, and write the split to a file you commit.
Expect the honest number to be lower
Why: A tile-level split on 100 tiles might report 0.55 instead of 0.87. That is not a regression — it is the first number you have that means anything, and every later improvement is measured against it.
A believable 0.55 is worth more than a flattering 0.89, because you can improve the first one on purpose.
Concept
Two separate ideas, and both matter.
The second point is the one people skip. If your split is a function of the directory listing, then annotating tile 101 reshuffles everything and last week's number is no longer comparable to this week's. Write it down once and read it thereafter.
Worked example
Short, boring, and load-bearing. Note that it operates on tile ids, never on file paths.
import json, numpy as np
from pathlib import Path
def build_splits(tile_ids, out_path, val_frac=0.2, seed=0):
"""Freeze a tile-level train/val split. Tiles, never observations."""
tiles = sorted(set(tile_ids)) # sorted => seed alone decides
rng = np.random.default_rng(seed)
order = rng.permutation(len(tiles))
n_val = max(1, round(val_frac * len(tiles)))
val = sorted(tiles[i] for i in order[:n_val])
train = sorted(tiles[i] for i in order[n_val:])
assert not set(train) & set(val), 'tile in both splits'
Path(out_path).write_text(json.dumps(
{'seed': seed, 'val_frac': val_frac, 'train': train, 'val': val},
indent=2))
return train, val| Line | Why it is there |
|---|---|
sorted(set(...)) | makes the result depend on the seed alone, not on filesystem order |
assert not set(train) & set(val) | the leak this whole part is about, checked rather than assumed |
writes seed into the file | the split can be regenerated and audited later |
| returns tile ids | the Dataset expands a tile into its observations, so leakage cannot re-enter |
Fill the middle
The split gives you tiles. The Dataset has to turn tiles into samples without letting the leak back in.
Fill in the blanks
class ParcelDataset(torch.utils.data.Dataset):
def __init__(self, tile_ids, ann_dir, img_dir):
# one entry per (tile, observation) — samples, not labels
self.index = [(t, o) for t in tile_ids
for o in range(n_observations(t))]
def __getitem__(self, i):
tile, obs = self.index[i]
image = load_observation(tile, obs) # varies with obs
target = load_annotation(tile) # constant per tile
return image, target
Why: Blank A takes only the tile ids handed to this split, so a tile can never appear in two datasets. Blank B keys the annotation on the tile alone — not on (tile, obs) — which is the line that encodes the real structure of the data: the image varies with the observation, the label does not. Once the loader is written this way, the 61 duplicate files on disk have no reason to exist.
Cost model
Annotate
Work through the storage arithmetic. The numbers are from your own tile.
uint8 is already the right choice — a boolean mask in bool dtype is also 1 byte per pixel in NumPy, so nothing is wasted here.Storage is cheap and this is still worth fixing — not because 8 GB is expensive, but because the ratio is a measurement of a design error. When a redundancy factor in your storage exactly matches a redundancy factor in your labels, the storage layout is telling you what the data actually is.
Comparison
Comparison matrix
Fill the After column from what you now know. The shape of the fix should be predictable by this point.
| Property | Before | After |
|---|---|---|
| annotation files per tile | ~183 | 3 |
| mask bytes, 100 tiles | ~8.3 GB | ~136 MB |
| unit of the train/val split | observation (file) | tile |
| independent val labels | 0 — all seen in train | 20 unseen tiles |
| training samples | ~6,100 | ~6,100 — unchanged |
The last row is the one worth pausing on. You do not lose training data by fixing this — every observation is still a sample, and the reflectance really does vary across dates. What changes is only the bookkeeping: which of those samples are allowed to be used to judge the model.
Check
One of these is a real defence. The others are things that feel like defences.
Check your understanding
You want to be confident there is no geometry leakage between your train and validation sets. Which check actually establishes that?
Answer: A
Why: Leakage here is defined at the tile level, because the tile is what determines the label. Disjoint tile ids is therefore the exact statement of the property you want, and it is one line to assert. Everything else either checks a weaker property or checks nothing at all.
obs_007 and obs_031 are different filenames carrying identical labels.Real world
Discussion prompt
The pattern is: many samples share one latent source, and the label is a property of the source rather than of the sample.
Name three settings outside satellite imagery where this holds. For each, say what the group is — then say what a careless engineer would split on instead.
Hint: Anywhere a single real-world entity gets measured, photographed, or recorded more than once.
Answer:
| Setting | Group (correct unit) | What gets split on by mistake |
|---|---|---|
| Medical imaging | patient | image / slice |
| Speech recognition | speaker | utterance |
| Document NLP | document | sentence |
| Video understanding | clip | frame |
| Recommenders | user | interaction |
| Industrial inspection | physical part | photograph |
sklearn.model_selection names this directly: GroupKFold, GroupShuffleSplit, StratifiedGroupKFold. When a library gives a concept its own family of classes, that is a strong signal about how often it is needed.
Intuition
If you remember nothing else from this part:
The rule — Split on whatever generated the label. If two samples share a label because they share a source, they belong on the same side of the split — always.
It generalises because it is not really about datasets. It is about not letting the unit you happen to store masquerade as the unit you actually sampled.
Section
Part 3
Concept
Everything in this part follows from five numbers I measured from your S2_10000_obs_000_boxes.npy. None of them is a problem. All of them are configuration inputs you have not used yet.
| Measurement | Value | Why the model cares |
|---|---|---|
| instances per tile | 83 | detection cap, RPN sampling density |
| box side, median | 16 px | anchor sizes, FPN level assignment |
| box side, min | 4 px | the smallest thing you are asking it to find |
| box side, max | 38 px | the scale range it must span — narrow |
| Σ box area ÷ image area | 1.36× | boxes overlap; parcels tile the plane |
That last row is worth a second look. The bounding boxes cover 136% of the image, which is only possible because they overlap heavily — exactly what you expect from agricultural parcels, which tessellate rather than sit isolated on a background. Object detectors are mostly designed and benchmarked on the opposite assumption.
Prediction
torchvision.models.detection.maskrcnn_resnet50_fpn wraps your input in a GeneralizedRCNNTransform before the backbone ever sees it. Its defaults are min_size=800, max_size=1333.
Predict first
You pass a 128×128 PASTIS tile into that model, unmodified.
What does the transform do to it, and what does that do to your median 16-pixel parcel?
Correct: It upscales the tile 6.25× to 800×800 by bilinear interpolation, which turns the median 16 px box into a 100 px box — and, by accident, lands your objects squarely inside the default anchor range of 32–512.
The lesson is not "the defaults are wrong". It is that you did not know what they were doing, and a 6.25× resize is not something you want to discover by accident.
Why: This is the detail that makes the situation confusing rather than simply broken: the default resize accidentally rescues the default anchors. Your model will train, and the numbers will not be absurd. But you are now paying full 800×800 COCO-scale compute for a 128×128 tile, and the network is spending capacity on interpolated detail that carries no sensor information. Neither of those shows up as an error — only as a slow training run and a ceiling you cannot explain.
Invariant
Step through it
Track your median parcel — a 16-pixel box — from disk to loss. Watch which stage changes its size and which stage interprets that size.
Frame 3 is where the surprise is, and frame 5 is where the ceiling is. A 28×28 mask prediction over a genuinely 16-pixel object is mostly interpolation — so expect mask AP to trail box AP, and do not spend a week hunting for the bug that causes it. That gap is a property of the data.
Concept
An anchor is a hypothesis about size. The RPN lays a grid of boxes of fixed sizes over each FPN level and asks, for each one, is there an object roughly this big here?
Anchor–object scale match — The assumption that your objects' size distribution overlaps the anchor size distribution. When it does not, the RPN proposes boxes that never sufficiently overlap a ground-truth box, those anchors are labelled negative, and the detector learns to predict nothing.
Figure (svg): Two number lines. The top shows the student's object sizes clustered between 4 and 38 pixels at the far left. The bottom shows default anchor sizes at 32, 64, 128, 256 and 512, spread across the full range. A caption notes that the 6.25x resize is what moves the object band rightward to overlap the anchors.
Two ways to make those bands overlap: resize the image until the objects reach the anchors, or move the anchors down to the objects. The defaults do the first, invisibly. Doing the second, deliberately, is cheaper and says what you mean.
Trap
You read that small-object detection needs smaller anchors, so you set sizes=((8,), (16,), (32,), (64,), (128,)) and leave min_size=800 alone.
Change one thing without checking the other
Why: Anchors are compared against boxes after the transform has already scaled them 6.25x. Your 100 px transformed boxes now sit above almost every anchor you just configured, so you have moved the bands apart rather than together.
Recall drops, you conclude that small anchors do not help, and you revert a change that was pointing in the right direction.
Fix min_size first, then measure the box sizes the RPN will actually see, then set anchors from that measurement.
Print the box sizes after the transform, not before
Why: One print inside a GeneralizedRCNNTransform call tells you the real numbers. Configuration derived from a measurement survives a change of dataset; configuration copied from a blog post does not.
min_size=256 with anchors of roughly 8–128 is a defensible starting point for these tiles — and you will be able to say why, which is the part that matters.
Worked example
Nothing exotic. The point is that every number here traces back to something printed from your own data.
from torchvision.models.detection import maskrcnn_resnet50_fpn
from torchvision.models.detection.rpn import AnchorGenerator
model = maskrcnn_resnet50_fpn(weights='DEFAULT')
# 1. Stop the 6.25x upscale. 128 -> 256 is a deliberate 2x.
model.transform.min_size = (256,)
model.transform.max_size = 256
# 2. Anchors spanning 2x the measured range (4-38 px -> 8-76 px).
model.rpn.anchor_generator = AnchorGenerator(
sizes=((8,), (16,), (32,), (64,), (128,)), # one per FPN level
aspect_ratios=((0.5, 1.0, 2.0),) * 5, # parcels are irregular
)
# 3. You measured 83 instances per tile. The default cap is 100.
model.roi_heads.detections_per_img = 300| Setting | Default | Yours | Measured from |
|---|---|---|---|
| min_size | 800 | 256 | 128 px tiles — a 2× upscale, chosen not inherited |
| anchor sizes | 32–512 | 8–128 | box sides 4–38 px, doubled by the resize |
| detections_per_img | 100 | 300 | 83 instances per tile, with headroom |
| aspect_ratios | 0.5/1/2 | unchanged | parcels are irregular; no reason to narrow it |
Treat these as a starting hypothesis, not an answer. The habit being taught is that each one has a reason you can state out loud, so when a number needs to change you know which measurement to go back to.
Tweak it
Parameter explorer
Three knobs, all interacting. Move one at a time and predict the effect on recall of small parcels, on GPU memory, and on step time — before you reason it through.
The interaction that catches people: min_size and anchor size are not independent knobs. Doubling the resize halves the effective anchor size relative to your objects. Change one, and you have implicitly changed the other.
Trade off
Comparison matrix
Fill in the blanks. There is no single right answer here — there is a right way to decide.
| min_size | Effective object size | Compute vs 128 | When this is the right call |
|---|---|---|---|
| 128 (none) | 4–38 px | 1× | fastest iteration; risks the 4 px tail being unlearnable |
| 256 | 8–76 px | 4× | sane default: small objects get real pixels, cost stays moderate |
| 512 | 16–152 px | 16× | if the smallest parcels are demonstrably being missed |
| 800 (default) | 25–238 px | 39× | almost never here — inherited, not chosen |
Compute scales with the square of the resize, which is why 800 costs 39× rather than 6.25×. That is the number to have in mind when a training run feels inexplicably slow.
Discrimination
Sort into buckets
Each symptom points at one stage. Sort them by which stage you would actually change — this is the diagnosis-first habit applied to configuration.
detections_per_img=100 against 83+ ground-truth instances is a ceiling you will hit on dense tiles.Check
One symptom, four explanations, one mechanism.
Check your understanding
You keep min_size=800 but set anchor sizes to ((4,), (8,), (16,), (32,), (64,)) to match your measured 4–38 px boxes. Recall gets dramatically worse. Why?
Answer: A
Why: Anchor sizes are expressed in the coordinate space the RPN operates in — which is after GeneralizedRCNNTransform has resized both image and boxes. Setting anchors from your native 128 px measurements while leaving an 800 px resize in place puts the two distributions on opposite sides of each other. The fix is to make min_size and the anchor sizes agree, in that order.
Notation
Annotate
Every argument below is a claim about your data. Read each one as a sentence and decide whether it is true of PASTIS parcels.
num_classes=2 says: there is exactly one kind of foreground thing. Index 0 is always background in torchvision, which is why your labels = np.ones(...) is correct and not off by one.min_size=max_size=256 says: every tile is square and the same size. True for PASTIS, false for most datasets — worth a comment so the next reader knows it is deliberate.sizes says: there are five FPN levels. The tuple length is not free — it must match the number of feature maps the backbone emits.box_detections_per_img=300 says: I expect up to 300 parcels in a tile. You measured 83. Headroom is cheap; truncation is silent.box_score_thresh=0.05 says: keep low-confidence detections. Correct for computing mAP, wrong for a demo — mAP integrates over thresholds and wants the full ranking.rpn_batch_size_per_image=256 says: sample 256 anchors per image for the RPN loss. With 83 dense ground-truth objects your positive/negative balance differs from COCO's — worth checking before you tune anything else.Section
Part 4
Concept
This one is currently harmless, which is exactly what makes it worth an entire section. Here is the caller:
annotation_size = 850
display_rgb = cv2.resize(
rgb, (annotation_size, annotation_size),
interpolation=cv2.INTER_LANCZOS4)
polygons = create_mask.draw_polygons(
rgb=display_rgb, file_name=file_path.name,
observation_index=observation_index)
boxes, labels, masks = create_mask.create_training_targets(
polygons=polygons,
display_size=(annotation_size, annotation_size), # <- 850
native_size=(width, height))| Value | Where it comes from | Where it is used |
|---|---|---|
annotation_size = 850 | the caller | resizing the image the annotator sees |
display_size=(850, 850) | the same variable | scaling clicks back down to 128 |
850 (hardcoded) | inside draw_polygons | resizing the canvas that receives the clicks |
Three places, two sources. Two of them move together when you edit the variable. The third does not.
Error analysis
Annotate
This is the drawing loop inside draw_polygons. The image arriving in rgb has already been resized to 850×850 by the caller. Find the line that will break the moment annotation_size changes.
850, with no reference to the annotation_size the caller used.annotation_size is also 850, so this resize is a no-op and everything works. The bug is dormant, not absent.annotation_size = 900 and the caller hands in a 900×900 image, which this line squashes back to 850. Clicks are then recorded in 850-space.create_training_targets is told display_size=(900, 900), so it scales those 850-space clicks by 128/900 instead of 128/850.while True loop, which is wasted work even in the version that is correct.Worked example
It is tempting to think a small systematic offset is a small problem. Work it through against your measured object sizes.
annotation_size = 900 # you changed this
clicks_recorded_in = 850 # draw_polygons still resizes to 850
assumed_scale = 128 / 900 # what create_training_targets uses
actual_scale = 128 / 850 # what the clicks actually needed
error = 1 - (assumed_scale / actual_scale) # 0.0556
median_box = 16 # px, measured from your data
print(median_box * error) # 0.89 px on a 16 px object
print(38 * error) # 2.1 px on the largest| Object | Native size | Shift | As a fraction of the object |
|---|---|---|---|
| smallest parcel | 4 px | 0.22 px | 5.6% |
| median parcel | 16 px | 0.89 px | 5.6% |
| largest parcel | 38 px | 2.1 px | 5.6% |
| tile corner | 128 px | 7.1 px | — |
A 5.6% systematic shift toward the origin, on every parcel, in every tile annotated after the change. Mixed in with correctly-annotated tiles it becomes label noise with a direction — which is worse than random noise, because the model can learn it.
Concept
Coincidental correctness — Code that produces the right answer because two independent values happen to be equal, rather than because it computes the relationship between them.
There is a second instance of exactly this in the same file:
height = img_array.shape[2]
width = img_array.shape[3]
create_training_targets(
polygons=selected_polygons,
display_size=(annotation_size, annotation_size),
native_size=(width, height))
# and inside:
# native_width, native_height = native_size
# points[:, 0] *= native_width / display_width
# points[:, 1] *= native_height / display_height| Question | Answer |
|---|---|
| Is the (width, height) order correct here? | Yes — it matches the unpacking inside |
| Would a transposed order be caught? | No — PASTIS tiles are 128×128 |
| What makes it correct? | H == W, not the code |
| When does it break? | The first non-square input, silently |
Both bugs share a signature worth learning to recognise: a value used in two places, where only one of them derives it from the other. When you see a literal that duplicates a variable, that is the shape.
Trap
The overlay looks right. The masks land on parcels. You conclude the annotation code is correct and move on to the model.
Treat passing output as proof of a correct mechanism
Why: Your test exercised exactly one configuration — the one where the two 850s coincide. A test that cannot distinguish 'correct' from 'coincidentally correct' has not tested the thing you think it has.
Three weeks later you bump the annotation window size for a larger monitor, and quietly corrupt every tile you annotate after that.
Remove the duplicated value so the two cannot disagree, then re-run the overlay to confirm nothing moved.
Make the relationship explicit, then vary the input
Why: Pass annotation_size into draw_polygons and derive the canvas from it. Then annotate one tile at 900 and run the overlay: if the masks still land on parcels, the relationship is real and not a coincidence.
One parameter, one source of truth, and a check that varies it. That is the whole fix.
Faded example
Three blanks. The goal is a function where the display size can be changed from one place without silently corrupting anything.
Fill in the blanks
def draw_polygons(rgb, file_name, observation_index,
display_size): # <- passed in, not assumed
display_image = np.clip(rgb * 255, 0, 255).astype(np.uint8)
display_image = cv2.cvtColor(display_image, cv2.COLOR_RGB2BGR)
# Assert the contract instead of hoping for it.
if display_image.shape[0] != display_size:
raise ValueError('image does not match the annotator display size')
while True:
canvas = display_image.copy() # no resize: it already fits
# ... draw polygons, read clicks off canvas ...
# caller — one source of truth for the number:
polygons = draw_polygons(rgb=display_rgb, file_name=name,
observation_index=i,
display_size=annotation_size)
Why: Blank A turns the silent assumption into a loud failure: the function now asserts the contract it depends on instead of hoping for it. Blank B deletes the second resize entirely — the image already arrives at the right size, so copying is all the frame loop needs, and the per-frame resize disappears as a bonus. Blank C passes the single source of truth down from the caller, so annotation_size now controls every place the number is used. The variable and the literal can no longer disagree because the literal is gone.
Two truths and a lie
Sort into buckets
Some of these are true about the annotation code as it stands today. Others are the kind of thing that sounds true and would get you into trouble. Sort them.
native_size=(width, height) argument is passed in the correct order.; The per-frame cv2.resize inside the draw loop is wasted work.annotation_size alone is safe, since it is used consistently.; A mask's bounding box matching its saved box proves the polygon was recorded correctly.annotation_size is precisely the variable that is not used consistently.Check
The general skill, not the specific bug.
Check your understanding
Which of these is the strongest evidence that a piece of geometric code is correct rather than coincidentally correct?
Answer: A
Why: Coincidental correctness means two values happen to be equal. The only test that distinguishes it from real correctness is one where they are not equal — so you vary the parameter and check the output again. Everything else holds the coincidence fixed and therefore cannot see it.
Section
Part 5
Concept
create_training_targets takes your polygons, rasterises them, and returns arrays. The polygons themselves are never saved. They exist only inside the function call that consumes them.
for polygon in polygons:
points = np.asarray(polygon, dtype=np.float32)
points[:, 0] *= native_width / display_width
points[:, 1] *= native_height / display_height
points = np.rint(points).astype(np.int32)
mask = np.zeros((native_height, native_width), dtype=np.uint8)
cv2.fillPoly(mask, [points], 1) # <- polygon becomes pixels here
instance_masks.append(mask)
# `polygons` goes out of scope. The clicks are gone.| Artifact | Cost to produce | Saved? | Recoverable? |
|---|---|---|---|
| polygon vertices | hours of human attention | no | only by redrawing |
| binary masks | microseconds of fillPoly | yes | trivially, from polygons |
| bounding boxes | microseconds of np.where | yes | trivially, from masks |
Read that table twice. You are persisting the two artifacts that cost nothing to recreate, and discarding the one that cost you an afternoon.
Counterexample
Discussion prompt
Here is a claim, stated as though it always holds:
"Masks are all I need. The polygons were just an input method."
Do one of two things: produce a concrete situation where that claim costs you real work, or say precisely what rules such a situation out. "It's just good practice" is not on the menu.
Hint: Think about every operation you might want to perform on an annotation that is not "feed it to this exact model at this exact resolution".
Answer:
.npy masks.pycocotools wants COCO polygons or RLE. Without them you write and debug your own mAP.Six of them, and every one costs hours you already spent once.
Concept
Source artifact — The thing a human produced or a decision created. Irreplaceable without redoing the work. Version-controlled, backed up, treated as precious.
Derived artifact — Anything a script can regenerate from source. Disposable by design, and safe to delete at any time — because deleting it is a re-run, not a loss.
You already apply this instinct elsewhere without thinking about it. You commit .py files and not __pycache__. You commit requirements.txt and not site-packages. Annotations are the same category of thing, and they are far more expensive than either.
| In your project | Source | Derived |
|---|---|---|
| annotation | polygon vertices + class | masks, boxes, areas, RLE |
| dataset composition | splits.json | the DataLoader's index |
| configuration | the config file | the instantiated model |
| results | the run log | plots and tables in the README |
Worked example
The format is unremarkable, which is the point — it is the one every tool already speaks. Note that geometry is stored per tile, not per observation, which folds in the Part 2 fix.
import json
from datetime import datetime, timezone
def polygons_to_coco(tile_polygons, tile_size=128, out_path='annotations.json'):
"""tile_polygons: {tile_id: [[(x, y), ...], ...]} in NATIVE coords."""
images, annotations, ann_id = [], [], 1
for image_id, (tile_id, polys) in enumerate(sorted(tile_polygons.items()), 1):
images.append({'id': image_id, 'file_name': f'{tile_id}.npy',
'width': tile_size, 'height': tile_size})
for poly in polys:
xs = [float(x) for x, _ in poly]
ys = [float(y) for _, y in poly]
annotations.append({
'id': ann_id, 'image_id': image_id, 'category_id': 1,
'segmentation': [[c for xy in poly for c in map(float, xy)]],
'bbox': [min(xs), min(ys), max(xs) - min(xs), max(ys) - min(ys)],
'area': float(polygon_area(poly)), 'iscrowd': 0})
ann_id += 1
json.dump({'info': {'date_created': datetime.now(timezone.utc).isoformat()},
'images': images, 'annotations': annotations,
'categories': [{'id': 1, 'name': 'parcel'}]},
open(out_path, 'w'), indent=2)| COCO field | Convention | Watch out for |
|---|---|---|
segmentation | flat [x1,y1,x2,y2,...] | flat, not pairs — a common first-time error |
bbox | [x, y, width, height] | NOT [x1,y1,x2,y2] like torchvision wants |
category_id | starts at 1 | 0 is reserved; matches your labels = 1 |
area | polygon area | the true area, not the bbox area — mAP splits on it |
iscrowd | 0 for normal instances | 1 means RLE region and changes how mAP scores it |
The bbox row is the one that bites everyone exactly once. COCO stores width and height; torchvision's boxes tensor stores x2, y2. Converting in the wrong direction produces boxes that are plausible, wrong, and completely silent — the same failure signature as Part 4.
Concept
This is not a hypothetical future benefit. Three things become available the day you write the file.
pycocotools gives you COCO-standard AP, AP50, AP75, and the small/medium/large breakdown — that last one matters enormously given your 4-38 px objects.The AP-by-size breakdown deserves emphasis. With a median object of 16 px, essentially every parcel you own falls into COCO's small bucket — so the aggregate AP you would otherwise report is dominated by a category most detection papers treat as the hard tail.
Ranking
Put in order
Put the rebuilt pipeline in the order it should run. One of these is the step that used to be missing entirely.
splits.json.Why: Human work first, then persist it in a lossless standard format, then verify it — the QA step goes immediately after saving, because that is the earliest point at which a saved artifact exists to check. Only then do you fix the split, because a split over unverified data is a split you will have to redo. Rasterisation drops to load time, where it belongs, since it is derived. Training comes last and reads from artifacts every one of which has been checked.
Matching
Match the pairs
Match each artifact to its role in the rebuilt pipeline. If you can do this without hesitating, the distinction has landed.
annotations.json (COCO)splits.jsonDECISIONS.mdannotation_qa.py outputWhy: The test for 'source' is not importance — it is whether a script could regenerate it. Masks and QA reports are reproducible from other artifacts, so they are derived no matter how useful they are. The split is interesting precisely because it is cheap to compute and still source: it encodes a decision about which tiles are allowed to judge the model, and recomputing it would silently invalidate every comparison made before.
Check
One line of reasoning separates the right answer from three plausible ones.
Check your understanding
You store polygons in COCO and rasterise inside the Dataset. A colleague argues this is wasteful: rasterising 83 polygons every time a tile is loaded burns CPU that a cached .npy would save. What is the strongest response?
Answer: A
Why: The source/derived distinction is about which artifact is authoritative, not about whether caching is allowed. Caching a derived artifact is fine and often correct; the discipline is that the cache must be reconstructible and safe to delete. That keeps the performance option open without giving up the ability to fix one parcel or change resolution later.
Section
Part 6
Concept
This appears three times in your file, once per display path. For display it is exactly right.
rgb = img_array[index, [2, 1, 0], :, :].astype(np.float32)
channel_min = rgb.min(axis=(1, 2), keepdims=True)
channel_max = rgb.max(axis=(1, 2), keepdims=True)
rgb = (rgb - channel_min) / (channel_max - channel_min + 1e-8)
rgb = np.clip(rgb, 0, 1)| Purpose | Per-image min–max | Verdict |
|---|---|---|
| see a cloudy tile clearly | stretches whatever is there to full range | correct |
| compare two tiles' brightness | destroys the comparison | wrong |
| train a model | every image gets its own scale | wrong |
| run inference reproducibly | output depends on the image's own extremes | wrong |
It is going to be copied into the Dataset, because it is written, it is nearby, and it produces nice-looking arrays in [0, 1]. That is how this particular mistake almost always happens.
Anomaly
Suppose you train with per-image min–max normalisation and the model works well on your validation tiles.
Predict first
Then you deploy it on a single new tile.
The same parcel, in the same field, on the same date, produces a different prediction depending on what else is in the tile — a bright cloud in one corner, say, or a dark water body.
Why? And why did validation never show this?
Correct: Per-image min-max makes every pixel's value depend on the minimum and maximum of the whole image, so an unrelated bright cloud rescales every other pixel. Validation missed it because your val tiles come from the same distribution and mostly contain the same kind of content.
Fixed per-band statistics remove the coupling entirely: pixel value in, same number out, every time, regardless of what shares the frame.
Why: Normalisation is part of the model's input contract. Per-image statistics make that contract depend on the image itself, which means the same physical surface maps to different network inputs depending on its neighbours. This is not merely a robustness issue: it makes single-tile inference non-reproducible in a way that is very hard to debug from the outside, because nothing about the parcel changed.
Worked example
PASTIS ships NORM_S2_patch.json for exactly this. If you were not using PASTIS you would compute the same thing yourself — over the training split only, which is a Part 2 idea reappearing here.
import json, numpy as np
def compute_norm_stats(train_tile_ids, img_dir, n_bands=10):
"""Per-band mean/std over the TRAINING tiles only.
Using val tiles here leaks their statistics into training."""
total = np.zeros(n_bands, np.float64)
total_sq = np.zeros(n_bands, np.float64)
count = 0
for tile in train_tile_ids:
arr = np.load(img_dir / f'{tile}.npy').astype(np.float64)
flat = arr.transpose(1, 0, 2, 3).reshape(n_bands, -1)
total += flat.sum(axis=1)
total_sq += (flat ** 2).sum(axis=1)
count += flat.shape[1]
mean = total / count
std = np.sqrt(total_sq / count - mean ** 2)
json.dump({'mean': mean.tolist(), 'std': std.tolist()},
open('norm_stats.json', 'w'), indent=2)
return mean, std
# in the Dataset, identical at train and inference:
# image = (image - mean[:, None, None]) / std[:, None, None]| Decision | Reason |
|---|---|
| training tiles only | val statistics are val information; using them is leakage |
float64 accumulators | summing squares over ~10^8 pixels overflows float32 precision |
| saved to JSON | a source artifact — inference must use the identical numbers |
| per band, not per image | the transform no longer depends on the image's own content |
| all bands, not just RGB | you display 3; the model should see all 10 |
Intuition
The rule is short: any transform that depends on the data must be fitted on training data only, saved, and reapplied unchanged at inference.
Normalisation statistics are the obvious case. The same rule covers class weights, anchor sizes derived from box statistics, and any threshold you tune. If a number came from looking at data, it belongs to the split it came from.
Notice that this is the Part 2 idea again, wearing different clothes. Leakage is not only about which samples go where — it is about which information crosses the boundary.
Section
Part 7
Concept
Worth saying plainly, because the adaptation is not trivial and the difference is where all the difficulty lives.
| Classification (Noisy Student) | Instance segmentation (yours) | |
|---|---|---|
| teacher output | one label + one confidence | N boxes, N masks, N scores, N unknown |
| "confident" | a single softmax value | per-instance, and badly calibrated |
| a wrong pseudo-label | one wrong class | a phantom parcel, or a missed one |
| missing output | impossible — always one label | a missed parcel becomes a false negative, taught as truth |
That last row is the one that makes detection self-training genuinely harder. In classification every sample gets a label. In detection, a parcel the teacher fails to find is not merely absent — it is actively taught to the student as background.
The literature to know by name: STAC, Unbiased Teacher, Soft Teacher, and — for instance segmentation specifically — Polite Teacher. You do not need to implement them. You need to be able to say what problem each was solving when someone asks.
Missing information
Discussion prompt
You asked how to set the confidence threshold for accepting pseudo-labels.
Before answering: what information would you need in order to choose one well? List everything. Then mark which items you can actually obtain.
Hint: A threshold trades false positives against false negatives. What would you need to know to price that trade?
Answer:
| You would need | Can you get it? |
|---|---|
| precision/recall vs score on unlabelled data | No — that needs the labels you do not have |
| whether the score is calibrated | Rarely; detector scores usually are not |
| the cost of a phantom vs a missed parcel | Yes — but it is a judgement, not a measurement |
| how the threshold interacts with round number | Only by running the loop and watching |
Which is the honest answer to your original question: you cannot choose a good global score threshold from information you have. Everyone picks 0.8 or 0.9 because everyone else does.
So the interesting move is to stop needing one.
Concept
The duplication from Part 2 — the thing that was a liability for splitting — is an asset here, and it is specific to your problem in a way that is worth exploiting.
Temporal support — The number of observations, out of the tile's time series, in which a matched detection appears. A property of the world rather than of the model's calibration — which is precisely what a raw confidence score is not.
Two things follow, and they are the reason this is worth building. Your "noise" for the student becomes genuine sensor and phenology variation rather than synthetic colour jitter. And your acceptance criterion becomes an agreement count rather than a guessed number.
Worked example
Run the teacher across all dates, match detections spatially, and keep what recurs.
from collections import defaultdict
def temporal_pseudo_labels(model, tile, n_obs, min_support=0.5, iou_thr=0.6):
"""Keep instances the teacher re-finds across the time series."""
clusters = [] # each: {'boxes': [...], 'obs': set()}
for obs in range(n_obs):
pred = model(load_observation(tile, obs))[0]
for box, score in zip(pred['boxes'], pred['scores']):
if score < 0.3: # a floor, not a decision
continue
hit = next((c for c in clusters
if iou(box, mean_box(c)) > iou_thr), None)
if hit is None:
clusters.append({'boxes': [box], 'obs': {obs}})
else:
hit['boxes'].append(box)
hit['obs'].add(obs)
# support = fraction of observations in which this instance was found
keep = [c for c in clusters if len(c['obs']) / n_obs >= min_support]
return [mean_box(c) for c in keep] # averaging also denoises| Detection | Found in | Support | Verdict |
|---|---|---|---|
| cluster A | 54 / 61 obs | 0.89 | keep — stable across the season |
| cluster B | 38 / 61 obs | 0.62 | keep — likely obscured on cloudy dates |
| cluster C | 9 / 61 obs | 0.15 | drop — appears and vanishes |
| cluster D | 2 / 61 obs | 0.03 | drop — almost certainly a hallucination |
The score < 0.3 floor is deliberately loose. It exists only to keep the clustering tractable — the real decision is made by support, which is the whole point. And averaging the boxes within a cluster gives you a better box than any single observation produced, because independent errors partially cancel.
Scale up
Step through it
Watch what happens to a real parcel and to a hallucination as you add more observations. At what point do they become distinguishable?
This is the argument in one picture. You cannot separate the two by score at any number of observations. You can separate them by support at three. Threshold-insensitivity is the goal — when the answer stops depending on where you drew the line, you have found the right quantity to measure.
Trade off
Comparison matrix
Fill in the consequences. Notice how much less sensitive this choice is than a score cutoff — that insensitivity is the feature.
| min_support | Keeps | Risk | When to use it |
|---|---|---|---|
| 0.25 | anything seen in 15+ of 61 | admits noise from persistent hallucinations | very cloudy series where real parcels are often hidden |
| 0.50 | seen in ~half the series | balanced | the sensible default to start from |
| 0.75 | seen in 46+ of 61 | drops parcels obscured by cloud | round 1, when purity matters more than volume |
| 0.95 | near-universal detections only | very few pseudo-labels survive | when you have evidence the loop is drifting |
A refinement worth mentioning once you have this working: cloud cover is not random across the series, so weighting each observation by a cloud score before computing support would be more principled than treating all 61 dates as equal evidence. That is a good second iteration, not a first one.
Trap
You have 100 manually labelled tiles. Round 1 works, so you pseudo-label the unlabelled pool and add the confident ones to training. Round 2 scores higher. You report the improvement.
Evaluate round 2 against a set the teacher has influenced
Why: If any pseudo-labelled tile reached your validation set — or if you re-split after adding them — you are measuring agreement between the student and its own teacher, which rises every round regardless of whether the model got better.
The number goes up on every round, forever, and it means nothing.
Ring-fence a manually labelled validation set before round 1 and never let a pseudo-label near it.
Freeze val once, evaluate everything against it
Why: It is the only fixed point in a loop where everything else is changing. Write those tile ids into splits.json and treat them as untouchable — no pseudo-labels, no re-splitting, no exceptions.
A self-training loop without a frozen human-labelled reference is not an experiment. It is a feedback circuit with a score attached.
Commit first
Predict first
You run three rounds of self-training with the temporal filter, and you track mean predicted instances per tile on the frozen validation set.
Round 1 gives 78 instances per tile. Round 2 gives 71. Round 3 gives 64. Your ground truth averages 83.
How confident are you that round 4 will improve things — and what is this trend telling you? Commit to an answer before reading on.
Correct: Not confident at all. The count is drifting steadily downward and away from the ground truth, which is the signature of confirmation bias: parcels the teacher misses become background in the student's training data, so each round the student learns to find fewer.
Stop at round 3, keep the round-2 weights, and look at what is being lost. A drifting count is a much earlier warning than a falling mAP.
Why: This is the specific way detection self-training collapses, and it is why mean instance count is the diagnostic to watch. A missed detection is not a neutral absence — it is taught to the student as confirmed background. The student then misses it more reliably, its own pseudo-labels contain it less often, and the loop reinforces itself. Monotonic drift in either direction means stop; stability near the ground-truth count means the loop is healthy.
Concept
Self-training fails quietly. Five rules keep it honest, and none of them is expensive.
The fifth is the one that turns this from a script into an experiment. Without a log you cannot say which change caused which effect, and a self-training loop has too many interacting knobs to reconstruct from memory.
Check
The mechanism, not the slogan.
Check your understanding
What is the fundamental reason temporal support is a better acceptance criterion than the detector's confidence score, for this dataset specifically?
Answer: A
Why: The parcel geometry is genuinely constant across the time series while the imagery genuinely varies, so each observation is close to an independent test of the same hypothesis. Agreement across many such tests is far stronger evidence than one model's self-reported confidence — especially since detector scores are known to be poorly calibrated. It also means the acceptance decision stops depending on where you put a threshold, which is the practical payoff.
scores tensor at inference — the code in this deck uses it as a loose floor. The argument is about which signal to trust, not which is available.Edge cases
Parameter explorer
Temporal consistency depends on geometry being static across the series. Push each parameter toward its extreme and find where the assumption breaks.
Being able to say where your own idea stops working is worth more in an interview than the idea itself. It is the difference between having read about a technique and having understood it.
Section
Part 8
Concept
Every part of this deck ended in a decision. A reader of your repo — an interviewer, a collaborator, you in three months — cannot tell your decisions from your accidents unless you say.
| Decision | The alternative | Why you chose it |
|---|---|---|
| one class (parcel) | 18 PASTIS crop types | geometry first; classification is a separable second problem |
| single-frame Mask R-CNN | U-TAE + PaPs | a deliberately simpler baseline, with temporal work planned |
min_size=256 | the default 800 | 2× upscale on 128 px tiles; 39× compute is not justified |
| tile-level split | random over observations | the tile is the unit of independence |
| temporal support filter | a fixed score threshold | support is evidence; a score is self-report |
That table is most of a DECISIONS.md. It took five minutes to write and it changes how the whole project reads: not "here is a Mask R-CNN tutorial I followed" but "here are five decisions I made and can defend."
Sorting
Sort into buckets
Each of these describes something you might do. Sort them by whether they actually produce evidence, or only produce the feeling of rigour.
Concept
The same move appeared in four different places in this deck. Worth naming so you can reuse it deliberately.
| Instead of | Measure | Then set |
|---|---|---|
| default anchor sizes | your box-side distribution | anchors that span it |
detections_per_img=100 | instances per image | a cap above your maximum |
| a guessed score threshold | temporal support | a support fraction |
| ImageNet normalisation | per-band statistics on train | fixed mean and std |
| "80/20 seems standard" | your independence structure | a split over the right unit |
Defaults are someone else's measurements. They are a reasonable starting point precisely to the extent that your data resembles theirs — and 128×128 tiles with 83 tessellating 16-pixel objects do not resemble COCO in any respect that matters.
Explain it
Discussion prompt
Pick the finding from this deck that surprised you most.
Explain it to an imaginary teammate who is competent but has never touched satellite data — in five sentences or fewer, with no jargon they would have to look up. Then answer their obvious follow-up: "so how would I have caught that myself?"
Hint: If you cannot explain the mechanism without saying "leakage" or "pseudo-label", you are relying on the word to carry meaning you have not yet supplied.
Answer:
The second half is the real test. A finding you can only recognise is a fact; a finding you can teach someone to detect is a skill — and only the second one transfers to the next project.
Real world
Be honest about what reads as strong to someone hiring for robotics or perception work.
| Reads as | Because |
|---|---|
| a tutorial follow-along | "I trained Mask R-CNN on a public dataset" |
| a competent engineer | "I found a leak in my own eval and fixed the split" |
| someone worth interviewing | "I used temporal consistency as a pseudo-label filter and ablated it against a score threshold" |
Discussion prompt
Write the two-sentence version of this project that you would put at the top of the README.
Constraint: neither sentence may name an architecture. If removing the model name empties the description, that is the finding.
Hint: Lead with the problem and the thing you did that someone else would not have.
Answer:
Something in this direction: "Hand-annotated parcel instance segmentation on Sentinel-2 time series, with a tile-level split and a round-trip annotation QA harness. Pseudo-labels are accepted by temporal consistency across the series rather than by a confidence threshold, which removes the tuned cutoff and improves precision by X% over a score-threshold baseline."
Note what is doing the work there: a decision, a check, and a measured comparison. The architecture is an implementation detail and belongs further down.
Concept
Sequenced so that nothing is built on something unverified. Roughly one session each.
splits.json at the tile level, and commit it.min_size, detection cap.Steps 1 to 3 produce no model and no metric, and they are the ones that determine whether anything after them means anything. That is normal, and being willing to spend three sessions there is itself a signal about how you work.
Connect it up
Draw it
One page. Put these six on it — round-trip QA · tile-level splitting · measure-then-configure · source vs derived · fixed normalisation · temporal support — and draw an arrow wherever one of them is what makes another possible. Label every arrow with the reason.
Two of them are the same underlying idea wearing different clothes. Find that pair and mark it.
The pair worth finding: tile-level splitting and fixed normalisation are both the rule that information must not cross the train/val boundary. One is about which samples, the other about which statistics — but it is one principle, and seeing that is worth more than memorising either.
Exit ticket
Predict first
Two answers, written down, before you open your editor:
1. Which finding from this deck would have cost you the most time had you not found it until after training?
2. What is the first command you will run tomorrow?
Correct: Most people say the tile-level leak for the first — it would have invalidated every number produced up to that point, including the comparisons used to choose everything else. The second should be the QA audit over your full annotation directory, because every other fix depends on knowing the current state of the data.
python annotation_qa.py audit --ann <your annotations> --masks <your masks>
Send me the output before the next session and we will start from what it says.
Why: The first question is about consequence, and the leak wins because it corrupts the measurement apparatus rather than the model: with a leaked validation set, every subsequent decision is made on the basis of a number that does not mean what it claims. The second is about sequencing — the audit is first because it is the only step that tells you which of the other problems you actually have, and how bad they are on your real data rather than on the five files you sent.
Recap
You started with working annotation code and a reasonable plan. What changed is the layer above both.
| Problem | Fix | How you will know it worked |
|---|---|---|
| 61 copies, 1 label | tile-level splits.json | disjoint tile ids; val mAP drops and becomes real |
| objects vs defaults | min_size + anchors from measurement | small-object AP rises; step time falls |
| invisible coordinate bug | one source of truth + overlay | output survives changing annotation_size |
| polygons discarded | COCO as source | you can fix one parcel without redrawing a tile |
| per-image normalisation | fixed per-band stats | same tile, same prediction, every time |
| guessed threshold | temporal support | result stops depending on the cutoff |
One more thing worth saying: none of this was a criticism of your annotator. It works, and the hard part — 83 careful polygons — was already done. Everything here was about making that work provable, which is the difference between a project that runs and a project someone else can trust.
Want this taught 1-on-1? Alexander tutors Computer Vision — $55/session, free consultation.