Dataset Engineering for Instance Segmentation (PASTIS Parcels)

A working-session deck for a hand-annotated PASTIS parcel instance-segmentation project, built around six findings measured from the student's own files: 83 instances per 128x128 tile, box sides of 4 to 38 px, and annotations byte-identical across all 61 observations of a tile. It covers round-trip QA of written artifacts, why per-observation duplication makes a random split leak an entire validation set, configuring Mask R-CNN anchors and min_size from measured object scale instead of COCO defaults, the dormant coordinate bug created by a duplicated 850 literal, polygons as source versus rasters as build output via COCO, fixed per-band normalisation, and a temporal-consistency filter that replaces a guessed pseudo-label confidence threshold with agreement across the time series.

Subject: Computer Vision · 84 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Your Labels Are Fine. Your Dataset Is Not.

Title

Computer Vision · Instance Segmentation

You drew 83 parcels by hand and the polygons are good. The problem is one layer up — in what the pipeline writes to disk, and in what you are about to ask a model to learn from it.

2. What you will be able to do

Objectives

Your annotator works. It has undo, per-polygon commit, an escape hatch for cloudy dates, and it rasterises correctly. This deck is not about fixing it — it is about the layer above it, where the expensive mistakes live.

3. Before we open anything: what does your pipeline produce?

Warm-up

Answer from memory. You wrote this code, so this should be quick — and if it is not quick, that is itself the finding.

Discussion prompt

For one PASTIS tile, annotate_first_apply_to_all finishes and writes files to disk.

Without looking: how many files, and what is in each one? Then the harder half — how many of those files contain different information from one another?

Hint: Count the loop. save_annotations_for_all_observations runs once per observation, and the tile has around 61 of them.

Answer:

Three files per observation — _boxes.npy, _labels.npy, _masks.npy — times ~61 observations, so roughly 183 files per tile.

The second half is the one that matters. All 61 observations receive the same arrays. I hashed the files you sent: obs 000 through 004 are byte-identical. So 183 files carry the information content of three.

4. The four questions a dataset has to answer

Concept

A model is the easy part. Datasets are where projects quietly fail, and they fail in a small number of repeatable ways. Every one of them is a question you can ask before you train.

QuestionWhat goes wrong when you skip itWhere we answer it
Are the labels correct?You optimise toward noise and never find outPart 1
How many independent labels are there?Validation measures memorisationPart 2
What do the objects look like, numerically?Framework defaults silently mismatch your dataPart 3
Can you regenerate everything from source?One fix means redoing all the labellingPart 5

Notice that none of these mention architecture. You picked Mask R-CNN, which is a reasonable choice — and it is very nearly the least important decision on this list.

5. One question, asked seriously

Socratic

Discussion prompt

You have annotated one tile: 83 parcels, drawn by hand, saved to disk.

How do you know they are correct?

Not "I was careful" — that is a claim about your intent, not about the bytes on disk. What procedure would show you a mistake if one existed?

Hint: You looked at an image and clicked. Then a series of transformations happened to those clicks. Which of those transformations have you ever seen the output of?

Answer:

Right now there is no such procedure. Nothing in the pipeline reads a saved .npy back and puts it next to the image it describes. The annotations are write-only.

That is the single most important structural gap in the project, and everything in Part 1 follows from closing it.

6. Round-Trip Everything

Section

Part 1

7. Write-only data is unverifiable data

Concept

Your pipeline is a one-way street. Clicks go in at 850×850, get scaled to 128×128, get rasterised, get saved. At no point does anything load a saved file and ask does this match the image it came from?

Round-trip check — Load the artifact you wrote, render it against the input it describes, and look. If the two disagree, you have found a bug that no unit test was ever going to catch.

This matters more in vision than almost anywhere else, because the failure mode is geometric. A coordinate scaling bug does not raise an exception. It does not produce a NaN. It produces perfectly well-formed arrays of plausible numbers that happen to describe the wrong parts of the image.

The model will then train happily on them, converge, and report a loss curve that looks exactly like a healthy one.

8. Five files, one tile

Anomaly

Here is what you sent me, as bytes rather than as filenames:

FileShapedtypeMD5 (first 12)
S2_10000_obs_000_boxes.npy(83, 4)float322eb1625e4fde
S2_10000_obs_002_boxes.npy(83, 4)float322eb1625e4fde
S2_10000_obs_004_boxes.npy(83, 4)float322eb1625e4fde
S2_10000_obs_000_labels.npy(83,)int648fc8dd18697e
S2_10000_obs_003_labels.npy(83,)int648fc8dd18697e

Predict first

Look at the last column before reading on.

Every hash in a given group is identical. Two questions:

1. Is this a bug?
2. Whether or not it is a bug, what does it do to the size of your training set?

Correct: It is not a bug — it is exactly what annotate_first_apply_to_all was written to do, and the reasoning behind it is sound. But it means 61 files per tile carry one tile's worth of information, so your dataset has ~61× fewer independent labels than it has samples.

Hold onto the second answer. Part 2 is entirely about it.

Why: The duplication is deliberate and correct as a modelling decision: parcel geometry does not move across a growing season, so re-drawing it per date would be wasted labour. The trap is that it is also a statement about your dataset's independence structure, and nothing downstream knows that. A random train/val split over observations will put the same geometry in both halves.

9. What a round-trip check actually tells you

Worked example

This is annotation_qa.py run against your files. It loads every saved array, regroups them by tile and observation, and checks each one against the others and against itself.

5 annotation records across 1 tiles

--- duplication ---
1/1 tiles have byte-identical annotations across all their observations.

--- object scale (this drives your anchor config) ---
box side px, native 128:  min 4  p25 12  median 16  p75 20  max 38
instances per tile: median 83, max 83

--- hard checks ---
  FAIL  S2_10000 obs 1: labels/masks present but no boxes
  FAIL  S2_10000 obs 3: labels/masks present but no boxes
LineWhat it isWhat you do about it
1 tile, 5 recordsindependence structuresplit by tile, not by record
byte-identicalexpected duplicationstore once, serve many
side 4–38 pxobject scaleset anchors from this
83 instancesdensityraise box_detections_per_img
2 FAILsmissing filesin this case an upload artifact

The two FAIL lines are not a real defect — you simply did not send boxes for observations 1 and 3. But notice what happened: the check found the gap without being told what to look for. That is the whole argument for having one.

10. Recipe: the data check you run before every training run

Pattern

Five steps. Run them as a script, not as a notebook cell you scroll past, and make the script exit non-zero when something fails so it can move into CI later.

  1. Load what you wrote — read the saved arrays back from disk, not from memory.
  2. Overlay onto the source image — the only check that catches geometric errors.
  3. Recompute the derived fields — a mask's own bounding box must equal the box you saved beside it.
  4. Print the distributions — object sizes, instance counts, class balance. These become configuration inputs.
  5. Assert the independence structure — how many genuinely independent labels do you have?

Step 3 is the sharpest one and the cheapest to write. You already have two representations of the same fact — the polygon-derived box and the mask — so make them check each other.

def boxes_from_masks(masks):
    """Recompute boxes with the same +1 exclusive convention as
    create_training_targets, so the two are directly comparable."""
    out = np.zeros((len(masks), 4), dtype=np.float32)
    for i, mask in enumerate(masks):
        ys, xs = np.where(mask == 1)
        out[i] = [xs.min(), ys.min(), xs.max() + 1, ys.max() + 1]
    return out

drift = np.abs(boxes_from_masks(masks) - saved_boxes).max()
assert drift <= 0.5, f"coordinate bug: {drift:.1f}px drift"
DriftMeaning
0 pxmasks and boxes agree — geometry is internally consistent
± a few pxa rounding or convention mismatch between the two paths
large / erratica scaling bug — the two were computed in different coordinate spaces

11. Why none of this surfaced on its own

Concept

It is worth being precise about why these problems are invisible, because the reason generalises far past this project.

No exception
Every array has the right shape and dtype. NumPy has nothing to complain about.
No test
There is nothing to assert against — the correct answer lives in an image, not in a fixture.
No symptom
Training runs. Loss decreases. Metrics are computed. All of it is consistent with the bug.

This is the defining property of data bugs: they are silent by construction. Code bugs announce themselves; data bugs just quietly shift what your model learns. The only defence is to build the observation channel yourself, deliberately, before you need it.

12. Check: what a round-trip catches

Check

Reason about the mechanism, not the vibe.

Check your understanding

You add a check that recomputes each mask's bounding box and compares it to the saved box. On your current data it reports 0 px drift everywhere. What have you actually proven?

  • A. That the mask and the box were computed in the same coordinate space — but not that either of them lands on a real parcel, since both derive from the same polygon. (correct)
  • B. That the annotations are correct, since the two independent representations agree.
  • C. Nothing at all — the check is circular and cannot fail.
  • D. That the display-to-native scaling is correct, since a scaling error would change the box.

Answer: A

Why: Both artifacts are computed from the same polygon inside create_training_targets, so agreement between them rules out a bug between those two steps and nothing more. If the polygon itself was recorded in the wrong coordinate space, the mask and the box will be wrong together and agree perfectly. That is exactly why the overlay against the image is a separate, non-negotiable check — it is the only one that tests the polygon itself.

Why B tempts people
They are not independent. Both are derived from the same points array a few lines apart, so their agreement is a weaker claim than it looks.
Why C tempts people
It can fail — a convention mismatch (inclusive vs exclusive max, a stray rounding) would show up here, and those are real bugs worth catching.
Why D tempts people
A scaling error applies to the polygon before either artifact is built, so it moves the mask and the box together and the comparison stays clean.

13. Say it back

Explain it to yourself

Discussion prompt

In your own words, and without using the phrase "good practice":

Why is a check that compares two things you derived from the same source still worth writing — and what is it specifically unable to tell you?

Hint: Think about which pairs of pipeline stages the check sits between, and which stages it sits entirely downstream of.

Answer:

A good answer names the boundary: the check covers everything from points onward, and covers nothing before it. Cheap, genuinely useful, and honest about its limit.

Being able to state what a test cannot see is most of what separates a useful test suite from a reassuring one.

14. 61 Copies, One Label

Section

Part 2

15. The duplication is correct. The bookkeeping is not.

Concept

Start by giving the design its due, because the reasoning behind it is right:

def save_annotations_for_all_observations(
    file_path, number_of_images, boxes, labels, masks,
    annotation_path, mask_path,
):
    for observation_index in range(number_of_images):
        annotation_name = f"{file_path.stem}_obs_{observation_index:03d}"
        np.save(annotation_path / f"{annotation_name}_boxes.npy", boxes)
        np.save(annotation_path / f"{annotation_name}_labels.npy", labels)
        np.save(mask_path / f"{annotation_name}_masks.npy", masks)
ClaimTrue?Why
Parcel geometry is static across a seasonYesfield boundaries do not move between satellite passes
So one annotation serves every dateYesre-drawing per date would be wasted labour
So every date is a training sampleYesdifferent reflectance, same geometry — real augmentation
So every date is an independent labelNothis is the step that does not follow

The first three are sound. The fourth is the one nothing in the code disagrees with, because nothing in the code has any opinion about independence at all.

16. Put a number on it

Estimation

Estimate before you calculate. Being roughly right here is the skill; being precise is arithmetic.

Predict first

You finish your plan: 100 tiles, annotated by hand, each with about 61 observations.

1. How many training samples is that?
2. How many independent labels is that?
3. If you use a standard 80/20 random split over samples, roughly what fraction of your validation tiles have their geometry already memorised from training?

Correct: About 6,100 samples; exactly 100 independent labels; and essentially 100% of validation tiles will have appeared in training.

0.2^61 ≈ 10^-43. There is no sampling luck that rescues this — it is structural.

Why: With 61 observations per tile and a random 80/20 split over the ~6,100 samples, the chance that a given tile contributes zero of its 61 observations to the training set is 0.2^61 — indistinguishable from zero. So effectively every tile in your validation set has had its exact parcel geometry seen during training. Your val mAP would then be measuring how well the model recalls 100 specific field layouts, not how well it segments fields.

17. Sample ≠ independent label

Concept

This distinction has a name in statistics and a different name in every applied field, which is part of why it keeps being rediscovered the hard way.

Grouped data — Observations that share a latent source. Within a group, samples are correlated; across groups they are not. The group — not the sample — is the unit of independence.

Figure (svg): Two split strategies. In a random split over observations, coloured squares representing tile A and tile B appear in both the train and validation boxes. In a tile-level split, every square from tile A is in train and every square from tile B is in validation.

The same structure, under other names: multiple photographs of one patient; several sentences from one document; repeated measurements on one subject; several frames from one video. In every case, splitting at the sample level and reporting the resulting score is the same mistake.

18. The split that looks careful and is not

Worked example

This is what almost everyone writes first. It uses a fixed seed, it shuffles properly, it even stratifies. It is still wrong.

# WRONG for grouped data, however careful it looks.
records = sorted(annotation_dir.glob('*_boxes.npy'))   # ~6100 entries
rng = np.random.default_rng(0)
rng.shuffle(records)

cut = int(0.8 * len(records))
train, val = records[:cut], records[cut:]

# S2_10000_obs_007 -> train
# S2_10000_obs_031 -> val      <-- same 83 parcels, same geometry
Property of this codeReads asActually
fixed seedreproduciblereproducibly wrong
shuffledunbiasedunbiased within the wrong unit
80/20standardstandard for independent samples
one file = one sampleobviousthe assumption that fails

Every line of it is defensible in isolation. The error is in a premise that is never written down anywhere: that one file is one independent thing.

19. Trap: trusting a validation score you did not earn

Trap

The trap

You train, and validation mAP comes back at 0.87. You are pleased. You tune a little, get 0.89, and write it in the README.

Accept the number because it is high

Why: A high validation score feels like evidence that everything upstream is fine, so it quietly certifies the split, the labels, and the config all at once — none of which it actually tested.

Then the model meets a region it has not seen and collapses to 0.3, and you have no idea which of the twenty things you changed is responsible — because your baseline was never real.

The fix

Split by tile first, before the first training run, and write the split to a file you commit.

Expect the honest number to be lower

Why: A tile-level split on 100 tiles might report 0.55 instead of 0.87. That is not a regression — it is the first number you have that means anything, and every later improvement is measured against it.

A believable 0.55 is worth more than a flattering 0.89, because you can improve the first one on purpose.

20. Split by group, and commit the split

Concept

Two separate ideas, and both matter.

The second point is the one people skip. If your split is a function of the directory listing, then annotating tile 101 reshuffles everything and last week's number is no longer comparable to this week's. Write it down once and read it thereafter.

21. A split worth committing

Worked example

Short, boring, and load-bearing. Note that it operates on tile ids, never on file paths.

import json, numpy as np
from pathlib import Path

def build_splits(tile_ids, out_path, val_frac=0.2, seed=0):
    """Freeze a tile-level train/val split. Tiles, never observations."""
    tiles = sorted(set(tile_ids))          # sorted => seed alone decides
    rng = np.random.default_rng(seed)
    order = rng.permutation(len(tiles))
    n_val = max(1, round(val_frac * len(tiles)))

    val = sorted(tiles[i] for i in order[:n_val])
    train = sorted(tiles[i] for i in order[n_val:])
    assert not set(train) & set(val), 'tile in both splits'

    Path(out_path).write_text(json.dumps(
        {'seed': seed, 'val_frac': val_frac, 'train': train, 'val': val},
        indent=2))
    return train, val
LineWhy it is there
sorted(set(...))makes the result depend on the seed alone, not on filesystem order
assert not set(train) & set(val)the leak this whole part is about, checked rather than assumed
writes seed into the filethe split can be regenerated and audited later
returns tile idsthe Dataset expands a tile into its observations, so leakage cannot re-enter

22. Complete the loader

Fill the middle

The split gives you tiles. The Dataset has to turn tiles into samples without letting the leak back in.

Fill in the blanks

class ParcelDataset(torch.utils.data.Dataset):
def __init__(self, tile_ids, ann_dir, img_dir):
# one entry per (tile, observation) — samples, not labels
self.index = [(t, o) for t in tile_ids
for o in range(n_observations(t))]

def __getitem__(self, i):
tile, obs = self.index[i]
image = load_observation(tile, obs) # varies with obs
target = load_annotation(tile) # constant per tile
return image, target

Why: Blank A takes only the tile ids handed to this split, so a tile can never appear in two datasets. Blank B keys the annotation on the tile alone — not on (tile, obs) — which is the line that encodes the real structure of the data: the image varies with the observation, the label does not. Once the loader is written this way, the 61 duplicate files on disk have no reason to exist.

23. What the duplication costs, in bytes

Cost model

Annotate

Work through the storage arithmetic. The numbers are from your own tile.

  • uint8 is already the right choice — a boolean mask in bool dtype is also 1 byte per pixel in NumPy, so nothing is wasted here.
  • 1.36 MB per observation is unremarkable on its own. Every storage problem starts out unremarkable.
  • 83 MB per tile is the first number that should feel wrong, because you only drew 83 polygons once.
  • 8.3 GB versus 136 MB — a factor of 61, which is exactly the duplication factor. The disk is telling you the independence structure if you listen.
  • The read cost matters more than the disk cost: every epoch pulls 61× more bytes through the loader than the information warrants, and that is dataloader time you never get back.

Storage is cheap and this is still worth fixing — not because 8 GB is expensive, but because the ratio is a measurement of a design error. When a redundancy factor in your storage exactly matches a redundancy factor in your labels, the storage layout is telling you what the data actually is.

24. Before and after, in the terms that matter

Comparison

Comparison matrix

Fill the After column from what you now know. The shape of the fix should be predictable by this point.

PropertyBeforeAfter
annotation files per tile~1833
mask bytes, 100 tiles~8.3 GB~136 MB
unit of the train/val splitobservation (file)tile
independent val labels0 — all seen in train20 unseen tiles
training samples~6,100~6,100 — unchanged

The last row is the one worth pausing on. You do not lose training data by fixing this — every observation is still a sample, and the reflectance really does vary across dates. What changes is only the bookkeeping: which of those samples are allowed to be used to judge the model.

25. Check: grouped splits

Check

One of these is a real defence. The others are things that feel like defences.

Check your understanding

You want to be confident there is no geometry leakage between your train and validation sets. Which check actually establishes that?

  • A. Assert that the sets of tile ids in train and val are disjoint. (correct)
  • B. Assert that no filename appears in both splits.
  • C. Confirm the split used a fixed seed and was shuffled before slicing.
  • D. Confirm validation mAP is lower than training mAP, which shows the model has not memorised val.

Answer: A

Why: Leakage here is defined at the tile level, because the tile is what determines the label. Disjoint tile ids is therefore the exact statement of the property you want, and it is one line to assert. Everything else either checks a weaker property or checks nothing at all.

Why B tempts people
Filenames are already unique per observation, so this assertion passes trivially in exactly the broken setup — obs_007 and obs_031 are different filenames carrying identical labels.
Why C tempts people
Reproducibility and leakage are unrelated. A seeded shuffle over the wrong unit gives you the same leak every time, which is worse than an unreproducible one because it looks disciplined.
Why D tempts people
A train/val gap is consistent with leakage, not evidence against it. With duplicated geometry the val score is inflated and can still sit below the train score.

26. Where else this exact bug lives

Real world

Discussion prompt

The pattern is: many samples share one latent source, and the label is a property of the source rather than of the sample.

Name three settings outside satellite imagery where this holds. For each, say what the group is — then say what a careless engineer would split on instead.

Hint: Anywhere a single real-world entity gets measured, photographed, or recorded more than once.

Answer:

SettingGroup (correct unit)What gets split on by mistake
Medical imagingpatientimage / slice
Speech recognitionspeakerutterance
Document NLPdocumentsentence
Video understandingclipframe
Recommendersuserinteraction
Industrial inspectionphysical partphotograph

sklearn.model_selection names this directly: GroupKFold, GroupShuffleSplit, StratifiedGroupKFold. When a library gives a concept its own family of classes, that is a strong signal about how often it is needed.

27. The one-sentence version

Intuition

If you remember nothing else from this part:

The rule — Split on whatever generated the label. If two samples share a label because they share a source, they belong on the same side of the split — always.

It generalises because it is not really about datasets. It is about not letting the unit you happen to store masquerade as the unit you actually sampled.

28. Measure, Then Configure

Section

Part 3

29. Your objects, as numbers

Concept

Everything in this part follows from five numbers I measured from your S2_10000_obs_000_boxes.npy. None of them is a problem. All of them are configuration inputs you have not used yet.

MeasurementValueWhy the model cares
instances per tile83detection cap, RPN sampling density
box side, median16 pxanchor sizes, FPN level assignment
box side, min4 pxthe smallest thing you are asking it to find
box side, max38 pxthe scale range it must span — narrow
Σ box area ÷ image area1.36×boxes overlap; parcels tile the plane

That last row is worth a second look. The bounding boxes cover 136% of the image, which is only possible because they overlap heavily — exactly what you expect from agricultural parcels, which tessellate rather than sit isolated on a background. Object detectors are mostly designed and benchmarked on the opposite assumption.

30. What happens to a 128-pixel tile

Prediction

torchvision.models.detection.maskrcnn_resnet50_fpn wraps your input in a GeneralizedRCNNTransform before the backbone ever sees it. Its defaults are min_size=800, max_size=1333.

Predict first

You pass a 128×128 PASTIS tile into that model, unmodified.

What does the transform do to it, and what does that do to your median 16-pixel parcel?

Correct: It upscales the tile 6.25× to 800×800 by bilinear interpolation, which turns the median 16 px box into a 100 px box — and, by accident, lands your objects squarely inside the default anchor range of 32–512.

The lesson is not "the defaults are wrong". It is that you did not know what they were doing, and a 6.25× resize is not something you want to discover by accident.

Why: This is the detail that makes the situation confusing rather than simply broken: the default resize accidentally rescues the default anchors. Your model will train, and the numbers will not be absurd. But you are now paying full 800×800 COCO-scale compute for a 128×128 tile, and the network is spending capacity on interpolated detail that carries no sensor information. Neither of those shows up as an error — only as a slow training run and a ceiling you cannot explain.

31. Follow one box through the pipeline

Invariant

Step through it

Track your median parcel — a 16-pixel box — from disk to loss. Watch which stage changes its size and which stage interprets that size.

  1. Truth on disk
  2. Loader: unchanged
  3. Transform: 6.25x upscale
  4. RPN: anchor matching
  5. Mask head: resolution exceeds signal

Frame 3 is where the surprise is, and frame 5 is where the ceiling is. A 28×28 mask prediction over a genuinely 16-pixel object is mostly interpolation — so expect mask AP to trail box AP, and do not spend a week hunting for the bug that causes it. That gap is a property of the data.

32. What anchors are actually claiming

Concept

An anchor is a hypothesis about size. The RPN lays a grid of boxes of fixed sizes over each FPN level and asks, for each one, is there an object roughly this big here?

Anchor–object scale match — The assumption that your objects' size distribution overlaps the anchor size distribution. When it does not, the RPN proposes boxes that never sufficiently overlap a ground-truth box, those anchors are labelled negative, and the detector learns to predict nothing.

Figure (svg): Two number lines. The top shows the student's object sizes clustered between 4 and 38 pixels at the far left. The bottom shows default anchor sizes at 32, 64, 128, 256 and 512, spread across the full range. A caption notes that the 6.25x resize is what moves the object band rightward to overlap the anchors.

Two ways to make those bands overlap: resize the image until the objects reach the anchors, or move the anchors down to the objects. The defaults do the first, invisibly. Doing the second, deliberately, is cheaper and says what you mean.

33. Trap: copying anchor sizes from a blog post

Trap

The trap

You read that small-object detection needs smaller anchors, so you set sizes=((8,), (16,), (32,), (64,), (128,)) and leave min_size=800 alone.

Change one thing without checking the other

Why: Anchors are compared against boxes after the transform has already scaled them 6.25x. Your 100 px transformed boxes now sit above almost every anchor you just configured, so you have moved the bands apart rather than together.

Recall drops, you conclude that small anchors do not help, and you revert a change that was pointing in the right direction.

The fix

Fix min_size first, then measure the box sizes the RPN will actually see, then set anchors from that measurement.

Print the box sizes after the transform, not before

Why: One print inside a GeneralizedRCNNTransform call tells you the real numbers. Configuration derived from a measurement survives a change of dataset; configuration copied from a blog post does not.

min_size=256 with anchors of roughly 8–128 is a defensible starting point for these tiles — and you will be able to say why, which is the part that matters.

34. Configuring from the measurement

Worked example

Nothing exotic. The point is that every number here traces back to something printed from your own data.

from torchvision.models.detection import maskrcnn_resnet50_fpn
from torchvision.models.detection.rpn import AnchorGenerator

model = maskrcnn_resnet50_fpn(weights='DEFAULT')

# 1. Stop the 6.25x upscale. 128 -> 256 is a deliberate 2x.
model.transform.min_size = (256,)
model.transform.max_size = 256

# 2. Anchors spanning 2x the measured range (4-38 px -> 8-76 px).
model.rpn.anchor_generator = AnchorGenerator(
    sizes=((8,), (16,), (32,), (64,), (128,)),   # one per FPN level
    aspect_ratios=((0.5, 1.0, 2.0),) * 5,        # parcels are irregular
)

# 3. You measured 83 instances per tile. The default cap is 100.
model.roi_heads.detections_per_img = 300
SettingDefaultYoursMeasured from
min_size800256128 px tiles — a 2× upscale, chosen not inherited
anchor sizes32–5128–128box sides 4–38 px, doubled by the resize
detections_per_img10030083 instances per tile, with headroom
aspect_ratios0.5/1/2unchangedparcels are irregular; no reason to narrow it

Treat these as a starting hypothesis, not an answer. The habit being taught is that each one has a reason you can state out loud, so when a number needs to change you know which measurement to go back to.

35. Feel the trade-off

Tweak it

Parameter explorer

Three knobs, all interacting. Move one at a time and predict the effect on recall of small parcels, on GPU memory, and on step time — before you reason it through.

  • min_size (input resize) — from 128 to 800
  • smallest anchor (px) — from 4 to 64
  • detections_per_img — from 50 to 500

The interaction that catches people: min_size and anchor size are not independent knobs. Doubling the resize halves the effective anchor size relative to your objects. Change one, and you have implicitly changed the other.

36. Choosing min_size on purpose

Trade off

Comparison matrix

Fill in the blanks. There is no single right answer here — there is a right way to decide.

min_sizeEffective object sizeCompute vs 128When this is the right call
128 (none)4–38 px1×fastest iteration; risks the 4 px tail being unlearnable
2568–76 px4×sane default: small objects get real pixels, cost stays moderate
51216–152 px16×if the smallest parcels are demonstrably being missed
800 (default)25–238 px39×almost never here — inherited, not chosen

Compute scales with the square of the resize, which is why 800 costs 39× rather than 6.25×. That is the number to have in mind when a training run feels inexplicably slow.

37. Symptom to knob

Discrimination

Sort into buckets

Each symptom points at one stage. Sort them by which stage you would actually change — this is the diagnosis-first habit applied to configuration.

Anchors / resize
The smallest parcels are never detected at all, at any score threshold.; Training is far slower per step than the 128×128 tile size suggests.
Detection cap
A full tile returns exactly 100 instances, never more.
Inherent to the data
Box AP is respectable but mask AP is far behind it.
NMS / IoU thresholds
Adjacent parcels are merged into one detection.
b1
Both scale symptoms trace to the same pair of interacting settings: objects that never overlap an anchor cannot be proposed, and an oversized resize is pure compute you did not ask for.
b2
A count that stops at exactly the default is a cap, not a model failure. detections_per_img=100 against 83+ ground-truth instances is a ceiling you will hit on dense tiles.
b3
A 28×28 mask head over a 16-pixel object cannot resolve more detail than the sensor captured. Expect this gap, measure it, and do not chase it as a bug.
b4
Parcels tessellate, so their boxes overlap heavily — Σ box area is 1.36× the image. Default NMS IoU thresholds assume more separation than agricultural fields provide.

38. Check: reading a configuration mismatch

Check

One symptom, four explanations, one mechanism.

Check your understanding

You keep min_size=800 but set anchor sizes to ((4,), (8,), (16,), (32,), (64,)) to match your measured 4–38 px boxes. Recall gets dramatically worse. Why?

  • A. The transform scales boxes by 6.25× before the RPN sees them, so the ground truth is 25–238 px while the anchors top out at 64 — almost nothing matches. (correct)
  • B. Anchors below 8 px are smaller than the FPN's finest stride, so they cannot be placed on the feature grid.
  • C. Smaller anchors always reduce recall because each one covers less area, so fewer objects are proposed.
  • D. The pretrained RPN head weights expect the default anchor sizes and must be retrained from scratch.

Answer: A

Why: Anchor sizes are expressed in the coordinate space the RPN operates in — which is after GeneralizedRCNNTransform has resized both image and boxes. Setting anchors from your native 128 px measurements while leaving an 800 px resize in place puts the two distributions on opposite sides of each other. The fix is to make min_size and the anchor sizes agree, in that order.

Why B tempts people
Stride does constrain how finely anchors can be placed, but it is not what is happening here: the mismatch is a factor of 6.25 in scale, which is far larger than any stride effect.
Why C tempts people
Smaller anchors improve recall on small objects — that is the entire reason to shrink them. The failure here is the mismatch with the resize, not the direction of the change.
Why D tempts people
The RPN head is convolutional and its weights are not tied to specific anchor sizes; replacing the anchor generator is a supported and common change.

39. Read the config like a sentence

Notation

Annotate

Every argument below is a claim about your data. Read each one as a sentence and decide whether it is true of PASTIS parcels.

  • num_classes=2 says: there is exactly one kind of foreground thing. Index 0 is always background in torchvision, which is why your labels = np.ones(...) is correct and not off by one.
  • min_size=max_size=256 says: every tile is square and the same size. True for PASTIS, false for most datasets — worth a comment so the next reader knows it is deliberate.
  • Five anchor sizes says: there are five FPN levels. The tuple length is not free — it must match the number of feature maps the backbone emits.
  • box_detections_per_img=300 says: I expect up to 300 parcels in a tile. You measured 83. Headroom is cheap; truncation is silent.
  • box_score_thresh=0.05 says: keep low-confidence detections. Correct for computing mAP, wrong for a demo — mAP integrates over thresholds and wants the full ranking.
  • rpn_batch_size_per_image=256 says: sample 256 anchors per image for the RPN loss. With 83 dense ground-truth objects your positive/negative balance differs from COCO's — worth checking before you tune anything else.

40. The Bug You Cannot See

Section

Part 4

41. Two resizes, one assumption

Concept

This one is currently harmless, which is exactly what makes it worth an entire section. Here is the caller:

annotation_size = 850

display_rgb = cv2.resize(
    rgb, (annotation_size, annotation_size),
    interpolation=cv2.INTER_LANCZOS4)

polygons = create_mask.draw_polygons(
    rgb=display_rgb, file_name=file_path.name,
    observation_index=observation_index)

boxes, labels, masks = create_mask.create_training_targets(
    polygons=polygons,
    display_size=(annotation_size, annotation_size),   # <- 850
    native_size=(width, height))
ValueWhere it comes fromWhere it is used
annotation_size = 850the callerresizing the image the annotator sees
display_size=(850, 850)the same variablescaling clicks back down to 128
850 (hardcoded)inside draw_polygonsresizing the canvas that receives the clicks

Three places, two sources. Two of them move together when you edit the variable. The third does not.

42. Find it yourself

Error analysis

Annotate

This is the drawing loop inside draw_polygons. The image arriving in rgb has already been resized to 850×850 by the caller. Find the line that will break the moment annotation_size changes.

  • Line 7: the canvas is resized to a literal 850, with no reference to the annotation_size the caller used.
  • Today annotation_size is also 850, so this resize is a no-op and everything works. The bug is dormant, not absent.
  • Set annotation_size = 900 and the caller hands in a 900×900 image, which this line squashes back to 850. Clicks are then recorded in 850-space.
  • But create_training_targets is told display_size=(900, 900), so it scales those 850-space clicks by 128/900 instead of 128/850.
  • Every polygon shrinks by about 5.5% toward the origin. No exception, no warning — just 83 slightly wrong parcels, saved and trained on.
  • The resize also runs every frame inside the while True loop, which is wasted work even in the version that is correct.

43. What a 5.5% shift actually costs

Worked example

It is tempting to think a small systematic offset is a small problem. Work it through against your measured object sizes.

annotation_size = 900          # you changed this
clicks_recorded_in = 850       # draw_polygons still resizes to 850

assumed_scale = 128 / 900      # what create_training_targets uses
actual_scale  = 128 / 850      # what the clicks actually needed

error = 1 - (assumed_scale / actual_scale)    # 0.0556

median_box = 16                # px, measured from your data
print(median_box * error)      # 0.89 px on a 16 px object
print(38 * error)              # 2.1 px on the largest
ObjectNative sizeShiftAs a fraction of the object
smallest parcel4 px0.22 px5.6%
median parcel16 px0.89 px5.6%
largest parcel38 px2.1 px5.6%
tile corner128 px7.1 px—

A 5.6% systematic shift toward the origin, on every parcel, in every tile annotated after the change. Mixed in with correctly-annotated tiles it becomes label noise with a direction — which is worse than random noise, because the model can learn it.

44. The class of bug this belongs to

Concept

Coincidental correctness — Code that produces the right answer because two independent values happen to be equal, rather than because it computes the relationship between them.

There is a second instance of exactly this in the same file:

height = img_array.shape[2]
width  = img_array.shape[3]

create_training_targets(
    polygons=selected_polygons,
    display_size=(annotation_size, annotation_size),
    native_size=(width, height))

# and inside:
#     native_width, native_height = native_size
#     points[:, 0] *= native_width  / display_width
#     points[:, 1] *= native_height / display_height
QuestionAnswer
Is the (width, height) order correct here?Yes — it matches the unpacking inside
Would a transposed order be caught?No — PASTIS tiles are 128×128
What makes it correct?H == W, not the code
When does it break?The first non-square input, silently

Both bugs share a signature worth learning to recognise: a value used in two places, where only one of them derives it from the other. When you see a literal that duplicates a variable, that is the shape.

45. Trap: "it works, so it is correct"

Trap

The trap

The overlay looks right. The masks land on parcels. You conclude the annotation code is correct and move on to the model.

Treat passing output as proof of a correct mechanism

Why: Your test exercised exactly one configuration — the one where the two 850s coincide. A test that cannot distinguish 'correct' from 'coincidentally correct' has not tested the thing you think it has.

Three weeks later you bump the annotation window size for a larger monitor, and quietly corrupt every tile you annotate after that.

The fix

Remove the duplicated value so the two cannot disagree, then re-run the overlay to confirm nothing moved.

Make the relationship explicit, then vary the input

Why: Pass annotation_size into draw_polygons and derive the canvas from it. Then annotate one tile at 900 and run the overlay: if the masks still land on parcels, the relationship is real and not a coincidence.

One parameter, one source of truth, and a check that varies it. That is the whole fix.

46. Repair it

Faded example

Three blanks. The goal is a function where the display size can be changed from one place without silently corrupting anything.

Fill in the blanks

def draw_polygons(rgb, file_name, observation_index,
display_size): # <- passed in, not assumed
display_image = np.clip(rgb * 255, 0, 255).astype(np.uint8)
display_image = cv2.cvtColor(display_image, cv2.COLOR_RGB2BGR)

# Assert the contract instead of hoping for it.
if display_image.shape[0] != display_size:
raise ValueError('image does not match the annotator display size')

while True:
canvas = display_image.copy() # no resize: it already fits
# ... draw polygons, read clicks off canvas ...

# caller — one source of truth for the number:
polygons = draw_polygons(rgb=display_rgb, file_name=name,
observation_index=i,
display_size=annotation_size)

Why: Blank A turns the silent assumption into a loud failure: the function now asserts the contract it depends on instead of hoping for it. Blank B deletes the second resize entirely — the image already arrives at the right size, so copying is all the frame loop needs, and the per-frame resize disappears as a bonus. Blank C passes the single source of truth down from the caller, so annotation_size now controls every place the number is used. The variable and the literal can no longer disagree because the literal is gone.

47. Which claims survive?

Two truths and a lie

Sort into buckets

Some of these are true about the annotation code as it stands today. Others are the kind of thing that sounds true and would get you into trouble. Sort them.

True
The polygons currently saved to disk are geometrically correct.; The native_size=(width, height) argument is passed in the correct order.; The per-frame cv2.resize inside the draw loop is wasted work.
Sounds true, is not
Because the current output is correct, the scaling code is correct.; Changing annotation_size alone is safe, since it is used consistently.; A mask's bounding box matching its saved box proves the polygon was recorded correctly.
b1
These hold right now and can be demonstrated: the overlay confirms the geometry, the argument order matches the unpacking inside the function, and the resize genuinely runs every frame for no benefit.
b2
Each of these confuses an outcome with a mechanism. The output is correct because two values coincide; the mask/box agreement tests only the steps after the polygon; and annotation_size is precisely the variable that is not used consistently.

48. Check: coincidental correctness

Check

The general skill, not the specific bug.

Check your understanding

Which of these is the strongest evidence that a piece of geometric code is correct rather than coincidentally correct?

  • A. It still produces correct output when you vary the parameter that the two duplicated values depend on. (correct)
  • B. It produces correct output on every tile in the dataset.
  • C. Its output passes an internal consistency check between two derived representations.
  • D. It has been code-reviewed by someone who understands the coordinate conventions.

Answer: A

Why: Coincidental correctness means two values happen to be equal. The only test that distinguishes it from real correctness is one where they are not equal — so you vary the parameter and check the output again. Everything else holds the coincidence fixed and therefore cannot see it.

Why B tempts people
Every tile in the dataset shares the same annotation window size, so this is one configuration tested many times. Volume does not substitute for variation.
Why C tempts people
Internal consistency between two things derived from the same source is a useful but strictly weaker check — both can be wrong together, which is exactly what happens here.
Why D tempts people
Review is valuable and catches plenty, but the question asks for evidence from the system's behaviour. A reviewer looking at the code has the same blind spot the author does unless they vary the parameter.

49. Source, Not Output

Section

Part 5

50. You are deleting the expensive part

Concept

create_training_targets takes your polygons, rasterises them, and returns arrays. The polygons themselves are never saved. They exist only inside the function call that consumes them.

for polygon in polygons:
    points = np.asarray(polygon, dtype=np.float32)
    points[:, 0] *= native_width / display_width
    points[:, 1] *= native_height / display_height
    points = np.rint(points).astype(np.int32)

    mask = np.zeros((native_height, native_width), dtype=np.uint8)
    cv2.fillPoly(mask, [points], 1)      # <- polygon becomes pixels here
    instance_masks.append(mask)

# `polygons` goes out of scope. The clicks are gone.
ArtifactCost to produceSaved?Recoverable?
polygon verticeshours of human attentionnoonly by redrawing
binary masksmicroseconds of fillPolyyestrivially, from polygons
bounding boxesmicroseconds of np.whereyestrivially, from masks

Read that table twice. You are persisting the two artifacts that cost nothing to recreate, and discarding the one that cost you an afternoon.

51. Break the claim

Counterexample

Discussion prompt

Here is a claim, stated as though it always holds:

"Masks are all I need. The polygons were just an input method."

Do one of two things: produce a concrete situation where that claim costs you real work, or say precisely what rules such a situation out. "It's just good practice" is not on the menu.

Hint: Think about every operation you might want to perform on an annotation that is not "feed it to this exact model at this exact resolution".

Answer:

  • One bad parcel. You spot a mis-drawn field in tile 40. With polygons you drag one vertex; with masks you redraw the tile.
  • Resolution change. You move to the 10 m S2 bands or crop to 64×64. Polygons re-rasterise exactly; masks resample and lose edges.
  • Another tool. CVAT, Labelme, Roboflow all import polygons. None of them import a directory of .npy masks.
  • Standard eval. pycocotools wants COCO polygons or RLE. Without them you write and debug your own mAP.
  • Class labels later. You decide to add the 18 PASTIS crop types. With polygons you relabel; with masks you re-annotate.
  • Boundary conventions. Adjacent parcels share edges. Whether a boundary pixel belongs to one parcel or both is a decision you can revisit with polygons and cannot with a rasterised mask.

Six of them, and every one costs hours you already spent once.

52. Source and derived, as a discipline

Concept

Source artifact — The thing a human produced or a decision created. Irreplaceable without redoing the work. Version-controlled, backed up, treated as precious.

Derived artifact — Anything a script can regenerate from source. Disposable by design, and safe to delete at any time — because deleting it is a re-run, not a loss.

You already apply this instinct elsewhere without thinking about it. You commit .py files and not __pycache__. You commit requirements.txt and not site-packages. Annotations are the same category of thing, and they are far more expensive than either.

In your projectSourceDerived
annotationpolygon vertices + classmasks, boxes, areas, RLE
dataset compositionsplits.jsonthe DataLoader's index
configurationthe config filethe instantiated model
resultsthe run logplots and tables in the README

53. Write COCO instead

Worked example

The format is unremarkable, which is the point — it is the one every tool already speaks. Note that geometry is stored per tile, not per observation, which folds in the Part 2 fix.

import json
from datetime import datetime, timezone

def polygons_to_coco(tile_polygons, tile_size=128, out_path='annotations.json'):
    """tile_polygons: {tile_id: [[(x, y), ...], ...]} in NATIVE coords."""
    images, annotations, ann_id = [], [], 1

    for image_id, (tile_id, polys) in enumerate(sorted(tile_polygons.items()), 1):
        images.append({'id': image_id, 'file_name': f'{tile_id}.npy',
                       'width': tile_size, 'height': tile_size})

        for poly in polys:
            xs = [float(x) for x, _ in poly]
            ys = [float(y) for _, y in poly]
            annotations.append({
                'id': ann_id, 'image_id': image_id, 'category_id': 1,
                'segmentation': [[c for xy in poly for c in map(float, xy)]],
                'bbox': [min(xs), min(ys), max(xs) - min(xs), max(ys) - min(ys)],
                'area': float(polygon_area(poly)), 'iscrowd': 0})
            ann_id += 1

    json.dump({'info': {'date_created': datetime.now(timezone.utc).isoformat()},
               'images': images, 'annotations': annotations,
               'categories': [{'id': 1, 'name': 'parcel'}]},
              open(out_path, 'w'), indent=2)
COCO fieldConventionWatch out for
segmentationflat [x1,y1,x2,y2,...]flat, not pairs — a common first-time error
bbox[x, y, width, height]NOT [x1,y1,x2,y2] like torchvision wants
category_idstarts at 10 is reserved; matches your labels = 1
areapolygon areathe true area, not the bbox area — mAP splits on it
iscrowd0 for normal instances1 means RLE region and changes how mAP scores it

The bbox row is the one that bites everyone exactly once. COCO stores width and height; torchvision's boxes tensor stores x2, y2. Converting in the wrong direction produces boxes that are plausible, wrong, and completely silent — the same failure signature as Part 4.

54. What the format buys you immediately

Concept

This is not a hypothetical future benefit. Three things become available the day you write the file.

Standard mAP
pycocotools gives you COCO-standard AP, AP50, AP75, and the small/medium/large breakdown — that last one matters enormously given your 4-38 px objects.
Any annotation tool
CVAT, Labelme and Roboflow all import COCO. If you ever want a second pair of eyes on your labels, this is what makes that possible.
Comparability
PASTIS reports panoptic metrics. Speaking a standard format is the first step toward putting your number next to a published one honestly.

The AP-by-size breakdown deserves emphasis. With a median object of 16 px, essentially every parcel you own falls into COCO's small bucket — so the aggregate AP you would otherwise report is dominated by a category most detection papers treat as the hard tail.

55. Order the pipeline

Ranking

Put in order

Put the rebuilt pipeline in the order it should run. One of these is the step that used to be missing entirely.

  1. Draw polygons in the annotator (human work — the only irreplaceable step).
  2. Save polygons to COCO JSON, keyed by tile.
  3. Run the QA overlay and the audit; fail loudly if anything drifts.
  4. Freeze a tile-level train/val split to splits.json.
  5. Rasterise polygons to masks on load, inside the Dataset.
  6. Train, and evaluate with pycocotools against the frozen split.

Why: Human work first, then persist it in a lossless standard format, then verify it — the QA step goes immediately after saving, because that is the earliest point at which a saved artifact exists to check. Only then do you fix the split, because a split over unverified data is a split you will have to redo. Rasterisation drops to load time, where it belongs, since it is derived. Training comes last and reads from artifacts every one of which has been checked.

56. Artifact to role

Matching

Match the pairs

Match each artifact to its role in the rebuilt pipeline. If you can do this without hesitating, the distinction has landed.

  • l1. annotations.json (COCO)
  • l2. splits.json
  • l3. in-memory masks
  • l4. DECISIONS.md
  • l5. annotation_qa.py output
  • r1. Source. Hours of human work; back it up.
  • r2. Source. A decision, not a computation — freeze and commit it.
  • r3. Derived. Rebuilt every epoch; never written to disk.
  • r4. Source. The reasoning, which no artifact can reconstruct.
  • r5. Derived. Regenerate on demand; read it, do not archive it.

Why: The test for 'source' is not importance — it is whether a script could regenerate it. Masks and QA reports are reproducible from other artifacts, so they are derived no matter how useful they are. The split is interesting precisely because it is cheap to compute and still source: it encodes a decision about which tiles are allowed to judge the model, and recomputing it would silently invalidate every comparison made before.

57. Check: source and derived

Check

One line of reasoning separates the right answer from three plausible ones.

Check your understanding

You store polygons in COCO and rasterise inside the Dataset. A colleague argues this is wasteful: rasterising 83 polygons every time a tile is loaded burns CPU that a cached .npy would save. What is the strongest response?

  • A. Cache the rasters if profiling shows the loader is the bottleneck — but keep polygons as the source of truth, so the cache is disposable and can be regenerated. (correct)
  • B. They are right; write the masks back to disk after the first epoch and read them thereafter.
  • C. Rasterisation is negligible compared to a forward pass, so the concern does not apply.
  • D. Keep polygons only, and never cache — caching derived artifacts is what created the duplication problem in Part 2.

Answer: A

Why: The source/derived distinction is about which artifact is authoritative, not about whether caching is allowed. Caching a derived artifact is fine and often correct; the discipline is that the cache must be reconstructible and safe to delete. That keeps the performance option open without giving up the ability to fix one parcel or change resolution later.

Why B tempts people
Not wrong about the performance, but it makes the raster authoritative by accident — once masks are the thing on disk that gets read, they become the thing people edit, and the polygons rot.
Why C tempts people
Probably true and still the weaker answer: it dismisses a legitimate concern rather than resolving it, and it stops being true if the loader is CPU-bound with many workers.
Why D tempts people
This over-corrects. Part 2's problem was not that a derived artifact was stored — it was that 61 identical copies were stored and then mistaken for independent labels.

58. Looking vs Learning

Section

Part 6

59. The normalisation you already wrote

Concept

This appears three times in your file, once per display path. For display it is exactly right.

rgb = img_array[index, [2, 1, 0], :, :].astype(np.float32)

channel_min = rgb.min(axis=(1, 2), keepdims=True)
channel_max = rgb.max(axis=(1, 2), keepdims=True)

rgb = (rgb - channel_min) / (channel_max - channel_min + 1e-8)
rgb = np.clip(rgb, 0, 1)
PurposePer-image min–maxVerdict
see a cloudy tile clearlystretches whatever is there to full rangecorrect
compare two tiles' brightnessdestroys the comparisonwrong
train a modelevery image gets its own scalewrong
run inference reproduciblyoutput depends on the image's own extremeswrong

It is going to be copied into the Dataset, because it is written, it is nearby, and it produces nice-looking arrays in [0, 1]. That is how this particular mistake almost always happens.

60. The inference-time surprise

Anomaly

Suppose you train with per-image min–max normalisation and the model works well on your validation tiles.

Predict first

Then you deploy it on a single new tile.

The same parcel, in the same field, on the same date, produces a different prediction depending on what else is in the tile — a bright cloud in one corner, say, or a dark water body.

Why? And why did validation never show this?

Correct: Per-image min-max makes every pixel's value depend on the minimum and maximum of the whole image, so an unrelated bright cloud rescales every other pixel. Validation missed it because your val tiles come from the same distribution and mostly contain the same kind of content.

Fixed per-band statistics remove the coupling entirely: pixel value in, same number out, every time, regardless of what shares the frame.

Why: Normalisation is part of the model's input contract. Per-image statistics make that contract depend on the image itself, which means the same physical surface maps to different network inputs depending on its neighbours. This is not merely a robustness issue: it makes single-tile inference non-reproducible in a way that is very hard to debug from the outside, because nothing about the parcel changed.

61. Fixed statistics, computed once

Worked example

PASTIS ships NORM_S2_patch.json for exactly this. If you were not using PASTIS you would compute the same thing yourself — over the training split only, which is a Part 2 idea reappearing here.

import json, numpy as np

def compute_norm_stats(train_tile_ids, img_dir, n_bands=10):
    """Per-band mean/std over the TRAINING tiles only.
    Using val tiles here leaks their statistics into training."""
    total = np.zeros(n_bands, np.float64)
    total_sq = np.zeros(n_bands, np.float64)
    count = 0

    for tile in train_tile_ids:
        arr = np.load(img_dir / f'{tile}.npy').astype(np.float64)
        flat = arr.transpose(1, 0, 2, 3).reshape(n_bands, -1)
        total    += flat.sum(axis=1)
        total_sq += (flat ** 2).sum(axis=1)
        count    += flat.shape[1]

    mean = total / count
    std = np.sqrt(total_sq / count - mean ** 2)
    json.dump({'mean': mean.tolist(), 'std': std.tolist()},
              open('norm_stats.json', 'w'), indent=2)
    return mean, std

# in the Dataset, identical at train and inference:
#     image = (image - mean[:, None, None]) / std[:, None, None]
DecisionReason
training tiles onlyval statistics are val information; using them is leakage
float64 accumulatorssumming squares over ~10^8 pixels overflows float32 precision
saved to JSONa source artifact — inference must use the identical numbers
per band, not per imagethe transform no longer depends on the image's own content
all bands, not just RGByou display 3; the model should see all 10

62. One rule, two places

Intuition

The rule is short: any transform that depends on the data must be fitted on training data only, saved, and reapplied unchanged at inference.

Normalisation statistics are the obvious case. The same rule covers class weights, anchor sizes derived from box statistics, and any threshold you tune. If a number came from looking at data, it belongs to the split it came from.

Notice that this is the Part 2 idea again, wearing different clothes. Leakage is not only about which samples go where — it is about which information crosses the boundary.

63. Pseudo-Labels That Earn Trust

Section

Part 7

64. Noisy Student is a classification method

Concept

Worth saying plainly, because the adaptation is not trivial and the difference is where all the difficulty lives.

Classification (Noisy Student)Instance segmentation (yours)
teacher outputone label + one confidenceN boxes, N masks, N scores, N unknown
"confident"a single softmax valueper-instance, and badly calibrated
a wrong pseudo-labelone wrong classa phantom parcel, or a missed one
missing outputimpossible — always one labela missed parcel becomes a false negative, taught as truth

That last row is the one that makes detection self-training genuinely harder. In classification every sample gets a label. In detection, a parcel the teacher fails to find is not merely absent — it is actively taught to the student as background.

The literature to know by name: STAC, Unbiased Teacher, Soft Teacher, and — for instance segmentation specifically — Polite Teacher. You do not need to implement them. You need to be able to say what problem each was solving when someone asks.

65. What would you need to set a threshold?

Missing information

Discussion prompt

You asked how to set the confidence threshold for accepting pseudo-labels.

Before answering: what information would you need in order to choose one well? List everything. Then mark which items you can actually obtain.

Hint: A threshold trades false positives against false negatives. What would you need to know to price that trade?

Answer:

You would needCan you get it?
precision/recall vs score on unlabelled dataNo — that needs the labels you do not have
whether the score is calibratedRarely; detector scores usually are not
the cost of a phantom vs a missed parcelYes — but it is a judgement, not a measurement
how the threshold interacts with round numberOnly by running the loop and watching

Which is the honest answer to your original question: you cannot choose a good global score threshold from information you have. Everyone picks 0.8 or 0.9 because everyone else does.

So the interesting move is to stop needing one.

66. Your data has a signal a score does not

Concept

The duplication from Part 2 — the thing that was a liability for splitting — is an asset here, and it is specific to your problem in a way that is worth exploiting.

Temporal support — The number of observations, out of the tile's time series, in which a matched detection appears. A property of the world rather than of the model's calibration — which is precisely what a raw confidence score is not.

Two things follow, and they are the reason this is worth building. Your "noise" for the student becomes genuine sensor and phenology variation rather than synthetic colour jitter. And your acceptance criterion becomes an agreement count rather than a guessed number.

67. The temporal consistency filter

Worked example

Run the teacher across all dates, match detections spatially, and keep what recurs.

from collections import defaultdict

def temporal_pseudo_labels(model, tile, n_obs, min_support=0.5, iou_thr=0.6):
    """Keep instances the teacher re-finds across the time series."""
    clusters = []                       # each: {'boxes': [...], 'obs': set()}

    for obs in range(n_obs):
        pred = model(load_observation(tile, obs))[0]
        for box, score in zip(pred['boxes'], pred['scores']):
            if score < 0.3:             # a floor, not a decision
                continue
            hit = next((c for c in clusters
                        if iou(box, mean_box(c)) > iou_thr), None)
            if hit is None:
                clusters.append({'boxes': [box], 'obs': {obs}})
            else:
                hit['boxes'].append(box)
                hit['obs'].add(obs)

    # support = fraction of observations in which this instance was found
    keep = [c for c in clusters if len(c['obs']) / n_obs >= min_support]
    return [mean_box(c) for c in keep]      # averaging also denoises
DetectionFound inSupportVerdict
cluster A54 / 61 obs0.89keep — stable across the season
cluster B38 / 61 obs0.62keep — likely obscured on cloudy dates
cluster C9 / 61 obs0.15drop — appears and vanishes
cluster D2 / 61 obs0.03drop — almost certainly a hallucination

The score < 0.3 floor is deliberately loose. It exists only to keep the clustering tractable — the real decision is made by support, which is the whole point. And averaging the boxes within a cluster gives you a better box than any single observation produced, because independent errors partially cancel.

68. How support separates signal from noise

Scale up

Step through it

Watch what happens to a real parcel and to a hallucination as you add more observations. At what point do they become distinguishable?

  1. 1 obs: indistinguishable
  2. 3 obs: a gap opens
  3. 10 obs: clearly separated
  4. 61 obs: threshold-insensitive

This is the argument in one picture. You cannot separate the two by score at any number of observations. You can separate them by support at three. Threshold-insensitivity is the goal — when the answer stops depending on where you drew the line, you have found the right quantity to measure.

69. Choosing the support threshold

Trade off

Comparison matrix

Fill in the consequences. Notice how much less sensitive this choice is than a score cutoff — that insensitivity is the feature.

min_supportKeepsRiskWhen to use it
0.25anything seen in 15+ of 61admits noise from persistent hallucinationsvery cloudy series where real parcels are often hidden
0.50seen in ~half the seriesbalancedthe sensible default to start from
0.75seen in 46+ of 61drops parcels obscured by cloudround 1, when purity matters more than volume
0.95near-universal detections onlyvery few pseudo-labels survivewhen you have evidence the loop is drifting

A refinement worth mentioning once you have this working: cloud cover is not random across the series, so weighting each observation by a cloud score before computing support would be more principled than treating all 61 dates as equal evidence. That is a good second iteration, not a first one.

70. Trap: letting pseudo-labels touch your validation set

Trap

The trap

You have 100 manually labelled tiles. Round 1 works, so you pseudo-label the unlabelled pool and add the confident ones to training. Round 2 scores higher. You report the improvement.

Evaluate round 2 against a set the teacher has influenced

Why: If any pseudo-labelled tile reached your validation set — or if you re-split after adding them — you are measuring agreement between the student and its own teacher, which rises every round regardless of whether the model got better.

The number goes up on every round, forever, and it means nothing.

The fix

Ring-fence a manually labelled validation set before round 1 and never let a pseudo-label near it.

Freeze val once, evaluate everything against it

Why: It is the only fixed point in a loop where everything else is changing. Write those tile ids into splits.json and treat them as untouchable — no pseudo-labels, no re-splitting, no exceptions.

A self-training loop without a frozen human-labelled reference is not an experiment. It is a feedback circuit with a score attached.

71. Commit before you run it

Commit first

Predict first

You run three rounds of self-training with the temporal filter, and you track mean predicted instances per tile on the frozen validation set.

Round 1 gives 78 instances per tile. Round 2 gives 71. Round 3 gives 64. Your ground truth averages 83.

How confident are you that round 4 will improve things — and what is this trend telling you? Commit to an answer before reading on.

Correct: Not confident at all. The count is drifting steadily downward and away from the ground truth, which is the signature of confirmation bias: parcels the teacher misses become background in the student's training data, so each round the student learns to find fewer.

Stop at round 3, keep the round-2 weights, and look at what is being lost. A drifting count is a much earlier warning than a falling mAP.

Why: This is the specific way detection self-training collapses, and it is why mean instance count is the diagnostic to watch. A missed detection is not a neutral absence — it is taught to the student as confirmed background. The student then misses it more reliably, its own pseudo-labels contain it less often, and the loop reinforces itself. Monotonic drift in either direction means stop; stability near the ground-truth count means the loop is healthy.

72. Guardrails, stated once

Concept

Self-training fails quietly. Five rules keep it honest, and none of them is expensive.

  1. Freeze a human-labelled validation set before round 1. It is never pseudo-labelled and never re-split.
  2. Cap the loop at two or three rounds. Gains flatten fast; bias does not.
  3. Track mean instances per tile every round. Monotonic drift means stop.
  4. Keep every round's weights. The best model is often not the last one.
  5. Log the config-to-metric mapping. Support threshold, round, mAP, mean count — one row per run.

The fifth is the one that turns this from a script into an experiment. Without a log you cannot say which change caused which effect, and a self-training loop has too many interacting knobs to reconstruct from memory.

73. Check: why temporal support beats a score

Check

The mechanism, not the slogan.

Check your understanding

What is the fundamental reason temporal support is a better acceptance criterion than the detector's confidence score, for this dataset specifically?

  • A. Support is evidence from repeated independent observations of the same fixed geometry, whereas the score is a single, typically uncalibrated model output. (correct)
  • B. Support is bounded between 0 and 1, which makes the threshold easier to reason about than an unbounded score.
  • C. Support is computed after non-maximum suppression, so it is unaffected by duplicate detections.
  • D. Confidence scores are unavailable from Mask R-CNN at inference, so support is the only usable signal.

Answer: A

Why: The parcel geometry is genuinely constant across the time series while the imagery genuinely varies, so each observation is close to an independent test of the same hypothesis. Agreement across many such tests is far stronger evidence than one model's self-reported confidence — especially since detector scores are known to be poorly calibrated. It also means the acceptance decision stops depending on where you put a threshold, which is the practical payoff.

Why B tempts people
Detector scores are already bounded in [0, 1]. Boundedness is not the difference between the two signals.
Why C tempts people
Duplicate suppression is a detail of the clustering step, not the reason support is more informative. You still run NMS either way.
Why D tempts people
Mask R-CNN returns a scores tensor at inference — the code in this deck uses it as a loose floor. The argument is about which signal to trust, not which is available.

74. Where does the idea stop working?

Edge cases

Parameter explorer

Temporal consistency depends on geometry being static across the series. Push each parameter toward its extreme and find where the assumption breaks.

  • observations per tile — from 2 to 61
  • cloud-free fraction — from 0.1 to 1
  • days spanned by the series — from 7 to 730

Being able to say where your own idea stops working is worth more in an interview than the idea itself. It is the difference between having read about a technique and having understood it.

75. Habits That Compound

Section

Part 8

76. Write down what you decided and why

Concept

Every part of this deck ended in a decision. A reader of your repo — an interviewer, a collaborator, you in three months — cannot tell your decisions from your accidents unless you say.

DecisionThe alternativeWhy you chose it
one class (parcel)18 PASTIS crop typesgeometry first; classification is a separable second problem
single-frame Mask R-CNNU-TAE + PaPsa deliberately simpler baseline, with temporal work planned
min_size=256the default 8002× upscale on 128 px tiles; 39× compute is not justified
tile-level splitrandom over observationsthe tile is the unit of independence
temporal support filtera fixed score thresholdsupport is evidence; a score is self-report

That table is most of a DECISIONS.md. It took five minutes to write and it changes how the whole project reads: not "here is a Mask R-CNN tutorial I followed" but "here are five decisions I made and can defend."

77. Habits and their impostors

Sorting

Sort into buckets

Each of these describes something you might do. Sort them by whether they actually produce evidence, or only produce the feeling of rigour.

Produces evidence
Print the object-size distribution before configuring anchors.; Overlay saved annotations onto the source image and look.; Log config-to-metric for every run in one table.; Assert that train and val tile ids are disjoint.
Feels rigorous, proves nothing
Use a fixed random seed everywhere.; Report validation mAP to three decimal places.; Follow the reference implementation's hyperparameters exactly.
b1
Each one produces an artifact that could contradict you: a distribution that disagrees with your config, an overlay that shows drift, a log that shows a change made things worse, an assertion that fails. A check that cannot fail is not a check.
b2
A seed makes a result reproducible without making it correct. Three decimals on a leaked score is precision about a fiction. Copying hyperparameters tuned on COCO — 800 px images, ~7 objects each — onto 128 px tiles with 83 objects is inheritance, not reasoning.

78. Measure, then configure — the general form

Concept

The same move appeared in four different places in this deck. Worth naming so you can reuse it deliberately.

Instead ofMeasureThen set
default anchor sizesyour box-side distributionanchors that span it
detections_per_img=100instances per imagea cap above your maximum
a guessed score thresholdtemporal supporta support fraction
ImageNet normalisationper-band statistics on trainfixed mean and std
"80/20 seems standard"your independence structurea split over the right unit

Defaults are someone else's measurements. They are a reasonable starting point precisely to the extent that your data resembles theirs — and 128×128 tiles with 83 tessellating 16-pixel objects do not resemble COCO in any respect that matters.

79. Teach it to someone else

Explain it

Discussion prompt

Pick the finding from this deck that surprised you most.

Explain it to an imaginary teammate who is competent but has never touched satellite data — in five sentences or fewer, with no jargon they would have to look up. Then answer their obvious follow-up: "so how would I have caught that myself?"

Hint: If you cannot explain the mechanism without saying "leakage" or "pseudo-label", you are relying on the word to carry meaning you have not yet supplied.

Answer:

The second half is the real test. A finding you can only recognise is a fact; a finding you can teach someone to detect is a skill — and only the second one transfers to the next project.

80. What this project is actually a portfolio piece for

Real world

Be honest about what reads as strong to someone hiring for robotics or perception work.

Reads asBecause
a tutorial follow-along"I trained Mask R-CNN on a public dataset"
a competent engineer"I found a leak in my own eval and fixed the split"
someone worth interviewing"I used temporal consistency as a pseudo-label filter and ablated it against a score threshold"

Discussion prompt

Write the two-sentence version of this project that you would put at the top of the README.

Constraint: neither sentence may name an architecture. If removing the model name empties the description, that is the finding.

Hint: Lead with the problem and the thing you did that someone else would not have.

Answer:

Something in this direction: "Hand-annotated parcel instance segmentation on Sentinel-2 time series, with a tile-level split and a round-trip annotation QA harness. Pseudo-labels are accepted by temporal consistency across the series rather than by a confidence threshold, which removes the tuned cutoff and improves precision by X% over a score-threshold baseline."

Note what is doing the work there: a decision, a check, and a measured comparison. The architecture is an implementation detail and belongs further down.

81. The order to do this in

Concept

Sequenced so that nothing is built on something unverified. Roughly one session each.

  1. Run the QA audit on your full annotation directory. Fix whatever it reports before anything else.
  2. Add polygon saving, export to COCO, and delete the duplicate per-observation writes.
  3. Freeze splits.json at the tile level, and commit it.
  4. Measure and configure — box sizes after the transform, then anchors, min_size, detection cap.
  5. Train the baseline and record it honestly. This number is the thing everything later must beat.
  6. Add the temporal filter, run two rounds, and ablate it against a fixed score threshold.

Steps 1 to 3 produce no model and no metric, and they are the ones that determine whether anything after them means anything. That is normal, and being willing to spend three sessions there is itself a signal about how you work.

82. Draw the whole thing

Connect it up

Draw it

One page. Put these six on it — round-trip QA · tile-level splitting · measure-then-configure · source vs derived · fixed normalisation · temporal support — and draw an arrow wherever one of them is what makes another possible. Label every arrow with the reason.

Two of them are the same underlying idea wearing different clothes. Find that pair and mark it.

The pair worth finding: tile-level splitting and fixed normalisation are both the rule that information must not cross the train/val boundary. One is about which samples, the other about which statistics — but it is one principle, and seeing that is worth more than memorising either.

83. Before you close this

Exit ticket

Predict first

Two answers, written down, before you open your editor:

1. Which finding from this deck would have cost you the most time had you not found it until after training?
2. What is the first command you will run tomorrow?

Correct: Most people say the tile-level leak for the first — it would have invalidated every number produced up to that point, including the comparisons used to choose everything else. The second should be the QA audit over your full annotation directory, because every other fix depends on knowing the current state of the data.

python annotation_qa.py audit --ann <your annotations> --masks <your masks>

Send me the output before the next session and we will start from what it says.

Why: The first question is about consequence, and the leak wins because it corrupts the measurement apparatus rather than the model: with a leaked validation set, every subsequent decision is made on the basis of a number that does not mean what it claims. The second is about sequencing — the audit is first because it is the only step that tells you which of the other problems you actually have, and how bad they are on your real data rather than on the five files you sent.

84. What you can now do

Recap

You started with working annotation code and a reasonable plan. What changed is the layer above both.

ProblemFixHow you will know it worked
61 copies, 1 labeltile-level splits.jsondisjoint tile ids; val mAP drops and becomes real
objects vs defaultsmin_size + anchors from measurementsmall-object AP rises; step time falls
invisible coordinate bugone source of truth + overlayoutput survives changing annotation_size
polygons discardedCOCO as sourceyou can fix one parcel without redrawing a tile
per-image normalisationfixed per-band statssame tile, same prediction, every time
guessed thresholdtemporal supportresult stops depending on the cutoff

One more thing worth saying: none of this was a criticism of your annotator. It works, and the hard part — 83 careful polygons — was already done. Everything here was about making that work provable, which is the difference between a project that runs and a project someone else can trust.

Sources

  1. PASTIS: Panoptic Segmentation of Satellite Image Time Series — Garnot, V. S. F., & Landrieu, L. (2021). Panoptic Segmentation of Satellite Image Time Series with Convolutional Temporal Attention Networks. ICCV 2021.
  2. torchvision.models.detection — Mask R-CNN reference implementation
  3. COCO dataset format specification
  4. Self-training with Noisy Student improves ImageNet classification — Xie, Q., Luong, M.-T., Hovy, E., & Le, Q. V. (2020). CVPR 2020.
  5. End-to-End Semi-Supervised Object Detection with Soft Teacher — Xu, M., Zhang, Z., Hu, H., et al. (2021). ICCV 2021.
  6. Polite Teacher: Semi-Supervised Instance Segmentation with Pseudo-Labels — Filipiak, D., Zapala, A., Tempczyk, P., et al. (2023). CVPR Workshops 2023.

Want this taught 1-on-1? Alexander tutors Computer Vision — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108