Lesson 106: Named Entity Recognition — BERT + CRF

USAAIO Lesson 106, from Phase 3. It covers the BIO tagging scheme for named-entity recognition, a BERT encoder feeding a per-token linear head, and a CRF transition matrix that enforces valid label sequences. Viterbi decoding is traced on a three-token toy example, with the good path scoring 11.70 and the bad path 4.80, and the lesson contrasts token-level with entity-level F1 using a concrete span-mismatch example. The shapes of TinyBertNER - 81,991 parameters at hidden size 64 - were verified with torch 2.7.1+cpu. The lesson runs to 29 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Named Entity Recognition BERT + CRF

Title

USAAIO · Lesson 106 · Phase 3

Tag each token with its entity type. BIO scheme, BERT encoder, CRF for globally consistent label sequences, Viterbi decoding — every number verified in PyTorch.

2. By the end of this lesson you can

Objectives

  1. Apply the BIO tagging scheme to label a sentence and identify which transitions are valid
  2. Trace the BERT + linear head pipeline: token IDs → encoder hiddens (B,T,H) → logits (B,T,L) → greedy labels
  3. Explain what a CRF transition matrix captures and why it prevents invalid label sequences like O → I-PER
  4. Execute Viterbi decoding on a toy 3-token sequence and identify the globally best path
  5. Distinguish token-level F1 from entity-level F1 and give a case where one succeeds and the other fails

3. What survived from BLEU, ROUGE, and BERTScore?

Warm-up

Discussion prompt

Before we open Lesson 106: Named Entity Recognition — BERT + CRF: without looking back, what was the main idea of BLEU, ROUGE, and BERTScore, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

BLEU n-gram precision with brevity penalty, ROUGE recall-oriented n-gram overlap and ROUGE-L via LCS, BERTScore via contextual embedding cosine similarity, and when each metric fails. All computations implemented from scratch in Python and verified with real execution — p_1=5/6, p_2=3/5, BLEU-4=0.4209 on MT example, ROUGE-L F1=0.8333 on LCS example.

4. BIO Tagging — the sequence labeling contract

Section

Part 1 of 4

5. NER as token classification

Concept

Named Entity Recognition assigns a label to every token. NER is not a span-extraction task — it is a per-token classification problem, and the label set encodes span boundaries via the BIO scheme.

BIO scheme — Each token gets label O (outside), B-TYPE (beginning of an entity span of that type), or I-TYPE (inside/continuation of a span). A span is exactly the maximal run B followed by zero or more consecutive I of the same type.

TokenLabelSpan role
JohnB-PERstarts PER span
SmithI-PERcontinues PER span
worksOoutside
atOoutside
GoogleB-ORGstarts ORG span
inOoutside
NewB-LOCstarts LOC span
YorkI-LOCcontinues LOC span

6. Fill in: Label for NER as token classification

Comparison

Comparison matrix

From NER as token classification: refill the Label column from what you know. The rest of the table is as it appeared.

TokenLabelSpan role
JohnB-PERstarts PER span
SmithI-PERcontinues PER span
worksOoutside
atOoutside
GoogleB-ORGstarts ORG span
inOoutside
NewB-LOCstarts LOC span
YorkI-LOCcontinues LOC span

7. The 7-label BIO vocabulary

Concept

For 3 entity types (PER, ORG, LOC) the full label set is size 7: O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC. In general, L = 2K + 1 for K entity types.

\[ L = 2K + 1 \quad (K \text{ entity types}) \]

Label IDLabelMeaning
0OOutside any entity
1B-PERBegin PER span
2I-PERInside PER span
3B-ORGBegin ORG span
4I-ORGInside ORG span
5B-LOCBegin LOC span
6I-LOCInside LOC span

8. What each one costs: The 7-label BIO vocabulary

Trade off

Comparison matrix

From The 7-label BIO vocabulary: every row here is a choice with a cost. Fill the Label column, then say which row you would actually pick and what you give up for it.

Label IDLabelMeaning
0OOutside any entity
1B-PERBegin PER span
2I-PERInside PER span
3B-ORGBegin ORG span
4I-ORGInside ORG span
5B-LOCBegin LOC span
6I-LOCInside LOC span

9. Something is wrong here: I-label at the start of a span

Anomaly

Predict first

A student writes this, and it looks reasonable:

Predicted sequence for "New York": I-LOC I-LOC

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.

Correct sequence for "New York": B-LOC I-LOC

Why: The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.

10. Trap: I-label at the start of a span

Trap

The trap

Predicted sequence for "New York": I-LOC I-LOC

Accept I-LOC as the first label of the span

Why: The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.

Entity-level evaluators treat this as zero entities detected — no span can be extracted without a B-label.

The fix

Correct sequence for "New York": B-LOC I-LOC

Require B-LOC on the first token of every new span

Why: BIO contract: I-TYPE is only valid immediately after B-TYPE or I-TYPE of the same type. The CRF transition matrix learns to assign -inf score to O->I and B-X->I-Y (type mismatch).

With a CRF layer, invalid transitions are globally penalized even if the token's emission score strongly favors I-LOC.

11. Break it on purpose: I-label at the start of a span

Break the constraint

Discussion prompt

The rule this trap just fixed:

BIO contract: I-TYPE is only valid immediately after B-TYPE or I-TYPE of the same type. The CRF transition matrix learns to assign -inf score to O->I and B-X->I-Y (type mismatch).

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.

12. BERT + Linear Head — token-level logits

Section

Part 2 of 4

13. BERT encoder produces per-token hidden states

Concept

Unlike classification tasks that read only the [CLS] token (Lesson 88), NER reads every token's final hidden state and passes it through a shared linear head to produce label logits.

\[ \text{logits} = W \cdot h_t + b \quad \forall t \in \{0,\ldots,T-1\} \]

Result: a (B, T, L) tensor of unnormalized scores — one length-L vector per token per example in the batch.

14. Break it if you can: BERT encoder produces per-token hidden states

Counterexample

Discussion prompt

Unlike classification tasks that read only the [CLS] token (Lesson 88), NER reads every token's final hidden state and passes it through a shared linear head to produce label logits.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Result: a (B, T, L) tensor of unnormalized scores — one length-L vector per token per example in the batch.

15. Predict the next row: BERT + linear head: shape trace

Pattern

Predict first

The table runs: token IDs | (1, 8) | batch=1, seq_len=8 · embed + pos | (1, 8, 64) | token + positional embeddings · encoder output | (1, 8, 64) | contextualized hidden states · logits (head) | (1, 8, 7) | score per label per token

In BERT + linear head: shape trace, given the rows so far: what is the next one — the row where Tensor is argmax preds?

Correct: argmax preds | (1, 8) | greedy label IDs (0..6)

TensorShapeDescription
token IDs(1, 8)batch=1, seq_len=8
embed + pos(1, 8, 64)token + positional embeddings
encoder output(1, 8, 64)contextualized hidden states
logits (head)(1, 8, 7)score per label per token
argmax preds(1, 8)greedy label IDs (0..6)

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Embedding(100,64)=6400 + Embedding(128,64)=8192 + TransformerEncoder(2 layers, H=64, FFN=128, heads=4) + Linear(64,7)=455.

16. BERT + linear head: shape trace

Worked example

import torch, torch.nn as nn

class TinyBertNER(nn.Module):
    def __init__(self, vocab=100, H=64, heads=4, layers=2, L=7):
        super().__init__()
        self.embed = nn.Embedding(vocab, H, padding_idx=0)
        self.pos   = nn.Embedding(128, H)
        enc = nn.TransformerEncoderLayer(
            d_model=H, nhead=heads,
            dim_feedforward=128, batch_first=True, dropout=0.0)
        self.encoder = nn.TransformerEncoder(enc, num_layers=layers)
        self.head    = nn.Linear(H, L)

    def forward(self, ids):
        B, T = ids.shape
        pos  = torch.arange(T).unsqueeze(0).expand(B, -1)
        x    = self.embed(ids) + self.pos(pos)
        x    = self.encoder(x)       # (B, T, H)
        return self.head(x)          # (B, T, L)

torch.manual_seed(42)
model = TinyBertNER()
print(sum(p.numel() for p in model.parameters()))  # 81991
tokens = torch.tensor([[5, 12, 8, 23, 45, 7, 61, 19]])
logits = model(tokens)
print(logits.shape)                  # torch.Size([1, 8, 7])
preds  = logits.argmax(-1).squeeze()
print(preds.tolist())                # per-token label IDs
TensorShapeDescription
token IDs(1, 8)batch=1, seq_len=8
embed + pos(1, 8, 64)token + positional embeddings
encoder output(1, 8, 64)contextualized hidden states
logits (head)(1, 8, 7)score per label per token
argmax preds(1, 8)greedy label IDs (0..6)

Count parameters: 81,991 total

Why: Embedding(100,64)=6400 + Embedding(128,64)=8192 + TransformerEncoder(2 layers, H=64, FFN=128, heads=4) + Linear(64,7)=455. Verified by running sum(p.numel() for p in model.parameters()).

Greedy decoding: argmax over the 7-class dimension

Why: With an untrained model, predictions are essentially random. After fine-tuning on labeled data, argmax often gives valid BIO sequences — but not always, motivating the CRF layer.

17. Which is which, by Shape

Discrimination

Sort into buckets

Sort these by Shape, from memory, without looking back at BERT + linear head: shape trace. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(1, 8)
token IDs; argmax preds
(1, 8, 64)
embed + pos; encoder output
(1, 8, 7)
logits (head)
g1
Shape is "(1, 8)" for token IDs, argmax preds — that is what the table on "BERT + linear head: shape trace" records, and it is the single property separating this group from the rest.
g2
Shape is "(1, 8, 64)" for embed + pos, encoder output — that is what the table on "BERT + linear head: shape trace" records, and it is the single property separating this group from the rest.
g3
Shape is "(1, 8, 7)" for logits (head) — that is what the table on "BERT + linear head: shape trace" records, and it is the single property separating this group from the rest.

18. CRF Layer — globally consistent label sequences

Section

Part 3 of 4

19. What the CRF adds: a transition matrix

Concept

A Conditional Random Field (CRF) adds a learnable transition matrix A of shape (L, L). Entry A[i,j] is the score of transitioning from label i to label j. The model scores a full label sequence, not each token independently.

\[ \text{score}(\mathbf{y}) = \sum_{t=0}^{T-1} \underbrace{e_t[y_t]}_{\text{emission}} + \sum_{t=1}^{T-1} \underbrace{A[y_{t-1}, y_t]}_{\text{transition}} \]

Training maximizes the log-likelihood of the gold sequence relative to all sequences via the forward algorithm. Inference finds the highest-scoring sequence via Viterbi decoding.

20. By analogy: What the CRF adds: a transition matrix

Analogy

Discussion prompt

Explain What the CRF adds: a transition matrix by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Training maximizes the log-likelihood of the gold sequence relative to all sequences via the forward algorithm. Inference finds the highest-scoring sequence via Viterbi decoding.

21. Valid vs invalid BIO transitions

Concept

The CRF does not hard-code BIO rules — it learns them from data. After training, A[O, I-PER] becomes very negative (never observed), while A[B-PER, I-PER] becomes very positive.

TransitionValid?Reason
O -> B-PERYESStart a new PER entity
B-PER -> I-PERYESContinue the PER entity
B-PER -> I-ORGNOType mismatch inside span
I-ORG -> I-PERNOType mismatch inside span
O -> I-PERNOI-label without preceding B-label
I-LOC -> B-PERYESEnd LOC, start new PER

The key advantage over greedy token-by-token decoding: even if the emission score strongly favors I-ORG after an O, the learned transition penalty A[O, I-ORG] can override it globally.

22. Which is which, by Valid?

Discrimination

Sort into buckets

Sort these by Valid?, from memory, without looking back at Valid vs invalid BIO transitions. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

YES
O -> B-PER; B-PER -> I-PER; I-LOC -> B-PER
NO
B-PER -> I-ORG; I-ORG -> I-PER; O -> I-PER
g1
Valid? is "YES" for O -> B-PER, B-PER -> I-PER, I-LOC -> B-PER — that is what the table on "Valid vs invalid BIO transitions" records, and it is the single property separating this group from the rest.
g2
Valid? is "NO" for B-PER -> I-ORG, I-ORG -> I-PER, O -> I-PER — that is what the table on "Valid vs invalid BIO transitions" records, and it is the single property separating this group from the rest.

23. What has to happen first: Viterbi decoding — 3-token trace

Ranking

Put in order

Put the moves of Viterbi decoding — 3-token trace into the order they have to happen.

  1. Initialize: V[0] = emission scores at t=0 = [0.10, 3.00, 0.50, 0.20]
  2. t=1: V[j] = max_i(V_prev[i] + A[i,j]) + emit[1,j]
  3. t=2: best path score for O = 11.70; backtrack gives [B-PER, I-PER, O]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. No prior label exists, so only emission scores contribute.

24. Viterbi decoding — 3-token trace

Worked example

import torch
# 3 tokens, 4 labels: O=0 B-PER=1 I-PER=2 B-ORG=3
# Emission scores (unnormalized log probs per token per label)
emit = torch.tensor([
    [0.10, 3.00, 0.50, 0.20],  # tok0: high B-PER
    [0.20, 0.30, 3.50, 0.10],  # tok1: high I-PER
    [3.20, 0.10, 0.20, 0.30],  # tok2: high O
])
# Transition matrix A[from, to]
trans = torch.tensor([
    [ 1.0,  0.5, -1.0,  0.5],  # from O
    [ 0.5,  0.5,  1.5, -2.0],  # from B-PER
    [ 0.5,  0.5,  0.5, -2.0],  # from I-PER
    [ 0.5,  0.5, -2.0,  0.5],  # from B-ORG
])
# Viterbi forward pass
V = emit[0].clone()          # t=0: no transition yet
bptr = torch.zeros(3, 4, dtype=torch.long)
for t in range(1, 3):
    scores = V.unsqueeze(1) + trans   # (4,4)
    best   = scores.max(0)            # max over 'from' dim
    bptr[t] = best.indices
    V = best.values + emit[t]
# Backtrack
path, cur = [], V.argmax().item()
path.append(cur)
for t in range(2, 0, -1):
    cur = bptr[t, cur].item(); path.append(cur)
path.reverse()
print(path)   # [1, 2, 0]  => B-PER, I-PER, O
tOB-PERI-PERB-ORGChosen
0 (Alice)0.103.000.500.20B-PER
1 (Smith)3.703.808.001.10I-PER
2 (left)11.708.608.706.30O

Initialize: V[0] = emission scores at t=0 = [0.10, 3.00, 0.50, 0.20]

Why: No prior label exists, so only emission scores contribute. B-PER leads with 3.00.

t=1: V[j] = max_i(V_prev[i] + A[i,j]) + emit[1,j]

Why: For I-PER: best predecessor is B-PER (3.00 + A[B-PER,I-PER]=1.5 = 4.50) + emit[1,I-PER]=3.50 = 8.00. This beats O path: 3.70.

t=2: best path score for O = 11.70; backtrack gives [B-PER, I-PER, O]

Why: The CRF scores the good path (emit=9.70, trans=2.00, total=11.70) against the bad path B-PER,B-ORG,O (emit=6.30, trans=-1.50, total=4.80). Good path wins by 6.90 points.

25. Inspect it line by line: Viterbi decoding — 3-token trace

Error analysis

Annotate

Walk the callouts on Viterbi decoding — 3-token trace. Each one is a place this is easy to get subtly wrong.

  • No prior label exists, so only emission scores contribute. B-PER leads with 3.00.
  • For I-PER: best predecessor is B-PER (3.00 + A[B-PER,I-PER]=1.5 = 4.50) + emit[1,I-PER]=3.50 = 8.00. This beats O path: 3.70.
  • The CRF scores the good path (emit=9.70, trans=2.00, total=11.70) against the bad path B-PER,B-ORG,O (emit=6.30, trans=-1.50, total=4.80). Good path wins by 6.90 points.

26. Something is wrong here: greedy decoding vs Viterbi

Anomaly

Predict first

A student writes this, and it looks reasonable:

At each step, take argmax of the current token's emission scores only.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Greedy looks at one token at a time — no transition score considered.

Viterbi finds the globally best sequence by propagating accumulated scores forward and backtracking.

Why: Greedy looks at one token at a time — no transition score considered.

27. Trap: greedy decoding vs Viterbi

Trap

The trap

At each step, take argmax of the current token's emission scores only.

t=1: argmax of emit[1] = I-PER (3.50). Accept it immediately.

Why: Greedy looks at one token at a time — no transition score considered.

This accidentally produces correct output here, but fails whenever a locally high emission conflicts with global structure (e.g., I-ORG after O).

The fix

Viterbi finds the globally best sequence by propagating accumulated scores forward and backtracking.

t=1: V[I-PER] = max_i(V_prev[i] + A[i, I-PER]) + emit[1, I-PER] = 8.00

Why: The predecessor B-PER contributes a +1.5 transition bonus; greedy would never see this. Viterbi guarantees the path maximizing the joint sequence score.

CRF + Viterbi runs in O(T * L^2) — practical for the typical label set sizes in NER.

28. Which of these survive contact with Lesson 106: Named Entity Recognition — BERT…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
For 3 entity types (PER, ORG, LOC) the full label set is size 7: O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC. In general, L = 2K + 1 for K entity types.; Result: a (B, T, L) tensor of unnormalized scores — one length-L vector per token per example in the batch.; NER has two distinct evaluation metrics. They can diverge dramatically — a system that gets every boundary wrong can still report high token-level F1.
Breaks
Predicted sequence for "New York": I-LOC I-LOC; At each step, take argmax of the current token's emission scores only.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 106: Named Entity Recognition — BERT + CRF puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

29. Evaluation — two F1 metrics

Section

Part 4 of 4

30. Token-level vs entity-level F1

Concept

NER has two distinct evaluation metrics. They can diverge dramatically — a system that gets every boundary wrong can still report high token-level F1.

MetricUnit of measurementMatch criterion
Token-level F1individual tokenslabel of each token matches
Entity-level F1full entity spansstart, end, AND type all match exactly

Entity-level F1 (the CoNLL seqeval standard) counts a prediction correct only if the entire span aligns with ground truth — partial credit is zero.

31. Teach it back: Token-level vs entity-level F1

Explain it

Discussion prompt

Explain Token-level vs entity-level F1 to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

NER has two distinct evaluation metrics. They can diverge dramatically — a system that gets every boundary wrong can still report high token-level F1.

32. Finish it with less help: Token vs entity F1 — concrete divergence

Faded example

Fill in the blanks

Token vs entity F1 — concrete divergence, with the scaffolding fading: two lines are gone now — fill both.

# True: ['B-PER','I-PER','O'] => entity 'Alice Smith' at tokens [0,1]
# Prediction A: correct
predA = ['B-PER','I-PER','O']
# Prediction B: wrong interior label
predB = ['B-PER','B-PER','O']
true = ['B-PER','I-PER','O']

# Token-level accuracy
tokA = sum(p==t for p,t in zip(predA, true)) / 3 # 1.000
tokB = sum(p==t for p,t in zip(predB, true)) / 3 # 0.667
print(f'Tok-acc A: ___ Tok-acc B: ___')

# Entity-level: does the predicted span match exactly?
# predA: span B-PER at 0, I-PER at 1 => entity (PER, 0, 1) -> MATCH
# predB: span B-PER at 0 (len 1), B-PER at 1 (len 1) => (PER,0,0),(PER,1,1)
# neither matches gold (PER, 0, 1) -> 0 correct
entA_correct, entB_correct = 1, 0
print(f'Ent-F1 A: ___/1 Ent-F1 B: ___/1')
# predA: token-acc=1.000, entity-F1=1.0
# predB: token-acc=0.667, entity-F1=0.0

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Neither matches the gold span (PER,0,1).

33. Token vs entity F1 — concrete divergence

Worked example

# True: ['B-PER','I-PER','O']  => entity 'Alice Smith' at tokens [0,1]
# Prediction A: correct
predA = ['B-PER','I-PER','O']
# Prediction B: wrong interior label
predB = ['B-PER','B-PER','O']
true  = ['B-PER','I-PER','O']

# Token-level accuracy
tokA = sum(p==t for p,t in zip(predA, true)) / 3   # 1.000
tokB = sum(p==t for p,t in zip(predB, true)) / 3   # 0.667
print(f'Tok-acc A: {tokA:.3f}  Tok-acc B: {tokB:.3f}')

# Entity-level: does the predicted span match exactly?
# predA: span B-PER at 0, I-PER at 1 => entity (PER, 0, 1) -> MATCH
# predB: span B-PER at 0 (len 1), B-PER at 1 (len 1) => (PER,0,0),(PER,1,1)
#         neither matches gold (PER, 0, 1) -> 0 correct
entA_correct, entB_correct = 1, 0
print(f'Ent-F1  A: {entA_correct}/1  Ent-F1  B: {entB_correct}/1')
# predA: token-acc=1.000, entity-F1=1.0
# predB: token-acc=0.667, entity-F1=0.0
PredictionSequenceToken-accEntity-F1
A (correct)B-PER, I-PER, O1.0001.000
B (bad span)B-PER, B-PER, O0.6670.000

Prediction B produces two one-token spans (PER,0,0) and (PER,1,1)

Why: Neither matches the gold span (PER,0,1). Entity-level F1 gives zero. Yet token-level gives 2/3 — a misleadingly high score.

Always report entity-level F1 in NER papers and olympiad problems

Why: Token-level F1 is inflated by the dominant O label and fails to penalize span boundary errors. Entity-level is the standard used in CoNLL-2003 leaderboards.

34. Inspect it line by line: Token vs entity F1 — concrete divergence

Error analysis

Annotate

Walk the callouts on Token vs entity F1 — concrete divergence. Each one is a place this is easy to get subtly wrong.

  • Neither matches the gold span (PER,0,1). Entity-level F1 gives zero. Yet token-level gives 2/3 — a misleadingly high score.
  • Token-level F1 is inflated by the dominant O label and fails to penalize span boundary errors. Entity-level is the standard used in CoNLL-2003 leaderboards.

35. Annotation cost and weak supervision

Concept

NER requires token-level human annotation — far more expensive than sentence-level classification. CoNLL-2003 English took substantial expert effort.

SplitSentencesTokens
Train14,041203,621
Dev3,25051,362
Test3,68446,435

For new domains (medical, legal, financial) annotation is prohibitively expensive. Weak supervision (Snorkel, dictionary matching, heuristic labeling functions) and active learning (annotate the most informative examples first) are the two standard mitigations.

36. Fill in: Tokens for Annotation cost and weak supervision

Comparison

Comparison matrix

From Annotation cost and weak supervision: refill the Tokens column from what you know. The rest of the table is as it appeared.

SplitSentencesTokens
Train14,041203,621
Dev3,25051,362
Test3,68446,435

37. Without one step: Pattern: BERT + CRF NER pipeline

Constraint

Discussion prompt

Run Pattern: BERT + CRF NER pipeline with this step confiscated:

CRF score: add transition matrix A (L x L); full sequence score = sum(emissions) + sum(transitions)

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Define label set: L = 2K + 1 labels (BIO scheme, K entity types)
  2. Tokenize: WordPiece/BPE tokenization — track alignment to original words for evaluation
  3. Encode: pass token IDs through BERT; collect all T hidden states (not just CLS)
  4. Project: shared linear layer maps each hidden h_t in R^H to logits in R^L
  5. CRF score: add transition matrix A (L x L); full sequence score = sum(emissions) + sum(transitions)
  6. Train: maximize log P(gold sequence) via forward algorithm (log-sum-exp over all paths)
  7. Infer: Viterbi decoding — O(T * L^2) — returns globally optimal label path
  8. Evaluate: use entity-level F1 (seqeval); report precision, recall, and F1 per entity type

38. Pattern: BERT + CRF NER pipeline

Pattern

  1. Define label set: L = 2K + 1 labels (BIO scheme, K entity types)
  2. Tokenize: WordPiece/BPE tokenization — track alignment to original words for evaluation
  3. Encode: pass token IDs through BERT; collect all T hidden states (not just CLS)
  4. Project: shared linear layer maps each hidden h_t in R^H to logits in R^L
  5. CRF score: add transition matrix A (L x L); full sequence score = sum(emissions) + sum(transitions)
  6. Train: maximize log P(gold sequence) via forward algorithm (log-sum-exp over all paths)
  7. Infer: Viterbi decoding — O(T * L^2) — returns globally optimal label path
  8. Evaluate: use entity-level F1 (seqeval); report precision, recall, and F1 per entity type

39. Where does it stop working: Pattern: BERT + CRF NER pipeline

Edge cases

Discussion prompt

Pattern: BERT + CRF NER pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Define label set: L = 2K + 1 labels (BIO scheme, K entity types)
  2. Tokenize: WordPiece/BPE tokenization — track alignment to original words for evaluation
  3. Encode: pass token IDs through BERT; collect all T hidden states (not just CLS)
  4. Project: shared linear layer maps each hidden h_t in R^H to logits in R^L
  5. CRF score: add transition matrix A (L x L); full sequence score = sum(emissions) + sum(transitions)
  6. Train: maximize log P(gold sequence) via forward algorithm (log-sum-exp over all paths)
  7. Infer: Viterbi decoding — O(T * L^2) — returns globally optimal label path
  8. Evaluate: use entity-level F1 (seqeval); report precision, recall, and F1 per entity type

40. Rule out three: Check 1: BIO validity

Elimination

Eliminate the wrong options

Which of the following label sequences violates the BIO constraint?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. B-PER I-PER O
  • B. O B-ORG I-ORG B-LOC
  • C. O I-ORG B-PER
  • D. B-PER O B-PER

Survives elimination: C

Why: I-ORG at position 1 is invalid because the preceding label is O, not B-ORG or I-ORG. A valid I-label must immediately follow a B or I of the same type. The other sequences are all legal BIO.

41. Check 1: BIO validity

Check

Identify the invalid BIO label sequence.

Check your understanding

Which of the following label sequences violates the BIO constraint?

  • A. B-PER I-PER O
  • B. O B-ORG I-ORG B-LOC
  • C. O I-ORG B-PER (correct)
  • D. B-PER O B-PER

Answer: C

Why: I-ORG at position 1 is invalid because the preceding label is O, not B-ORG or I-ORG. A valid I-label must immediately follow a B or I of the same type. The other sequences are all legal BIO.

Why A tempts people
B-PER followed by I-PER is the canonical valid continuation; then O ends the span cleanly.
Why B tempts people
O then B-ORG starts a new ORG entity; I-ORG continues it; B-LOC begins a new LOC entity — all valid transitions.
Why D tempts people
B-PER then O ends the entity; B-PER again starts a new entity. Two singleton PER entities — valid BIO even though the same type repeats.

42. Answer it before you see the options: Check 2: Viterbi vs greedy

Prediction

Predict first

Token t has emission scores: O=3.0, I-PER=3.8. The previous Viterbi scores were: V[O]=5.0, V[B-PER]=4.0. Transitions A[O, I-PER]=-5.0 and A[B-PER, I-PER]=+1.5. What does Viterbi choose at this token?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: O (because 5.0 + (-5.0) + 3.8 = 3.8; vs B-PER path: 4.0 + 1.5 + 3.8 = 9.3; Viterbi picks I-PER via B-PER)

Why: Viterbi computes both predecessor paths for I-PER. Via O: 5.0 + A[O,I-PER] + emit = 5.0 + (-5.0) + 3.8 = 3.8. Via B-PER: 4.0 + A[B-PER,I-PER] + emit = 4.0 + 1.5 + 3.8 = 9.3. Best score is 9.3, so Viterbi chooses I-PER with predecessor B-PER — the globally valid continuation.

43. Check 2: Viterbi vs greedy

Check

Reason through what CRF + Viterbi gains over greedy argmax.

Check your understanding

Token t has emission scores: O=3.0, I-PER=3.8. The previous Viterbi scores were: V[O]=5.0, V[B-PER]=4.0. Transitions A[O, I-PER]=-5.0 and A[B-PER, I-PER]=+1.5. What does Viterbi choose at this token?

  • A. I-PER (greedy emission winner)
  • B. O (transition from O path: 5.0 + 3.0 = 8.0; I-PER path via B-PER: 4.0+1.5+3.8=9.3, so actually I-PER)
  • C. O (accumulated best: V[O]=5.0 + A[O,O] not given, so tie broken)
  • D. O (because 5.0 + (-5.0) + 3.8 = 3.8; vs B-PER path: 4.0 + 1.5 + 3.8 = 9.3; Viterbi picks I-PER via B-PER) (correct)

Answer: D

Why: Viterbi computes both predecessor paths for I-PER. Via O: 5.0 + A[O,I-PER] + emit = 5.0 + (-5.0) + 3.8 = 3.8. Via B-PER: 4.0 + A[B-PER,I-PER] + emit = 4.0 + 1.5 + 3.8 = 9.3. Best score is 9.3, so Viterbi chooses I-PER with predecessor B-PER — the globally valid continuation.

Why A tempts people
Greedy emission argmax (I-PER=3.8 > O=3.0) would pick I-PER, but for the wrong reason — it ignores transition costs entirely, and in this case happens to be correct only because the B-PER path is also large.
Why B tempts people
The text of choice B contains the right computation but misidentifies the winner as O. The computed value 9.3 for the I-PER via B-PER path is higher than the O path, so I-PER wins.
Why C tempts people
This answer is vague and ignores the quantitative comparison. The actual Viterbi computation requires evaluating all predecessor paths with their transition scores, not just V[O].

44. Rule out three: Check 3: token vs entity F1

Elimination

Eliminate the wrong options

Gold: [B-PER, I-PER, O]. Prediction: [B-PER, I-PER, I-PER]. Token-level accuracy is 2/3. What is the entity-level result (seqeval, exact span match)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Entity-level F1 = 1.0 because the start of the span is correct
  • B. Entity-level F1 = 0.5 because two of the three entity tokens match
  • C. Entity-level F1 = 0 because the predicted span (tokens 0-2) does not match the gold span (tokens 0-1)
  • D. Entity-level F1 = 0.667 matching the token-level accuracy

Survives elimination: C

Why: The gold PER entity is the span [0,1] (2 tokens). The predicted entity is [0,2] (3 tokens, extended by the extra I-PER). Exact span match requires identical start, end, and type — (PER,0,2) != (PER,0,1), so precision=0, recall=0, F1=0. Entity-level gives zero despite 2/3 token accuracy.

45. Check 3: token vs entity F1

Check

A NER system predicts [B-PER, I-PER, I-PER] but the gold is [B-PER, I-PER, O]. What are the token-level and entity-level scores?

Check your understanding

Gold: [B-PER, I-PER, O]. Prediction: [B-PER, I-PER, I-PER]. Token-level accuracy is 2/3. What is the entity-level result (seqeval, exact span match)?

  • A. Entity-level F1 = 1.0 because the start of the span is correct
  • B. Entity-level F1 = 0.5 because two of the three entity tokens match
  • C. Entity-level F1 = 0 because the predicted span (tokens 0-2) does not match the gold span (tokens 0-1) (correct)
  • D. Entity-level F1 = 0.667 matching the token-level accuracy

Answer: C

Why: The gold PER entity is the span [0,1] (2 tokens). The predicted entity is [0,2] (3 tokens, extended by the extra I-PER). Exact span match requires identical start, end, and type — (PER,0,2) != (PER,0,1), so precision=0, recall=0, F1=0. Entity-level gives zero despite 2/3 token accuracy.

Why A tempts people
Entity-level evaluation requires the full span to match, not just the B-label. A matching start is necessary but not sufficient.
Why B tempts people
Entity-level F1 does not give partial credit for matching tokens within a span. The entire span boundary must be exact — no fractional score exists.
Why D tempts people
Token-level and entity-level F1 are completely independent metrics and will in general differ, as this example demonstrates. They coincide only when every span boundary happens to be exact.

46. Does CRF help in practice?

Concept

With BERT encoders, CRF gains are smaller than in the LSTM era — BERT contextual representations already encode significant boundary information. But the CRF still provides structural guarantees.

ModelCoNLL-2003 entity F1 (approx)
BERT-base + linear head~91.1%
BERT-base + CRF~92.0%
BERT-large + CRF~92.8%
BiLSTM + CRF (pre-BERT)~87-89%

CRF contributes ~0.9 F1 points over BERT+linear. The gain is larger on low-resource settings where the model has seen fewer training examples of valid transitions.

47. Teach it back: Does CRF help in practice?

Explain it

Discussion prompt

Explain Does CRF help in practice? to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

With BERT encoders, CRF gains are smaller than in the LSTM era — BERT contextual representations already encode significant boundary information. But the CRF still provides structural guarantees.

48. Subword tokenization challenge in NER

Concept

BERT uses WordPiece tokenization — a word like 'playing' becomes ['play', '##ing'], two tokens. NER labels are defined at the word level, not subword level.

WordSubword tokensLabel strategy
Alice['Alice']B-PER on 'Alice'
playing['play','##ing']B-X on first; ignore ##ing or copy label
New['New']B-LOC on 'New'
York['York']I-LOC on 'York'

Standard practice: assign the word-level label to the first subword token, set all continuation subwords (##...) to a special -100 ignore label during training. Evaluation maps predictions back to word boundaries before computing F1.

49. By analogy: Subword tokenization challenge in NER

Analogy

Discussion prompt

Explain Subword tokenization challenge in NER by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

BERT uses WordPiece tokenization — a word like 'playing' becomes ['play', '##ing'], two tokens. NER labels are defined at the word level, not subword level.

50. Lesson callbacks: where NER sits

Concept

Lesson 88: Transformer encoder
Self-attention produces contextualized token representations — the encoder BERT NER reads.
Lesson 100: Text classification
Used CLS token only. NER uses all T token outputs — a dense prediction head.
Lesson 104: Sequence-to-Sequence
Seq2seq generates variable-length output. NER is input-aligned: exactly one label per token.
Lesson 107: Relation extraction
NER identifies entity spans; relation extraction then labels pairs of spans — next lesson.

51. Which is which: Lesson callbacks: where NER sits

Matching

Match the pairs

From Lesson callbacks: where NER sits — match each one to what it actually does. The descriptions have been shuffled.

  • c1. Lesson 88: Transformer encoder
  • c2. Lesson 100: Text classification
  • c3. Lesson 104: Sequence-to-Sequence
  • c4. Lesson 107: Relation extraction
  • b1. Self-attention produces contextualized token representations — the encoder BERT NER reads.
  • b2. Used CLS token only. NER uses all T token outputs — a dense prediction head.
  • b3. Seq2seq generates variable-length output. NER is input-aligned: exactly one label per token.
  • b4. NER identifies entity spans; relation extraction then labels pairs of spans — next lesson.

Why: Lesson 88: Transformer encoder, Lesson 100: Text classification, Lesson 104: Sequence-to-Sequence, Lesson 107: Relation extraction are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

52. Your Turn — project brief

Concept

Build a TinyBertNER system: fine-tune a small transformer encoder on a synthetic NER dataset and compare greedy vs Viterbi decoding on held-out examples.

53. Break it if you can: Your Turn — project brief

Counterexample

Discussion prompt

Build a TinyBertNER system: fine-tune a small transformer encoder on a synthetic NER dataset and compare greedy vs Viterbi decoding on held-out examples.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

54. Restore the missing line: Your Turn — Milestone 1: BIO encode/decode

Fill the middle

Fill in the blanks

From Your Turn — Milestone 1: BIO encode/decode — one line has had its right-hand side removed. Put it back.

def bio_encode(spans, seq_len, tag='O'):
"""spans: list of (type_str, start, end_exclusive)"""
labels = ['O'] * seq_len
for etype, start, end in spans:
labels[start] = f'B-labels[i][2:]'
for i in range(start + 1, end):
labels[i] = f'I-___'
return labels

def bio_decode(labels):
"""Returns list of (type_str, start, end_exclusive) tuples."""
spans, i = [], 0
while i < len(labels):
if labels[i].startswith('B-'):
etype = ___
j = i + 1
while j < len(labels) and labels[j] == f'I-___':
j += 1
spans.append((etype, i, j))
i = j
else:
i += 1
return spans

# Verify round-trip
spans_in = [('PER', 0, 2), ('ORG', 4, 5)]
labels = bio_encode(spans_in, seq_len=6)
spans_out = bio_decode(labels)
print(labels) # ['B-PER','I-PER','O','O','B-ORG','O']
print(spans_out) # [('PER', 0, 2), ('ORG', 4, 5)]
print(spans_in == spans_out) # True

Why: etype is what everything below it consumes, so the wrong expression here fails later and somewhere else. bio_encode produces ['B-PER','I-PER','O','O','B-ORG','O']; bio_decode reconstructs exactly [('PER',0,2),('ORG',4,5)].

55. Your Turn — Milestone 1: BIO encode/decode

Worked example

def bio_encode(spans, seq_len, tag='O'):
    """spans: list of (type_str, start, end_exclusive)"""
    labels = ['O'] * seq_len
    for etype, start, end in spans:
        labels[start] = f'B-{etype}'
        for i in range(start + 1, end):
            labels[i] = f'I-{etype}'
    return labels

def bio_decode(labels):
    """Returns list of (type_str, start, end_exclusive) tuples."""
    spans, i = [], 0
    while i < len(labels):
        if labels[i].startswith('B-'):
            etype = labels[i][2:]
            j = i + 1
            while j < len(labels) and labels[j] == f'I-{etype}':
                j += 1
            spans.append((etype, i, j))
            i = j
        else:
            i += 1
    return spans

# Verify round-trip
spans_in  = [('PER', 0, 2), ('ORG', 4, 5)]
labels    = bio_encode(spans_in, seq_len=6)
spans_out = bio_decode(labels)
print(labels)     # ['B-PER','I-PER','O','O','B-ORG','O']
print(spans_out)  # [('PER', 0, 2), ('ORG', 4, 5)]
print(spans_in == spans_out)  # True
Input spanseq_lenOutput labels
(PER, 0, 2)6B-PER I-PER O O B-ORG O
(ORG, 4, 5)6(same output above)
round-trip6spans_in == spans_out: True

Verify with seq_len=6, spans [(PER,0,2),(ORG,4,5)]

Why: bio_encode produces ['B-PER','I-PER','O','O','B-ORG','O']; bio_decode reconstructs exactly [('PER',0,2),('ORG',4,5)]. Round-trip identity confirms both functions are correct inverses.

56. What each one costs: Your Turn — Milestone 1: BIO encode/decode

Trade off

Comparison matrix

From Your Turn — Milestone 1: BIO encode/decode: every row here is a choice with a cost. Fill the Output labels column, then say which row you would actually pick and what you give up for it.

Input spanseq_lenOutput labels
(PER, 0, 2)6B-PER I-PER O O B-ORG O
(ORG, 4, 5)6(same output above)
round-trip6spans_in == spans_out: True

57. Predict the next row: Your Turn — full program

Pattern

Predict first

The table runs: Greedy argmax | per-token independent | No — ignores transitions · Viterbi | joint sequence search | Yes — O(T * L^2)

In Your Turn — full program, given the rows so far: what is the next one — the row where Decoder is Difference??

Correct: Difference? | likely on random trans | Run code to verify

DecoderOutput (label IDs)Globally optimal?
Greedy argmaxper-token independentNo — ignores transitions
Viterbijoint sequence searchYes — O(T * L^2)
Difference?likely on random transRun code to verify

Why: The relationship between the columns, not the individual numbers, is what generates the next row. With a random transition matrix, Viterbi's globally optimal path will typically differ from greedy label-by-label choices.

58. Your Turn — full program

Worked example

import torch, torch.nn as nn
# 1. Build model
class TinyBertNER(nn.Module):
    def __init__(self, V=200, H=64, heads=4, layers=2, L=7):
        super().__init__()
        self.embed = nn.Embedding(V, H, padding_idx=0)
        self.pos   = nn.Embedding(128, H)
        enc = nn.TransformerEncoderLayer(
            d_model=H, nhead=heads,
            dim_feedforward=128, batch_first=True, dropout=0.0)
        self.enc  = nn.TransformerEncoder(enc, num_layers=layers)
        self.head = nn.Linear(H, L)
    def forward(self, ids):
        B, T = ids.shape
        x = self.embed(ids) + self.pos(
                torch.arange(T).expand(B,-1))
        return self.head(self.enc(x))

torch.manual_seed(42)
model = TinyBertNER()
ids    = torch.randint(1, 200, (1, 6))
logits = model(ids)           # (1, 6, 7)
greedy = logits.argmax(-1)[0] # (6,)

# 2. Viterbi (random transition for demo)
L   = 7
trans = torch.randn(L, L)
e   = logits[0]               # (6, 7)
V_s = e[0].clone()
ptr = torch.zeros(6, L, dtype=torch.long)
for t in range(1, 6):
    sc = V_s.unsqueeze(1) + trans
    ptr[t] = sc.max(0).indices
    V_s = sc.max(0).values + e[t]
cur, path = V_s.argmax().item(), []
path.append(cur)
for t in range(5, 0, -1):
    cur = ptr[t, cur].item(); path.append(cur)
path.reverse()
print('Greedy:', greedy.tolist())
print('Viterbi:', path)
print('Same?', greedy.tolist() == path)
DecoderOutput (label IDs)Globally optimal?
Greedy argmaxper-token independentNo — ignores transitions
Viterbijoint sequence searchYes — O(T * L^2)
Difference?likely on random transRun code to verify

Run the full program and check Same? — expect False with random transitions

Why: With a random transition matrix, Viterbi's globally optimal path will typically differ from greedy label-by-label choices. The diff reveals exactly where transition penalties redirect the path.

59. Fill in: Globally optimal? for Your Turn — full program

Comparison

Comparison matrix

From Your Turn — full program: refill the Globally optimal? column from what you know. The rest of the table is as it appeared.

DecoderOutput (label IDs)Globally optimal?
Greedy argmaxper-token independentNo — ignores transitions
Viterbijoint sequence searchYes — O(T * L^2)
Difference?likely on random transRun code to verify

60. Connect it up: Lesson 106: Named Entity Recognition — BERT + CRF

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — BIO Tagging — the sequence labeling contract · BERT + Linear Head — token-level logits · CRF Layer — globally consistent label sequences · Evaluation — two F1 metrics. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. Lesson 106 recap

Recap

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 106 — NER: BIO Tagging, BERT + CRF, Evaluation — Barron · USAAIO Round 2 Preparation, 2026
  2. Lafferty et al. 'Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data' (ICML 2001) — ICML 2001, pp. 282-289
  3. Devlin et al. 'BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding' (NAACL 2019) — arXiv:1810.04805
  4. Sang & De Meulder, 'Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition' (2003) — CoNLL-2003 dataset statistics verified
  5. Viterbi trace, CRF scores, TinyBertNER (81991 params), entity vs token F1 verified with torch 2.7.1+cpu, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108