USAAIO Lesson 106, from Phase 3. It covers the BIO tagging scheme for named-entity recognition, a BERT encoder feeding a per-token linear head, and a CRF transition matrix that enforces valid label sequences. Viterbi decoding is traced on a three-token toy example, with the good path scoring 11.70 and the bad path 4.80, and the lesson contrasts token-level with entity-level F1 using a concrete span-mismatch example. The shapes of TinyBertNER - 81,991 parameters at hidden size 64 - were verified with torch 2.7.1+cpu. The lesson runs to 29 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 106 · Phase 3
Tag each token with its entity type. BIO scheme, BERT encoder, CRF for globally consistent label sequences, Viterbi decoding — every number verified in PyTorch.
Objectives
O → I-PERWarm-up
Discussion prompt
Before we open Lesson 106: Named Entity Recognition — BERT + CRF: without looking back, what was the main idea of BLEU, ROUGE, and BERTScore, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
BLEU n-gram precision with brevity penalty, ROUGE recall-oriented n-gram overlap and ROUGE-L via LCS, BERTScore via contextual embedding cosine similarity, and when each metric fails. All computations implemented from scratch in Python and verified with real execution — p_1=5/6, p_2=3/5, BLEU-4=0.4209 on MT example, ROUGE-L F1=0.8333 on LCS example.
Section
Part 1 of 4
Concept
Named Entity Recognition assigns a label to every token. NER is not a span-extraction task — it is a per-token classification problem, and the label set encodes span boundaries via the BIO scheme.
BIO scheme — Each token gets label O (outside), B-TYPE (beginning of an entity span of that type), or I-TYPE (inside/continuation of a span). A span is exactly the maximal run B followed by zero or more consecutive I of the same type.
| Token | Label | Span role |
|---|---|---|
| John | B-PER | starts PER span |
| Smith | I-PER | continues PER span |
| works | O | outside |
| at | O | outside |
| B-ORG | starts ORG span | |
| in | O | outside |
| New | B-LOC | starts LOC span |
| York | I-LOC | continues LOC span |
Comparison
Comparison matrix
From NER as token classification: refill the Label column from what you know. The rest of the table is as it appeared.
| Token | Label | Span role |
|---|---|---|
| John | B-PER | starts PER span |
| Smith | I-PER | continues PER span |
| works | O | outside |
| at | O | outside |
| B-ORG | starts ORG span | |
| in | O | outside |
| New | B-LOC | starts LOC span |
| York | I-LOC | continues LOC span |
Concept
For 3 entity types (PER, ORG, LOC) the full label set is size 7: O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC. In general, L = 2K + 1 for K entity types.
\[ L = 2K + 1 \quad (K \text{ entity types}) \]
| Label ID | Label | Meaning |
|---|---|---|
| 0 | O | Outside any entity |
| 1 | B-PER | Begin PER span |
| 2 | I-PER | Inside PER span |
| 3 | B-ORG | Begin ORG span |
| 4 | I-ORG | Inside ORG span |
| 5 | B-LOC | Begin LOC span |
| 6 | I-LOC | Inside LOC span |
Trade off
Comparison matrix
From The 7-label BIO vocabulary: every row here is a choice with a cost. Fill the Label column, then say which row you would actually pick and what you give up for it.
| Label ID | Label | Meaning |
|---|---|---|
| 0 | O | Outside any entity |
| 1 | B-PER | Begin PER span |
| 2 | I-PER | Inside PER span |
| 3 | B-ORG | Begin ORG span |
| 4 | I-ORG | Inside ORG span |
| 5 | B-LOC | Begin LOC span |
| 6 | I-LOC | Inside LOC span |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Predicted sequence for "New York": I-LOC I-LOC
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.
Correct sequence for "New York": B-LOC I-LOC
Why: The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.
Trap
Predicted sequence for "New York": I-LOC I-LOC
Accept I-LOC as the first label of the span
Why: The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.
Entity-level evaluators treat this as zero entities detected — no span can be extracted without a B-label.
Correct sequence for "New York": B-LOC I-LOC
Require B-LOC on the first token of every new span
Why: BIO contract: I-TYPE is only valid immediately after B-TYPE or I-TYPE of the same type. The CRF transition matrix learns to assign -inf score to O->I and B-X->I-Y (type mismatch).
With a CRF layer, invalid transitions are globally penalized even if the token's emission score strongly favors I-LOC.
Break the constraint
Discussion prompt
The rule this trap just fixed:
BIO contract: I-TYPE is only valid immediately after B-TYPE or I-TYPE of the same type. The CRF transition matrix learns to assign -inf score to O->I and B-X->I-Y (type mismatch).
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
The model saw high confidence for I-LOC on 'New' and output it — but no B-LOC precedes it, so the span has no defined start.
Section
Part 2 of 4
Concept
Unlike classification tasks that read only the [CLS] token (Lesson 88), NER reads every token's final hidden state and passes it through a shared linear head to produce label logits.
\[ \text{logits} = W \cdot h_t + b \quad \forall t \in \{0,\ldots,T-1\} \]
Result: a (B, T, L) tensor of unnormalized scores — one length-L vector per token per example in the batch.
Counterexample
Discussion prompt
Unlike classification tasks that read only the [CLS] token (Lesson 88), NER reads every token's final hidden state and passes it through a shared linear head to produce label logits.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Result: a (B, T, L) tensor of unnormalized scores — one length-L vector per token per example in the batch.
Pattern
Predict first
The table runs: token IDs | (1, 8) | batch=1, seq_len=8 · embed + pos | (1, 8, 64) | token + positional embeddings · encoder output | (1, 8, 64) | contextualized hidden states · logits (head) | (1, 8, 7) | score per label per token
In BERT + linear head: shape trace, given the rows so far: what is the next one — the row where Tensor is argmax preds?
Correct: argmax preds | (1, 8) | greedy label IDs (0..6)
| Tensor | Shape | Description |
|---|---|---|
| token IDs | (1, 8) | batch=1, seq_len=8 |
| embed + pos | (1, 8, 64) | token + positional embeddings |
| encoder output | (1, 8, 64) | contextualized hidden states |
| logits (head) | (1, 8, 7) | score per label per token |
| argmax preds | (1, 8) | greedy label IDs (0..6) |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Embedding(100,64)=6400 + Embedding(128,64)=8192 + TransformerEncoder(2 layers, H=64, FFN=128, heads=4) + Linear(64,7)=455.
Worked example
import torch, torch.nn as nn
class TinyBertNER(nn.Module):
def __init__(self, vocab=100, H=64, heads=4, layers=2, L=7):
super().__init__()
self.embed = nn.Embedding(vocab, H, padding_idx=0)
self.pos = nn.Embedding(128, H)
enc = nn.TransformerEncoderLayer(
d_model=H, nhead=heads,
dim_feedforward=128, batch_first=True, dropout=0.0)
self.encoder = nn.TransformerEncoder(enc, num_layers=layers)
self.head = nn.Linear(H, L)
def forward(self, ids):
B, T = ids.shape
pos = torch.arange(T).unsqueeze(0).expand(B, -1)
x = self.embed(ids) + self.pos(pos)
x = self.encoder(x) # (B, T, H)
return self.head(x) # (B, T, L)
torch.manual_seed(42)
model = TinyBertNER()
print(sum(p.numel() for p in model.parameters())) # 81991
tokens = torch.tensor([[5, 12, 8, 23, 45, 7, 61, 19]])
logits = model(tokens)
print(logits.shape) # torch.Size([1, 8, 7])
preds = logits.argmax(-1).squeeze()
print(preds.tolist()) # per-token label IDs| Tensor | Shape | Description |
|---|---|---|
| token IDs | (1, 8) | batch=1, seq_len=8 |
| embed + pos | (1, 8, 64) | token + positional embeddings |
| encoder output | (1, 8, 64) | contextualized hidden states |
| logits (head) | (1, 8, 7) | score per label per token |
| argmax preds | (1, 8) | greedy label IDs (0..6) |
Count parameters: 81,991 total
Why: Embedding(100,64)=6400 + Embedding(128,64)=8192 + TransformerEncoder(2 layers, H=64, FFN=128, heads=4) + Linear(64,7)=455. Verified by running sum(p.numel() for p in model.parameters()).
Greedy decoding: argmax over the 7-class dimension
Why: With an untrained model, predictions are essentially random. After fine-tuning on labeled data, argmax often gives valid BIO sequences — but not always, motivating the CRF layer.
Discrimination
Sort into buckets
Sort these by Shape, from memory, without looking back at BERT + linear head: shape trace. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Part 3 of 4
Concept
A Conditional Random Field (CRF) adds a learnable transition matrix A of shape (L, L). Entry A[i,j] is the score of transitioning from label i to label j. The model scores a full label sequence, not each token independently.
\[ \text{score}(\mathbf{y}) = \sum_{t=0}^{T-1} \underbrace{e_t[y_t]}_{\text{emission}} + \sum_{t=1}^{T-1} \underbrace{A[y_{t-1}, y_t]}_{\text{transition}} \]
Training maximizes the log-likelihood of the gold sequence relative to all sequences via the forward algorithm. Inference finds the highest-scoring sequence via Viterbi decoding.
Analogy
Discussion prompt
Explain What the CRF adds: a transition matrix by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Training maximizes the log-likelihood of the gold sequence relative to all sequences via the forward algorithm. Inference finds the highest-scoring sequence via Viterbi decoding.
Concept
The CRF does not hard-code BIO rules — it learns them from data. After training, A[O, I-PER] becomes very negative (never observed), while A[B-PER, I-PER] becomes very positive.
| Transition | Valid? | Reason |
|---|---|---|
| O -> B-PER | YES | Start a new PER entity |
| B-PER -> I-PER | YES | Continue the PER entity |
| B-PER -> I-ORG | NO | Type mismatch inside span |
| I-ORG -> I-PER | NO | Type mismatch inside span |
| O -> I-PER | NO | I-label without preceding B-label |
| I-LOC -> B-PER | YES | End LOC, start new PER |
The key advantage over greedy token-by-token decoding: even if the emission score strongly favors I-ORG after an O, the learned transition penalty A[O, I-ORG] can override it globally.
Discrimination
Sort into buckets
Sort these by Valid?, from memory, without looking back at Valid vs invalid BIO transitions. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Ranking
Put in order
Put the moves of Viterbi decoding — 3-token trace into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. No prior label exists, so only emission scores contribute.
Worked example
import torch
# 3 tokens, 4 labels: O=0 B-PER=1 I-PER=2 B-ORG=3
# Emission scores (unnormalized log probs per token per label)
emit = torch.tensor([
[0.10, 3.00, 0.50, 0.20], # tok0: high B-PER
[0.20, 0.30, 3.50, 0.10], # tok1: high I-PER
[3.20, 0.10, 0.20, 0.30], # tok2: high O
])
# Transition matrix A[from, to]
trans = torch.tensor([
[ 1.0, 0.5, -1.0, 0.5], # from O
[ 0.5, 0.5, 1.5, -2.0], # from B-PER
[ 0.5, 0.5, 0.5, -2.0], # from I-PER
[ 0.5, 0.5, -2.0, 0.5], # from B-ORG
])
# Viterbi forward pass
V = emit[0].clone() # t=0: no transition yet
bptr = torch.zeros(3, 4, dtype=torch.long)
for t in range(1, 3):
scores = V.unsqueeze(1) + trans # (4,4)
best = scores.max(0) # max over 'from' dim
bptr[t] = best.indices
V = best.values + emit[t]
# Backtrack
path, cur = [], V.argmax().item()
path.append(cur)
for t in range(2, 0, -1):
cur = bptr[t, cur].item(); path.append(cur)
path.reverse()
print(path) # [1, 2, 0] => B-PER, I-PER, O| t | O | B-PER | I-PER | B-ORG | Chosen |
|---|---|---|---|---|---|
| 0 (Alice) | 0.10 | 3.00 | 0.50 | 0.20 | B-PER |
| 1 (Smith) | 3.70 | 3.80 | 8.00 | 1.10 | I-PER |
| 2 (left) | 11.70 | 8.60 | 8.70 | 6.30 | O |
Initialize: V[0] = emission scores at t=0 = [0.10, 3.00, 0.50, 0.20]
Why: No prior label exists, so only emission scores contribute. B-PER leads with 3.00.
t=1: V[j] = max_i(V_prev[i] + A[i,j]) + emit[1,j]
Why: For I-PER: best predecessor is B-PER (3.00 + A[B-PER,I-PER]=1.5 = 4.50) + emit[1,I-PER]=3.50 = 8.00. This beats O path: 3.70.
t=2: best path score for O = 11.70; backtrack gives [B-PER, I-PER, O]
Why: The CRF scores the good path (emit=9.70, trans=2.00, total=11.70) against the bad path B-PER,B-ORG,O (emit=6.30, trans=-1.50, total=4.80). Good path wins by 6.90 points.
Error analysis
Annotate
Walk the callouts on Viterbi decoding — 3-token trace. Each one is a place this is easy to get subtly wrong.
Anomaly
Predict first
A student writes this, and it looks reasonable:
At each step, take argmax of the current token's emission scores only.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Greedy looks at one token at a time — no transition score considered.
Viterbi finds the globally best sequence by propagating accumulated scores forward and backtracking.
Why: Greedy looks at one token at a time — no transition score considered.
Trap
At each step, take argmax of the current token's emission scores only.
t=1: argmax of emit[1] = I-PER (3.50). Accept it immediately.
Why: Greedy looks at one token at a time — no transition score considered.
This accidentally produces correct output here, but fails whenever a locally high emission conflicts with global structure (e.g., I-ORG after O).
Viterbi finds the globally best sequence by propagating accumulated scores forward and backtracking.
t=1: V[I-PER] = max_i(V_prev[i] + A[i, I-PER]) + emit[1, I-PER] = 8.00
Why: The predecessor B-PER contributes a +1.5 transition bonus; greedy would never see this. Viterbi guarantees the path maximizing the joint sequence score.
CRF + Viterbi runs in O(T * L^2) — practical for the typical label set sizes in NER.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC. In general, L = 2K + 1 for K entity types.; Result: a (B, T, L) tensor of unnormalized scores — one length-L vector per token per example in the batch.; NER has two distinct evaluation metrics. They can diverge dramatically — a system that gets every boundary wrong can still report high token-level F1.I-LOC I-LOC; At each step, take argmax of the current token's emission scores only.Section
Part 4 of 4
Concept
NER has two distinct evaluation metrics. They can diverge dramatically — a system that gets every boundary wrong can still report high token-level F1.
| Metric | Unit of measurement | Match criterion |
|---|---|---|
| Token-level F1 | individual tokens | label of each token matches |
| Entity-level F1 | full entity spans | start, end, AND type all match exactly |
Entity-level F1 (the CoNLL seqeval standard) counts a prediction correct only if the entire span aligns with ground truth — partial credit is zero.
Explain it
Discussion prompt
Explain Token-level vs entity-level F1 to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
NER has two distinct evaluation metrics. They can diverge dramatically — a system that gets every boundary wrong can still report high token-level F1.
Faded example
Fill in the blanks
Token vs entity F1 — concrete divergence, with the scaffolding fading: two lines are gone now — fill both.
# True: ['B-PER','I-PER','O'] => entity 'Alice Smith' at tokens [0,1]
# Prediction A: correct
predA = ['B-PER','I-PER','O']
# Prediction B: wrong interior label
predB = ['B-PER','B-PER','O']
true = ['B-PER','I-PER','O']
# Token-level accuracy
tokA = sum(p==t for p,t in zip(predA, true)) / 3 # 1.000
tokB = sum(p==t for p,t in zip(predB, true)) / 3 # 0.667
print(f'Tok-acc A: ___ Tok-acc B: ___')
# Entity-level: does the predicted span match exactly?
# predA: span B-PER at 0, I-PER at 1 => entity (PER, 0, 1) -> MATCH
# predB: span B-PER at 0 (len 1), B-PER at 1 (len 1) => (PER,0,0),(PER,1,1)
# neither matches gold (PER, 0, 1) -> 0 correct
entA_correct, entB_correct = 1, 0
print(f'Ent-F1 A: ___/1 Ent-F1 B: ___/1')
# predA: token-acc=1.000, entity-F1=1.0
# predB: token-acc=0.667, entity-F1=0.0
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Neither matches the gold span (PER,0,1).
Worked example
# True: ['B-PER','I-PER','O'] => entity 'Alice Smith' at tokens [0,1]
# Prediction A: correct
predA = ['B-PER','I-PER','O']
# Prediction B: wrong interior label
predB = ['B-PER','B-PER','O']
true = ['B-PER','I-PER','O']
# Token-level accuracy
tokA = sum(p==t for p,t in zip(predA, true)) / 3 # 1.000
tokB = sum(p==t for p,t in zip(predB, true)) / 3 # 0.667
print(f'Tok-acc A: {tokA:.3f} Tok-acc B: {tokB:.3f}')
# Entity-level: does the predicted span match exactly?
# predA: span B-PER at 0, I-PER at 1 => entity (PER, 0, 1) -> MATCH
# predB: span B-PER at 0 (len 1), B-PER at 1 (len 1) => (PER,0,0),(PER,1,1)
# neither matches gold (PER, 0, 1) -> 0 correct
entA_correct, entB_correct = 1, 0
print(f'Ent-F1 A: {entA_correct}/1 Ent-F1 B: {entB_correct}/1')
# predA: token-acc=1.000, entity-F1=1.0
# predB: token-acc=0.667, entity-F1=0.0| Prediction | Sequence | Token-acc | Entity-F1 |
|---|---|---|---|
| A (correct) | B-PER, I-PER, O | 1.000 | 1.000 |
| B (bad span) | B-PER, B-PER, O | 0.667 | 0.000 |
Prediction B produces two one-token spans (PER,0,0) and (PER,1,1)
Why: Neither matches the gold span (PER,0,1). Entity-level F1 gives zero. Yet token-level gives 2/3 — a misleadingly high score.
Always report entity-level F1 in NER papers and olympiad problems
Why: Token-level F1 is inflated by the dominant O label and fails to penalize span boundary errors. Entity-level is the standard used in CoNLL-2003 leaderboards.
Error analysis
Annotate
Walk the callouts on Token vs entity F1 — concrete divergence. Each one is a place this is easy to get subtly wrong.
Concept
NER requires token-level human annotation — far more expensive than sentence-level classification. CoNLL-2003 English took substantial expert effort.
| Split | Sentences | Tokens |
|---|---|---|
| Train | 14,041 | 203,621 |
| Dev | 3,250 | 51,362 |
| Test | 3,684 | 46,435 |
For new domains (medical, legal, financial) annotation is prohibitively expensive. Weak supervision (Snorkel, dictionary matching, heuristic labeling functions) and active learning (annotate the most informative examples first) are the two standard mitigations.
Comparison
Comparison matrix
From Annotation cost and weak supervision: refill the Tokens column from what you know. The rest of the table is as it appeared.
| Split | Sentences | Tokens |
|---|---|---|
| Train | 14,041 | 203,621 |
| Dev | 3,250 | 51,362 |
| Test | 3,684 | 46,435 |
Constraint
Discussion prompt
Run Pattern: BERT + CRF NER pipeline with this step confiscated:
CRF score: add transition matrix A (L x L); full sequence score = sum(emissions) + sum(transitions)
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Edge cases
Discussion prompt
Pattern: BERT + CRF NER pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
Which of the following label sequences violates the BIO constraint?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: C
Why: I-ORG at position 1 is invalid because the preceding label is O, not B-ORG or I-ORG. A valid I-label must immediately follow a B or I of the same type. The other sequences are all legal BIO.
Check
Identify the invalid BIO label sequence.
Check your understanding
Which of the following label sequences violates the BIO constraint?
Answer: C
Why: I-ORG at position 1 is invalid because the preceding label is O, not B-ORG or I-ORG. A valid I-label must immediately follow a B or I of the same type. The other sequences are all legal BIO.
Prediction
Predict first
Token t has emission scores: O=3.0, I-PER=3.8. The previous Viterbi scores were: V[O]=5.0, V[B-PER]=4.0. Transitions A[O, I-PER]=-5.0 and A[B-PER, I-PER]=+1.5. What does Viterbi choose at this token?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: O (because 5.0 + (-5.0) + 3.8 = 3.8; vs B-PER path: 4.0 + 1.5 + 3.8 = 9.3; Viterbi picks I-PER via B-PER)
Why: Viterbi computes both predecessor paths for I-PER. Via O: 5.0 + A[O,I-PER] + emit = 5.0 + (-5.0) + 3.8 = 3.8. Via B-PER: 4.0 + A[B-PER,I-PER] + emit = 4.0 + 1.5 + 3.8 = 9.3. Best score is 9.3, so Viterbi chooses I-PER with predecessor B-PER — the globally valid continuation.
Check
Reason through what CRF + Viterbi gains over greedy argmax.
Check your understanding
Token t has emission scores: O=3.0, I-PER=3.8. The previous Viterbi scores were: V[O]=5.0, V[B-PER]=4.0. Transitions A[O, I-PER]=-5.0 and A[B-PER, I-PER]=+1.5. What does Viterbi choose at this token?
Answer: D
Why: Viterbi computes both predecessor paths for I-PER. Via O: 5.0 + A[O,I-PER] + emit = 5.0 + (-5.0) + 3.8 = 3.8. Via B-PER: 4.0 + A[B-PER,I-PER] + emit = 4.0 + 1.5 + 3.8 = 9.3. Best score is 9.3, so Viterbi chooses I-PER with predecessor B-PER — the globally valid continuation.
Elimination
Eliminate the wrong options
Gold: [B-PER, I-PER, O]. Prediction: [B-PER, I-PER, I-PER]. Token-level accuracy is 2/3. What is the entity-level result (seqeval, exact span match)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: C
Why: The gold PER entity is the span [0,1] (2 tokens). The predicted entity is [0,2] (3 tokens, extended by the extra I-PER). Exact span match requires identical start, end, and type — (PER,0,2) != (PER,0,1), so precision=0, recall=0, F1=0. Entity-level gives zero despite 2/3 token accuracy.
Check
A NER system predicts [B-PER, I-PER, I-PER] but the gold is [B-PER, I-PER, O]. What are the token-level and entity-level scores?
Check your understanding
Gold: [B-PER, I-PER, O]. Prediction: [B-PER, I-PER, I-PER]. Token-level accuracy is 2/3. What is the entity-level result (seqeval, exact span match)?
Answer: C
Why: The gold PER entity is the span [0,1] (2 tokens). The predicted entity is [0,2] (3 tokens, extended by the extra I-PER). Exact span match requires identical start, end, and type — (PER,0,2) != (PER,0,1), so precision=0, recall=0, F1=0. Entity-level gives zero despite 2/3 token accuracy.
Concept
With BERT encoders, CRF gains are smaller than in the LSTM era — BERT contextual representations already encode significant boundary information. But the CRF still provides structural guarantees.
| Model | CoNLL-2003 entity F1 (approx) |
|---|---|
| BERT-base + linear head | ~91.1% |
| BERT-base + CRF | ~92.0% |
| BERT-large + CRF | ~92.8% |
| BiLSTM + CRF (pre-BERT) | ~87-89% |
CRF contributes ~0.9 F1 points over BERT+linear. The gain is larger on low-resource settings where the model has seen fewer training examples of valid transitions.
Explain it
Discussion prompt
Explain Does CRF help in practice? to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
With BERT encoders, CRF gains are smaller than in the LSTM era — BERT contextual representations already encode significant boundary information. But the CRF still provides structural guarantees.
Concept
BERT uses WordPiece tokenization — a word like 'playing' becomes ['play', '##ing'], two tokens. NER labels are defined at the word level, not subword level.
| Word | Subword tokens | Label strategy |
|---|---|---|
| Alice | ['Alice'] | B-PER on 'Alice' |
| playing | ['play','##ing'] | B-X on first; ignore ##ing or copy label |
| New | ['New'] | B-LOC on 'New' |
| York | ['York'] | I-LOC on 'York' |
Standard practice: assign the word-level label to the first subword token, set all continuation subwords (##...) to a special -100 ignore label during training. Evaluation maps predictions back to word boundaries before computing F1.
Analogy
Discussion prompt
Explain Subword tokenization challenge in NER by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
BERT uses WordPiece tokenization — a word like 'playing' becomes ['play', '##ing'], two tokens. NER labels are defined at the word level, not subword level.
Concept
Matching
Match the pairs
From Lesson callbacks: where NER sits — match each one to what it actually does. The descriptions have been shuffled.
Why: Lesson 88: Transformer encoder, Lesson 100: Text classification, Lesson 104: Sequence-to-Sequence, Lesson 107: Relation extraction are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Concept
Build a TinyBertNER system: fine-tune a small transformer encoder on a synthetic NER dataset and compare greedy vs Viterbi decoding on held-out examples.
bio_encode(spans, seq_len) and bio_decode(labels) and verify they are inverses(B,T,7)seqeval or a manual span extractor; show the gap vs token-level accuracy on three adversarial examplesCounterexample
Discussion prompt
Build a TinyBertNER system: fine-tune a small transformer encoder on a synthetic NER dataset and compare greedy vs Viterbi decoding on held-out examples.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Fill the middle
Fill in the blanks
From Your Turn — Milestone 1: BIO encode/decode — one line has had its right-hand side removed. Put it back.
def bio_encode(spans, seq_len, tag='O'):
"""spans: list of (type_str, start, end_exclusive)"""
labels = ['O'] * seq_len
for etype, start, end in spans:
labels[start] = f'B-labels[i][2:]'
for i in range(start + 1, end):
labels[i] = f'I-___'
return labels
def bio_decode(labels):
"""Returns list of (type_str, start, end_exclusive) tuples."""
spans, i = [], 0
while i < len(labels):
if labels[i].startswith('B-'):
etype = ___
j = i + 1
while j < len(labels) and labels[j] == f'I-___':
j += 1
spans.append((etype, i, j))
i = j
else:
i += 1
return spans
# Verify round-trip
spans_in = [('PER', 0, 2), ('ORG', 4, 5)]
labels = bio_encode(spans_in, seq_len=6)
spans_out = bio_decode(labels)
print(labels) # ['B-PER','I-PER','O','O','B-ORG','O']
print(spans_out) # [('PER', 0, 2), ('ORG', 4, 5)]
print(spans_in == spans_out) # True
Why: etype is what everything below it consumes, so the wrong expression here fails later and somewhere else. bio_encode produces ['B-PER','I-PER','O','O','B-ORG','O']; bio_decode reconstructs exactly [('PER',0,2),('ORG',4,5)].
Worked example
def bio_encode(spans, seq_len, tag='O'):
"""spans: list of (type_str, start, end_exclusive)"""
labels = ['O'] * seq_len
for etype, start, end in spans:
labels[start] = f'B-{etype}'
for i in range(start + 1, end):
labels[i] = f'I-{etype}'
return labels
def bio_decode(labels):
"""Returns list of (type_str, start, end_exclusive) tuples."""
spans, i = [], 0
while i < len(labels):
if labels[i].startswith('B-'):
etype = labels[i][2:]
j = i + 1
while j < len(labels) and labels[j] == f'I-{etype}':
j += 1
spans.append((etype, i, j))
i = j
else:
i += 1
return spans
# Verify round-trip
spans_in = [('PER', 0, 2), ('ORG', 4, 5)]
labels = bio_encode(spans_in, seq_len=6)
spans_out = bio_decode(labels)
print(labels) # ['B-PER','I-PER','O','O','B-ORG','O']
print(spans_out) # [('PER', 0, 2), ('ORG', 4, 5)]
print(spans_in == spans_out) # True| Input span | seq_len | Output labels |
|---|---|---|
| (PER, 0, 2) | 6 | B-PER I-PER O O B-ORG O |
| (ORG, 4, 5) | 6 | (same output above) |
| round-trip | 6 | spans_in == spans_out: True |
Verify with seq_len=6, spans [(PER,0,2),(ORG,4,5)]
Why: bio_encode produces ['B-PER','I-PER','O','O','B-ORG','O']; bio_decode reconstructs exactly [('PER',0,2),('ORG',4,5)]. Round-trip identity confirms both functions are correct inverses.
Trade off
Comparison matrix
From Your Turn — Milestone 1: BIO encode/decode: every row here is a choice with a cost. Fill the Output labels column, then say which row you would actually pick and what you give up for it.
| Input span | seq_len | Output labels |
|---|---|---|
| (PER, 0, 2) | 6 | B-PER I-PER O O B-ORG O |
| (ORG, 4, 5) | 6 | (same output above) |
| round-trip | 6 | spans_in == spans_out: True |
Pattern
Predict first
The table runs: Greedy argmax | per-token independent | No — ignores transitions · Viterbi | joint sequence search | Yes — O(T * L^2)
In Your Turn — full program, given the rows so far: what is the next one — the row where Decoder is Difference??
Correct: Difference? | likely on random trans | Run code to verify
| Decoder | Output (label IDs) | Globally optimal? |
|---|---|---|
| Greedy argmax | per-token independent | No — ignores transitions |
| Viterbi | joint sequence search | Yes — O(T * L^2) |
| Difference? | likely on random trans | Run code to verify |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. With a random transition matrix, Viterbi's globally optimal path will typically differ from greedy label-by-label choices.
Worked example
import torch, torch.nn as nn
# 1. Build model
class TinyBertNER(nn.Module):
def __init__(self, V=200, H=64, heads=4, layers=2, L=7):
super().__init__()
self.embed = nn.Embedding(V, H, padding_idx=0)
self.pos = nn.Embedding(128, H)
enc = nn.TransformerEncoderLayer(
d_model=H, nhead=heads,
dim_feedforward=128, batch_first=True, dropout=0.0)
self.enc = nn.TransformerEncoder(enc, num_layers=layers)
self.head = nn.Linear(H, L)
def forward(self, ids):
B, T = ids.shape
x = self.embed(ids) + self.pos(
torch.arange(T).expand(B,-1))
return self.head(self.enc(x))
torch.manual_seed(42)
model = TinyBertNER()
ids = torch.randint(1, 200, (1, 6))
logits = model(ids) # (1, 6, 7)
greedy = logits.argmax(-1)[0] # (6,)
# 2. Viterbi (random transition for demo)
L = 7
trans = torch.randn(L, L)
e = logits[0] # (6, 7)
V_s = e[0].clone()
ptr = torch.zeros(6, L, dtype=torch.long)
for t in range(1, 6):
sc = V_s.unsqueeze(1) + trans
ptr[t] = sc.max(0).indices
V_s = sc.max(0).values + e[t]
cur, path = V_s.argmax().item(), []
path.append(cur)
for t in range(5, 0, -1):
cur = ptr[t, cur].item(); path.append(cur)
path.reverse()
print('Greedy:', greedy.tolist())
print('Viterbi:', path)
print('Same?', greedy.tolist() == path)| Decoder | Output (label IDs) | Globally optimal? |
|---|---|---|
| Greedy argmax | per-token independent | No — ignores transitions |
| Viterbi | joint sequence search | Yes — O(T * L^2) |
| Difference? | likely on random trans | Run code to verify |
Run the full program and check Same? — expect False with random transitions
Why: With a random transition matrix, Viterbi's globally optimal path will typically differ from greedy label-by-label choices. The diff reveals exactly where transition penalties redirect the path.
Comparison
Comparison matrix
From Your Turn — full program: refill the Globally optimal? column from what you know. The rest of the table is as it appeared.
| Decoder | Output (label IDs) | Globally optimal? |
|---|---|---|
| Greedy argmax | per-token independent | No — ignores transitions |
| Viterbi | joint sequence search | Yes — O(T * L^2) |
| Difference? | likely on random trans | Run code to verify |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — BIO Tagging — the sequence labeling contract · BERT + Linear Head — token-level logits · CRF Layer — globally consistent label sequences · Evaluation — two F1 metrics. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.