Lesson 17 proves that the generator and the recognizer describe the same class. It starts from the key observation that a leftmost sentential form splits into matched input and stack contents, then gives the three-state grammar-to-machine construction with its expansion and matching transitions and the push-order trap, the correctness proof by invariant, and why the stack forces leftmost derivations. It shows how to recover a parse tree from a computation, then gives the triple-variable encoding for the reverse direction with its never-dips-below condition, machine normalization, and the three rule schemas. It ends with the consequences: choosing whichever formalism makes a closure proof easier, cubic-time parsing via Chomsky normal form, and what is still missing before Lesson 18.
Subject: Theory of Computation · 112 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Theory of Computation · Lesson 17
The stack holds the unfinished derivation. That one observation converts a grammar into a machine — and, with more work, back again.
Objectives
Lesson 13 gave a generator and Lesson 16 a recognizer. This lesson proves they describe the same class. By the end you can:
Warm-up
Discussion prompt
Before we open Equivalence of PDAs & CFGs: without looking back, what was the main idea of Pushdown Automata, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Lesson 16 adds one unbounded stack to a finite automaton. It explains why a stack is exactly the right addition and what its last-in-first-out discipline still forbids, then covers transitions that read, pop, and push, with each part optional, and the bottom-marker trick for testing emptiness. It gives the formal seven-tuple and the six-tuple variant, configurations and the computation relation, and the final-state and empty-stack acceptance conventions with conversions in both directions. Designs follow for matched counts, palindromes, two kinds of bracket, and strict inequalities. It closes with the result that deterministic pushdown automata are strictly weaker, giving the standard witness language and explaining why the subset construction cannot be transferred.
Section
Section 1
Concept
The pairing that makes the chapter cohere: the grammars and the machines describe exactly the same languages.
the equivalence theorem — A language is generated by some context-free grammar if and only if it is recognized by some pushdown automaton.
\[ \exists G : L = L(G) \quad \iff \quad \exists P : L = L(P) \]
This is the analogue of Kleene's theorem from Lessons 8 and 9, one level up the hierarchy. Both directions need a construction, and as before one is much easier than the other.
Counterexample
Discussion prompt
The pairing that makes the chapter cohere: the grammars and the machines describe exactly the same languages.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Intuition
A grammar has inductive structure and a machine does not, exactly as in the regular case — but here the asymmetry lands differently.
| Direction | Difficulty | Reason |
|---|---|---|
| grammar to machine | easy | the stack can literally hold the derivation |
| machine to grammar | hard | a grammar must summarize every stack behaviour |
The easy direction needs only three transitions regardless of the grammar's size. The hard direction introduces a variable for every triple of state, state and stack symbol, which is where all the complexity lives.
Comparison
Comparison matrix
From Why the two directions differ in difficulty: refill the Difficulty column from what you know. The rest of the table is as it appeared.
| Direction | Difficulty | Reason |
|---|---|---|
| grammar to machine | easy | the stack can literally hold the derivation |
| machine to grammar | hard | a grammar must summarize every stack behaviour |
Concept
Everything in the easy direction follows from noticing what a leftmost derivation looks like at each moment.
In a leftmost derivation, the sentential form splits cleanly: a prefix of terminals already settled, then the leftmost variable, then the rest.
\[ S \Longrightarrow^{*} \underbrace{uAv}_{\text{sentential form}}, \quad u \in \Sigma^{*} \]
The terminal prefix is exactly what the machine has already matched against its input. The remainder — the variable and everything after it — is what the machine still owes, and that is what goes on the stack.
\[ \text{stack} = A\,v \quad\text{— the unmatched part} \]
Analogy
Discussion prompt
Explain The key observation by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Everything in the easy direction follows from noticing what a leftmost derivation looks like at each moment.
Ranking
Put in order
Put the moves of Watch a derivation and a stack move together into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The derivation replaces the variable by the rule's right-hand side; the stack pops the variable and pushes that same right-hand side.
Worked example
Before any construction, run the two side by side on a small grammar and see the correspondence.
\[ S \to aSb \;\mid\; \varepsilon \]
Start
Why: The derivation begins at the start variable; the stack holds exactly that variable.
\[ S \;\;\longleftrightarrow\;\; \text{stack } S \]
Apply the recursive rule
Why: The derivation replaces the variable by the rule's right-hand side; the stack pops the variable and pushes that same right-hand side.
\[ aSb \;\;\longleftrightarrow\;\; \text{stack } aSb \]
Match a terminal
Why: The machine reads an a from the input and pops the a from the top of the stack. The derivation's terminal prefix grows by one.
\[ \text{matched } a \;\;\longleftrightarrow\;\; \text{stack } Sb \]
Continue to the end
Why: Applying the empty rule pops the variable and pushes nothing; matching the b empties the stack.
\[ \text{matched } ab \;\;\longleftrightarrow\;\; \text{stack empty} \]
Verify the invariant held throughout
Why: At every stage, the input matched so far plus the stack contents spelled exactly the current sentential form. That correspondence is the construction, and the rest of this section only formalizes it.
\[ \text{matched} \;\cdot\; \text{stack} = \text{sentential form} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Watch a derivation and a stack move together", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: At every stage, the input matched so far plus the stack contents spelled exactly the current sentential form. That correspondence is the construction, and the rest of this section only formalizes it.
Concept
Two constructions, each with a proof, and each following the shape used in the regular chapters.
Sections 2 and 3 do the first with its proof; Sections 4 and 5 do the second. Section 6 collects what the equivalence buys.
Explain it
Discussion prompt
Explain The plan for both directions to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Two constructions, each with a proof, and each following the shape used in the regular chapters.
Ranking
Put in order
These are the steps of Reading a correspondence proof, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Both proofs here relate two very different objects, and they share a shape worth recognizing.
Step four is the one that takes the work. Showing every derivation gives a computation is usually easy; showing every computation gives a derivation needs the invariant stated carefully enough to run backwards.
Elimination
Eliminate the wrong options
In the grammar-to-machine construction, what do the stack contents represent?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A leftmost sentential form is a settled terminal prefix followed by the rest. The prefix is what the machine has already matched, so the stack holds precisely the remainder — the leftmost variable and everything after it.
Check
Think about what the stack is holding at each moment.
Check your understanding
In the grammar-to-machine construction, what do the stack contents represent?
Answer: A
Why: A leftmost sentential form is a settled terminal prefix followed by the rest. The prefix is what the machine has already matched, so the stack holds precisely the remainder — the leftmost variable and everything after it.
Intuition
The construction is not a theoretical curiosity — it is the top-down parsing algorithm, written as a machine.
A recursive-descent parser keeps its pending obligations on the call stack, expanding the leftmost unfinished nonterminal and matching terminals as it goes. That is exactly the machine built in the next section.
The difference is only that a real parser must choose which rule to expand by, using lookahead, while the machine guesses nondeterministically. Making that choice deterministic is what parser generators do, and why they need the restrictions of Lesson 16.
\[ \text{nondeterministic guess} \;\longleftrightarrow\; \text{lookahead in a real parser} \]
Section
Section 2
Concept
Remarkably, the construction needs no states at all beyond the bookkeeping ones. All the work happens on the stack.
The machine accepts by empty stack in the cleanest presentation, and by final state after the standard conversion of Lesson 16.
\[ \Gamma = V \cup \Sigma \cup \{Z_0\} \]
Concept
Every move belongs to one of three families, and the whole construction is these three lines.
| Situation | Move |
|---|---|
| a variable on top | pop it and push some right-hand side of one of its rules |
| a terminal on top | read that same symbol from the input and pop it |
| the bottom marker on top | accept, if the input is exhausted |
\[ \varepsilon, A \to \alpha \;\;\text{ for each rule } A \to \alpha; \qquad a, a \to \varepsilon \;\;\text{ for each terminal} \]
The first family consumes no input, which is why the machine can expand a whole derivation without reading anything. The second is where input is actually matched.
Trade off
Comparison matrix
From The three kinds of transition: every row here is a choice with a cost. Fill the Move column, then say which row you would actually pick and what you give up for it.
| Situation | Move |
|---|---|
| a variable on top | pop it and push some right-hand side of one of its rules |
| a terminal on top | read that same symbol from the input and pop it |
| the bottom marker on top | accept, if the input is exhausted |
Step zero
Discussion prompt
Convert a grammar to a machine — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Set up the stack
Answer:
Worked example
Apply the construction to the matched-counts grammar.
\[ S \to aSb \;\mid\; \varepsilon \]
Set up the stack
Why: A free move from the start state pushes the start variable above the bottom marker, then enters the simulation state.
\[ \varepsilon, \varepsilon \to S Z_0 \]
Add one expansion transition per rule
Why: Two rules, so two transitions. Each pops the variable and pushes the rule's right-hand side.
\[ \varepsilon, S \to aSb; \qquad \varepsilon, S \to \varepsilon \]
Add one matching transition per terminal
Why: Two terminals, so two transitions. Each reads a symbol and pops the same symbol.
\[ a, a \to \varepsilon; \qquad b, b \to \varepsilon \]
Add the acceptance check
Why: A free move popping the bottom marker enters the accepting state, which is reachable only when everything above has been consumed.
\[ \varepsilon, Z_0 \to \varepsilon \]
Verify by tracing a member and a non-member
Why: On aabb the machine expands twice, matches four symbols and empties the stack — accept. On aab, every expansion sequence leaves either a b unmatched on the stack or an unmatched input symbol, so no computation accepts. Both verdicts match the grammar.
\[ aabb \in L(P) \qquad aab \notin L(P) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Convert a grammar to a machine", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: On aabb the machine expands twice, matches four symbols and empties the stack — accept. On aab, every expansion sequence leaves either a b unmatched on the stack or an unmatched input symbol, so no computation accepts. Both verdicts match the grammar.
Estimation
Predict first
Follow every configuration, so the correspondence with the derivation is visible.
Commit before you compute: what does Trace the constructed machine in full come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify acceptance by matching the remaining terminals
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Two matching moves consume both b's and empty the stack down to the marker, which the acceptance move then pops.
Worked example
Follow every configuration, so the correspondence with the derivation is visible.
\[ \text{input } aabb \]
Initialize
Why: The stack holds the start variable above the marker; nothing has been read.
\[ (q, \; aabb, \; SZ_0) \]
Expand, then match
Why: A free move replaces the variable by the recursive right-hand side; then the leading a is matched and popped.
\[ (q, \; aabb, \; aSbZ_0) \vdash (q, \; abb, \; SbZ_0) \]
Expand and match again
Why: The same pair of moves consumes the second a.
\[ (q, \; abb, \; aSbbZ_0) \vdash (q, \; bb, \; SbbZ_0) \]
Apply the empty rule
Why: The variable is popped and nothing is pushed, leaving only terminals on the stack.
\[ (q, \; bb, \; bbZ_0) \]
Verify acceptance by matching the remaining terminals
Why: Two matching moves consume both b's and empty the stack down to the marker, which the acceptance move then pops. The input is exhausted at exactly that moment, so the string is accepted — and the derivation it mirrors is the leftmost one from the grammar.
\[ (q, \; \varepsilon, \; Z_0) \vdash (q_{\text{acc}}, \; \varepsilon, \; \varepsilon) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Trace the constructed machine in full", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Two matching moves consume both b's and empty the stack down to the marker, which the acceptance move then pops. The input is exhausted at exactly that moment, so the string is accepted — and the derivation it mirrors is the leftmost one from the grammar.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Add the expansion transition for a rule whose right-hand side has several symbols.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Pop the variable, then push the first symbol, then the second, then the third.
Add the expansion transition for a rule whose right-hand side has several symbols.
Why: Pop the variable, then push the first symbol, then the second, then the third.
Trap
Add the expansion transition for a rule whose right-hand side has several symbols.
Push the symbols one at a time, left to right
Why: Pop the variable, then push the first symbol, then the second, then the third.
\[ A \to XYZ \;\Longrightarrow\; \text{push } X, \text{ then } Y, \text{ then } Z \]
Inspect the stack
Why: The last symbol pushed sits on top, so the stack reads the right-hand side backwards. The machine will then try to match the last symbol first.
\[ \text{stack top} = Z \quad\text{but the input expects } X \]
Add the expansion transition for a rule whose right-hand side has several symbols.
Push the whole right-hand side in one move, leftmost on top
Why: The transition function returns a string to push, and the convention is that its leftmost symbol ends up on top.
\[ \varepsilon, A \to XYZ \]
Inspect the stack
Why: The first symbol of the right-hand side is now on top, so it is the first thing matched — which is what a leftmost derivation requires.
\[ \text{stack top} = X \;\checkmark \]
Note why the convention exists
Why: Lesson 16 defined the pushed component as a string precisely so a whole right-hand side goes on in one move, in the order that makes this construction work.
Notation
Annotate
From Trap: pushing the right-hand side in the wrong order — read this one piece at a time. What is each part doing?
On: \( A \to XYZ \;\Longrightarrow\; \text{push } X, \text{ then } Y, \text{ then } Z \)
Missing information
Discussion prompt
A second conversion, where the stack holds several pending obligations at once.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The start variable is pushed, then replaced by two variables in one move — the first of them on top.
Worked example
A second conversion, where the stack holds several pending obligations at once.
\[ S \to AB, \qquad A \to a, \qquad B \to b \]
Set up and expand the start variable
Why: The start variable is pushed, then replaced by two variables in one move — the first of them on top.
\[ (q, ab, SZ_0) \vdash (q, ab, ABZ_0) \]
Expand the leftmost variable
Why: Only the top symbol is visible, so the first variable is the one expanded. Its rule replaces it with a terminal.
\[ (q, ab, aBZ_0) \]
Match, then expand again
Why: The terminal on top is matched against the input; the second variable then becomes visible and is expanded.
\[ (q, b, BZ_0) \vdash (q, b, bZ_0) \]
Match and accept
Why: The last terminal matches, the marker is popped, and the input is exhausted at the same moment.
Verify the pending obligation was held correctly
Why: Between the first expansion and the second, the second variable sat on the stack below the first — invisible but preserved. That is exactly the role of a stack here: holding obligations in the order they will be needed.
\[ ab \in L(P) = L(G) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Convert a grammar with two variables", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Between the first expansion and the second, the second variable sat on the stack below the first — invisible but preserved. That is exactly the role of a stack here: holding obligations in the order they will be needed.
Intuition
It is surprising that a machine of three states can simulate a grammar of any size. The reason is worth stating.
A finite automaton keeps everything it knows in its state, so more to remember means more states. This machine keeps everything on the stack, and the stack is unbounded — so the state count never has to grow.
The grammar's size shows up instead in the transition count: one expansion transition per rule. Complexity moved from the states to the transitions and the stack alphabet.
\[ |Q| = 3, \qquad |\delta| = |R| + |\Sigma| + 2 \]
Pattern
Four lines, applicable to any grammar without modification.
Note what the construction does not need: no normal form, no removal of epsilon-rules, no analysis of the grammar at all. It works on any grammar as written.
Prediction
Predict first
A grammar has 12 variables, 40 rules and 5 terminals. How many states does the constructed machine have?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Three, regardless of the grammar's size
Why: All the grammar's structure lives on the stack, not in the states. The construction uses one state to run the simulation plus two for setup and acceptance, whatever the grammar looks like.
Check
Count what the construction produces.
Check your understanding
A grammar has 12 variables, 40 rules and 5 terminals. How many states does the constructed machine have?
Answer: A
Why: All the grammar's structure lives on the stack, not in the states. The construction uses one state to run the simulation plus two for setup and acceptance, whatever the grammar looks like.
Section
Section 3
Concept
The proof rests on a single claim relating the machine's configuration to the grammar's sentential forms.
At any point in a computation, the input already matched, followed by the stack contents read top to bottom, is exactly a sentential form of a leftmost derivation from the start variable.
\[ S \;\Longrightarrow^{*}_{\mathrm{lm}}\; u\,\gamma \quad\text{where } u \text{ is matched and } \gamma \text{ is the stack} \]
Prove that, and both containments follow immediately by looking at the two ends of a computation.
Step zero
Discussion prompt
Prove the invariant — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Expansion case
Answer:
Worked example
Induction on the number of moves, with two cases matching the two transition families.
Base case
Why: Before any move, nothing is matched and the stack holds the start variable alone. Their concatenation is the start variable, which is a sentential form in zero steps.
\[ u = \varepsilon, \; \gamma = S \;\Rightarrow\; u\gamma = S \]
Expansion case
Why: A free move pops a variable and pushes a right-hand side. The matched prefix is unchanged, and the concatenation changes by exactly one rule application to the leftmost variable — which is a leftmost derivation step.
\[ uA\beta \;\Longrightarrow_{\mathrm{lm}}\; u\alpha\beta \]
Matching case
Why: A reading move consumes one terminal and pops the same symbol. The matched prefix grows by that symbol and the stack loses it, so the concatenation is unchanged.
\[ u\,a\,\beta \;: \; \text{matched grows}, \; \text{stack shrinks}, \; \text{product fixed} \]
Note the leftmost condition is maintained
Why: Expansion always acts on the top of the stack, which is the leftmost symbol of the unmatched part. Since everything to its left is already terminal, it is the leftmost variable.
Verify both containments follow
Why: An accepting computation ends with an empty stack and the whole input matched, so the invariant gives a leftmost derivation of that input. Conversely, any leftmost derivation can be replayed as a computation, expanding when the top is a variable and matching when it is a terminal. Both directions hold, so the languages are equal.
\[ L(P) = L(G) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove the invariant", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: An accepting computation ends with an empty stack and the whole input matched, so the invariant gives a leftmost derivation of that input. Conversely, any leftmost derivation can be replayed as a computation, expanding when the top is a variable and matching when it is a terminal. Both directions hold, so the languages are equal.
Intuition
The proof used leftmost derivations specifically, and no other discipline would have worked.
A stack exposes only its top symbol. The top of the stack is the leftmost unmatched symbol, so the only variable the machine can act on is the leftmost one — which forces the derivation to be leftmost.
A rightmost derivation would need to expand the last variable, which sits at the bottom of the stack and is invisible. So the construction and the discipline are not independent choices: the stack picks the discipline.
\[ \text{stack top} \;=\; \text{leftmost unmatched symbol} \]
Sorting
Sort into buckets
These are the pieces of Equivalence of PDAs & CFGs, out of order. Put each one back under the part of the lesson it belongs to.
Concept
The constructed machine is heavily nondeterministic, and it is worth being precise about where.
Every variable with several alternatives gives several applicable expansion transitions in the same configuration. The machine guesses which rule the derivation used.
\[ A \to \alpha_1 \;\mid\; \alpha_2 \;\Rightarrow\; \text{two moves available} \]
Matching transitions involve no choice at all — the input symbol and the stack top must agree, or the branch dies. So all the guessing is about which rule, never about what to match.
That is exactly the choice a real parser resolves with lookahead, and exactly the choice that cannot always be resolved, which is why Lesson 16's deterministic machines are weaker.
Hypothesis
Predict first
Watch a wrong guess die is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Guess the empty rule immediately
Why: The machine pops the start variable and pushes nothing, leaving only the bottom marker.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Run the constructed machine down a branch that guesses badly, to see the failure mode.
\[ S \to aSb \;\mid\; \varepsilon, \qquad \text{input } ab \]
Guess the empty rule immediately
Why: The machine pops the start variable and pushes nothing, leaving only the bottom marker.
\[ (q, ab, SZ_0) \vdash (q, ab, Z_0) \]
Try to continue
Why: The only remaining move pops the marker into the accepting state — but the input still holds two unread symbols.
See the branch fail
Why: Acceptance requires the input to be exhausted, and it is not. This computation does not accept, and no move extends it usefully.
\[ (q_{\text{acc}}, ab, \varepsilon) \;: \text{ input not exhausted} \]
Find the branch that succeeds
Why: Guessing the recursive rule first pushes the three-symbol right-hand side, after which the a matches, the empty rule fires, and the b matches.
Verify acceptance depends on the good branch only
Why: One accepting computation exists, so the string is accepted — the failing branch places no constraint, exactly as Lesson 5 established. The machine's nondeterminism is doing precisely the work of searching for the right derivation.
\[ ab \in L(P) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Watch a wrong guess die", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: One accepting computation exists, so the string is accepted — the failing branch places no constraint, exactly as Lesson 5 established. The machine's nondeterminism is doing precisely the work of searching for the right derivation.
Concept
The construction is faithful in a stronger sense than mere language equality, and that strengthening is what makes it useful for parsing.
| Grammar object | Machine object |
|---|---|
| a leftmost derivation | an accepting computation |
| a rule application | an expansion transition |
| a terminal in the yield | a matching transition |
| a parse tree | the whole computation, up to scheduling |
So a machine that records which expansion it chose can output the parse tree, not just a verdict. That is exactly what a recursive-descent parser returns.
Ranking
Put in order
Put the moves of Recover a parse tree from a computation into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Every free move that pops a variable corresponds to one rule application.
Worked example
Use the correspondence in the direction that matters in practice.
Record each expansion
Why: Every free move that pops a variable corresponds to one rule application. Log the rule as it happens.
Note the order
Why: Because the derivation is leftmost, the logged rules arrive in leftmost order — which determines the tree uniquely.
Rebuild the tree
Why: Start with the start variable as the root and apply the logged rules in order, always to the leftmost unexpanded variable.
Check the matching moves need no logging
Why: They consume terminals and create leaves, but the leaves are determined by the rules already logged. So only expansions carry information.
Verify on the earlier trace
Why: The trace of the four-symbol string logged the recursive rule twice and the empty rule once, in that order. Replaying them leftmost from the start variable rebuilds the tree whose yield is that string, confirming nothing was lost.
\[ \text{expansions logged} \;\longrightarrow\; \text{unique leftmost derivation} \;\longrightarrow\; \text{tree} \ \checkmark \]
Notation
Annotate
From Recover a parse tree from a computation — read this one piece at a time. What is each part doing?
On: \( \text{expansions logged} \;\longrightarrow\; \text{unique leftmost derivation} \;\longrightarrow\; \text{tree} \ \checkmark \)
Check
Recall which part of the invariant each transition family changes.
Check your understanding
In the correctness proof, what does a matching transition do to the concatenation of the matched input and the stack?
Answer: A
Why: The transition consumes a terminal from the input and pops the identical symbol from the stack, so the symbol simply crosses from one part of the concatenation to the other. That is why matching moves preserve the sentential form while expansions advance it.
Section
Section 4
Concept
The reverse direction is harder because a grammar has no stack, so it must express stack behaviour through its variables.
The insight is to focus on the net effect of a stretch of computation, rather than on individual moves. Specifically: what does it take for the machine to go from one state to another while removing exactly one symbol from the stack?
If that question can be answered by a variable, then the whole computation decomposes, because every accepting computation empties the stack one symbol at a time.
\[ A_{pXq} \;: \; \text{from } p \text{ to } q, \text{ net-popping } X \]
Concept
triple variable — A grammar variable indexed by a start state, a stack symbol and an end state, generating exactly the input strings that take the machine between those states while net-removing that symbol.
There is one variable per triple, so the grammar has as many variables as the number of states squared times the size of the stack alphabet. That is the source of the construction's size.
\[ |V| = |Q|^{2} \cdot |\Gamma| \]
The start variable is the one for going from the start state to an accepting state while removing the initial stack symbol — which is exactly acceptance by empty stack.
Intuition
The choice of unit is what makes the decomposition work, and it is worth seeing why a cruder unit fails.
During the stretch, the stack may grow enormously — but it must return to exactly one symbol lower than it started, and it may never dip below that level in between. That makes the stretch self-contained: nothing below the symbol is touched.
Self-containment is what allows two stretches to be concatenated without interfering. A unit that allowed the stack to dip below its starting level would let stretches interact, and no context-free rule could describe that.
\[ \text{never dips below} \;\Rightarrow\; \text{stretches compose independently} \]
Concept
The grammar's rules come from decomposing a stretch in the two possible ways.
The second family is where the branching comes from, and it is why the construction is usually presented after converting the machine to a restricted form.
Concept
The construction is far simpler if the machine is first normalized, much as grammars were in Lesson 15.
Every machine can be put in this form by adding states and splitting transitions, and the language is unchanged. With it, the second rule family reduces to a single shape.
\[ A_{pXq} \to a \; A_{rYs} \; b \]
Step zero
Discussion prompt
Normalize a machine — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Force a single accepting state
Answer:
Worked example
Apply the three conditions to a machine that violates all of them.
Force a single accepting state
Why: Add a fresh accepting state, reachable by a free move from every old accepting state, and demote the old ones.
Force the stack to empty before accepting
Why: Add a draining state that pops any symbol without consuming input, placed between the old accepting states and the new one.
\[ \varepsilon, X \to \varepsilon \;\text{ for every } X \]
Split transitions that neither push nor pop
Why: Replace each by two moves through a fresh state: one pushing a dummy symbol, one popping it.
\[ a, \varepsilon \to \varepsilon \;\Longrightarrow\; a, \varepsilon \to D \;\text{ then }\; \varepsilon, D \to \varepsilon \]
Split transitions that both pop and push several
Why: Replace each by a chain: one pop, then one push per symbol of the pushed string, each through a fresh state.
Verify the language is unchanged
Why: Every added move consumes no input and every split preserves the net stack effect, so each original computation corresponds to exactly one normalized computation with the same input and the same verdict.
\[ L(P) = L(P') \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Normalize a machine", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Every added move consumes no input and every split preserves the net stack effect, so each original computation corresponds to exactly one normalized computation with the same input and the same verdict.
Commit first
Predict first
The variable indexed by states p and q and stack symbol X generates which strings?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: Those taking the machine from p to q while net-removing X and never dipping below it
Why: The stack may grow arbitrarily during the stretch, but it must end exactly one symbol lower and must never go below that level in between. That self-containment is what lets stretches be concatenated without interference.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Think about what the stack does during a stretch.
Check your understanding
The variable indexed by states p and q and stack symbol X generates which strings?
Answer: A
Why: The stack may grow arbitrarily during the stretch, but it must end exactly one symbol lower and must never go below that level in between. That self-containment is what lets stretches be concatenated without interference.
Intuition
The grammar-to-machine construction used three states; this one uses a variable per triple. The asymmetry has a clear cause.
A stack is a single unbounded object, and a machine can consult only its top. A grammar has no such object, so every fact about stack behaviour must be encoded in the finitely many variable names.
The triples are the smallest encoding that suffices: which state the stretch starts in, which symbol it removes, and which state it ends in. Dropping any component makes the decomposition unsound.
\[ |Q|^{2}|\Gamma| \text{ variables} \;\longleftrightarrow\; \text{one stack} \]
Section
Section 5
Concept
The reverse construction is expensive, and knowing the cost explains why it is rarely run in practice.
| Quantity | Size |
|---|---|
| variables | states squared, times the stack alphabet |
| rules from the push-pop schema | one per matched push-pop pair, times two states |
| rules from the splitting schema | variables times the state count |
| useless variables produced | typically most of them |
A ten-state machine with a five-symbol stack alphabet already gives five hundred variables, and the great majority generate nothing. Running the Lesson 15 cleanup afterwards is not optional in practice.
\[ |V| = |Q|^{2}|\Gamma| \]
Estimation
Predict first
Make the cost concrete before carrying out a conversion.
Commit before you compute: what does Count the variables for a small machine come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the cleanup is worth running
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Applying the generating and reachable closures of Lesson 15 removes every variable with no stretch, typically leaving a handful.
Worked example
Make the cost concrete before carrying out a conversion.
Take a machine
Why: Three states, and a stack alphabet of two symbols plus the bottom marker.
\[ |Q| = 3, \qquad |\Gamma| = 3 \]
Apply the formula
Why: One variable per ordered pair of states and each stack symbol.
\[ |V| = 3^{2} \cdot 3 = 27 \]
Estimate how many are useful
Why: A variable is useful only if the machine can actually get from the first state to the second while removing that symbol. For most triples no such stretch exists.
Note the practical consequence
Why: The generated grammar is unreadable and mostly dead. Its purpose is to establish the theorem, not to be used.
Verify the cleanup is worth running
Why: Applying the generating and reachable closures of Lesson 15 removes every variable with no stretch, typically leaving a handful. That is the grammar worth looking at, and it is usually recognizable as one you could have written directly.
\[ 27 \text{ variables} \;\longrightarrow\; \text{a handful, after cleanup} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Count the variables for a small machine", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Applying the generating and reachable closures of Lesson 15 removes every variable with no stretch, typically leaving a handful. That is the grammar worth looking at, and it is usually recognizable as one you could have written directly.
Concept
With the machine normalized, the grammar is generated by three schemas.
| Schema | When |
|---|---|
| push then later pop | a push move and a matching pop move exist |
| split a stretch | a stretch passes through an intermediate state |
| empty stretch | the start and end states coincide |
\[ A_{pXq} \to a\,A_{rYs}\,b, \qquad A_{pXq} \to A_{pXr}A_{rXq}, \qquad A_{pXp} \to \varepsilon \]
The third schema is where derivations terminate, and the first is where input is actually generated.
Comparison
Comparison matrix
From The three rule schemas: refill the When column from what you know. The rest of the table is as it appeared.
| Schema | When |
|---|---|
| push then later pop | a push move and a matching pop move exist |
| split a stretch | a stretch passes through an intermediate state |
| empty stretch | the start and end states coincide |
Estimation
Predict first
Convert a small normalized machine, one schema at a time.
Commit before you compute: what does Build the grammar from a machine come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the grammar generates the machine's language
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The rules produce a's on the left and b's on the right in matching numbers, terminating with the empty string — exactly the matched-counts language the machine recognizes.
Worked example
Convert a small normalized machine, one schema at a time.
\[ p: \; a, \varepsilon \to X \qquad q: \; b, X \to \varepsilon \]
List the variables
Why: Two states and one stack symbol beyond the marker, so there are four triple variables in principle, though not all will be useful.
\[ A_{pXp}, \; A_{pXq}, \; A_{qXp}, \; A_{qXq} \]
Apply the push-then-pop schema
Why: The push move reads an a and puts X on the stack; the pop move reads a b and removes it. Together they bracket a stretch.
\[ A_{pXq} \to a \; A_{pXq} \; b \]
Apply the empty-stretch schema
Why: Where the start and end states coincide, the empty string is generated.
\[ A_{pXp} \to \varepsilon \]
Identify the start variable
Why: It records going from the start state to the accepting state while removing the initial marker.
Verify the grammar generates the machine's language
Why: The rules produce a's on the left and b's on the right in matching numbers, terminating with the empty string — exactly the matched-counts language the machine recognizes. Deriving aabb and failing to derive aab confirms it.
\[ L(G) = \{a^{n}b^{n}\} = L(P) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Build the grammar from a machine", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The rules produce a's on the left and b's on the right in matching numbers, terminating with the empty string — exactly the matched-counts language the machine recognizes. Deriving aabb and failing to derive aab confirms it.
Concept
The proof is two inductions, mirroring the two directions of the language equality.
The second induction is the delicate one, and the never-dips-below condition is exactly what makes the decomposition point well defined.
\[ \text{first return to the starting level} \;=\; \text{the split point} \]
Step zero
Discussion prompt
Decompose a computation to find its derivation — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Take an accepting computation
Answer:
Worked example
Run the harder induction on a concrete computation, to see the decomposition happen.
Take an accepting computation
Why: It starts with the marker on the stack and ends with the stack empty, so it net-removes exactly that one symbol.
Find the first move
Why: In the normalized machine it either pops the marker outright, or pushes a symbol that must be removed later.
In the pushing case, find the matching pop
Why: Follow the computation until the stack first returns to its starting level. That move is the matching pop, and it is unique.
Split into two stretches
Why: The part strictly between the push and its matching pop is one self-contained stretch; the part after the pop is another. Each is shorter than the whole.
\[ A_{pXq} \to a \; A_{rYs} \; b \quad\text{applies here} \]
Verify the induction terminates
Why: Both stretches are strictly shorter than the original computation, so the induction is well founded and bottoms out at the empty-stretch schema. Every accepting computation therefore yields a derivation, completing the harder containment.
\[ \text{both pieces shorter} \;\Rightarrow\; \text{induction well founded} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Decompose a computation to find its derivation", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Both stretches are strictly shorter than the original computation, so the induction is well founded and bottoms out at the empty-stretch schema. Every accepting computation therefore yields a derivation, completing the harder containment.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Decompose a stretch that goes from one state to another while removing one symbol.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Choose any convenient moment when the depth matches, and cut the stretch there.
Decompose a stretch that goes from one state to another while removing one symbol.
Why: Choose any convenient moment when the depth matches, and cut the stretch there.
Trap
Decompose a stretch that goes from one state to another while removing one symbol.
Split at any point where the stack returns to its starting height
Why: Choose any convenient moment when the depth matches, and cut the stretch there.
See what goes wrong
Why: If the stack dipped below the starting level in between, the symbol being removed was already gone, so the two pieces are not independent — one of them touched what lay beneath.
\[ \text{dips below} \;\Rightarrow\; \text{pieces interact} \]
Decompose a stretch that goes from one state to another while removing one symbol.
Split at the FIRST return to the starting level
Why: Taking the earliest such moment guarantees the stack never went below it beforehand, so the symbol was untouched throughout the first piece.
Confirm both pieces are self-contained
Why: The first piece never sees below the symbol it is bracketed by, and the second starts fresh at the lower level. Neither can interfere with the other.
\[ \text{first return} \;\Rightarrow\; \text{both pieces self-contained} \ \checkmark \]
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Constraint
Discussion prompt
Run The machine-to-grammar recipe with this step confiscated:
Add a rule for each matched push-pop pair, bracketing an intermediate stretch.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Collected, though it is rarely carried out by hand on anything large.
The grammar produced is large and unreadable, and usually contains many useless variables. Running the cleanup pass of Lesson 15 afterwards is normal.
Edge cases
Discussion prompt
The machine-to-grammar recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Collected, though it is rarely carried out by hand on anything large.
Prediction
Predict first
Why must a stretch be split at the FIRST moment the stack returns to its starting level?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: A later split could allow the stack to have dipped below, so the pieces would not be independent
Why: The triple variables describe stretches that never go below their starting level, which is what makes them composable. Splitting at a later return would permit a dip in between, and the piece would no longer describe a self-contained stretch.
Check
Think about where the decomposition must cut.
Check your understanding
Why must a stretch be split at the FIRST moment the stack returns to its starting level?
Answer: A
Why: The triple variables describe stretches that never go below their starting level, which is what makes them composable. Splitting at a later return would permit a dip in between, and the piece would no longer describe a self-contained stretch.
Section
Section 6
Concept
With both constructions in hand, either formalism may be used for any argument about the class.
\[ \text{CFG} \;\equiv\; \text{PDA} \]
Ranking
Put in order
Put the moves of Use the equivalence to prove a closure property into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Grammars, because a union is one extra rule and a machine construction would need care about disjoint stacks.
Worked example
Show the context-free languages are closed under union, using whichever side is easier.
Choose the formalism
Why: Grammars, because a union is one extra rule and a machine construction would need care about disjoint stacks.
Take grammars for both languages
Why: By the theorem, each language has a grammar. Rename variables so the two sets are disjoint.
Add a fresh start variable
Why: One alternative per grammar, so a derivation commits to one branch immediately.
\[ S \to S_1 \;\mid\; S_2 \]
Read off the language
Why: A derivation goes entirely through one branch, so the generated strings are exactly those of one grammar or the other.
Verify the machine side would have been harder
Why: A machine construction would need a fresh start state with free moves into both machines, plus care that neither machine's stack symbols interfered with the other's — three conditions instead of one rule. The equivalence let the easier side be chosen, which is the whole point.
\[ L_1 \cup L_2 \text{ context-free} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Use the equivalence to prove a closure property", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: A machine construction would need a fresh start state with free moves into both machines, plus care that neither machine's stack symbols interfered with the other's — three conditions instead of one rule. The equivalence let the easier side be chosen, which is the whole point.
Concept
Two classes are fully described, each by a generator and a recognizer proved equivalent.
| Class | Generator | Recognizer | Equivalence proved in |
|---|---|---|---|
| regular | regular expression | finite automaton | Lessons 8 and 9 |
| context-free | context-free grammar | pushdown automaton | this lesson |
The pattern is deliberate, and it repeats once more: Lesson 20 introduces a machine, and Lesson 23 argues that it captures the informal notion of an algorithm — the analogue, at the top of the hierarchy, of these equivalence theorems.
Explain it
Discussion prompt
Explain The hierarchy, now with recognizers on both rows to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Two classes are fully described, each by a generator and a recognizer proved equivalent.
Concept
The equivalence also settles a practical question: given a grammar and a string, can membership be decided?
Yes. Convert the grammar to Chomsky normal form by Lesson 15, then fill a table indexed by substrings, combining pairs using the binary rules. The algorithm runs in time cubic in the string's length.
\[ O\big(n^{3} \cdot |G|\big) \]
The normal form is essential: the table combines exactly two pieces per entry, which needs every rule to have exactly two variables on the right. That is why Lesson 15 came before this one.
Analogy
Discussion prompt
Explain The parsing problem by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The equivalence also settles a practical question: given a grammar and a string, can membership be decided?
Intuition
Two equivalent descriptions, and the choice is never arbitrary in practice.
| Task | Formalism |
|---|---|
| specifying a language | grammar |
| proving closure under union, concatenation, star | grammar |
| proving a language is deterministic | machine |
| implementing a recognizer | machine, derived from the grammar |
| recovering structure from a string | grammar, via the parse tree |
The third row is the one that needs the machine: determinism is a property of computations, and grammars have no notion of one.
Trade off
Comparison matrix
From Where each formalism wins: every row here is a choice with a cost. Fill the Formalism column, then say which row you would actually pick and what you give up for it.
| Task | Formalism |
|---|---|
| specifying a language | grammar |
| proving closure under union, concatenation, star | grammar |
| proving a language is deterministic | machine |
| implementing a recognizer | machine, derived from the grammar |
| recovering structure from a string | grammar, via the parse tree |
Concept
The class is now described from the inside by two equivalent formalisms. Nothing yet describes it from the outside.
Every technique here proves membership by exhibiting an object. To prove a language is not context-free, none of it helps — exactly the situation Lesson 9 left the regular languages in.
Lesson 18 supplies the missing tool, and its shape is already visible: normal form bounds a parse tree's height by the string's length, a tall tree must repeat a variable on some path, and a repeated variable can be pumped.
\[ \text{long string} \;\Rightarrow\; \text{tall tree} \;\Rightarrow\; \text{repeated variable} \;\Rightarrow\; \text{pumping} \]
Counterexample
Discussion prompt
The class is now described from the inside by two equivalent formalisms. Nothing yet describes it from the outside.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Every technique here proves membership by exhibiting an object. To prove a language is not context-free, none of it helps — exactly the situation Lesson 9 left the regular languages in.
Elimination
Eliminate the wrong options
You want to prove the context-free languages are closed under concatenation. Which formalism gives the shorter proof?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: One rule sending a fresh start variable to the two old ones in sequence does it, and the correctness argument is a single sentence about derivations splitting. The machine construction needs the two stacks kept from interfering, which takes real care.
Check
Think about which side each property is naturally stated on.
Check your understanding
You want to prove the context-free languages are closed under concatenation. Which formalism gives the shorter proof?
Answer: A
Why: One rule sending a fresh start variable to the two old ones in sequence does it, and the correctness argument is a single sentence about derivations splitting. The machine construction needs the two stacks kept from interfering, which takes real care.
Intuition
Comparing this lesson with Lessons 8 and 9 shows the same argument shape twice, which makes both easier to remember.
| Regular class | Context-free class | |
|---|---|---|
| easy direction | expression to machine, by gadgets | grammar to machine, by stack simulation |
| what makes it easy | the expression is inductive | the derivation is inductive |
| hard direction | machine to expression, by state elimination | machine to grammar, by triple variables |
| what makes it hard | a graph has no inductive structure | a stack must be summarized finitely |
In both cases the generator has structure to recurse on and the recognizer does not. That asymmetry, not any accident of the models, is what makes one direction short and the other long.
Comparison
Comparison matrix
From Two equivalence theorems, one pattern: refill the Context-free class column from what you know. The rest of the table is as it appeared.
| Regular class | Context-free class | |
|---|---|---|
| easy direction | expression to machine, by gadgets | grammar to machine, by stack simulation |
| what makes it easy | the expression is inductive | the derivation is inductive |
| hard direction | machine to expression, by state elimination | machine to grammar, by triple variables |
| what makes it hard | a graph has no inductive structure | a stack must be summarized finitely |
Ranking
Put in order
These are the steps of Choosing a direction to convert, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Both conversions exist, but only one is routinely worth doing.
The third line is honest advice. The reverse construction is proof machinery, and a grammar written from an understanding of the language is invariably clearer than one generated from a machine.
Real world
Discussion prompt
Outside this lesson: where does Equivalence of PDAs & CFGs actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of Choosing a direction to convert is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
Lesson 17 proves that the generator and the recognizer describe the same class. It starts from the key observation that a leftmost sentential form splits into matched input and stack contents, then gives the three-state grammar-to-machine construction with its expansion and matching transitions and the push-order trap, the correctness proof by invariant, and why the stack forces leftmost derivations. It shows how to recover a parse tree from a computation, then gives the triple-variable encoding for the reverse direction with its never-dips-below condition, machine normalization, and the three rule schemas. It ends with the consequences: choosing whichever formalism makes a closure proof easier, cubic-time parsing via Chomsky normal form, and what is still missing before Lesson 18.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The Theorem · Grammar to Machine · Why the First Construction Works · Machine to Grammar: The Idea · Machine to Grammar: The Construction · Consequences. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
The generator and the recognizer are now proved equivalent, completing the pairing this chapter set out to establish.
| Situation | Move |
|---|---|
| a grammar to run | three states, expand on variables, match on terminals |
| a right-hand side to push | leftmost symbol on top, in one move |
| a machine to describe as a grammar | normalize first, then one variable per triple |
| a closure property to prove | grammars, unless the property is about computations |
| a language to show is not context-free | neither — that is Lesson 18 |
Lesson 18 supplies the pumping lemma for this class, built on the tree-height bound that Chomsky normal form made available.
Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.