Equivalence of PDAs & CFGs

Lesson 17 proves that the generator and the recognizer describe the same class. It starts from the key observation that a leftmost sentential form splits into matched input and stack contents, then gives the three-state grammar-to-machine construction with its expansion and matching transitions and the push-order trap, the correctness proof by invariant, and why the stack forces leftmost derivations. It shows how to recover a parse tree from a computation, then gives the triple-variable encoding for the reverse direction with its never-dips-below condition, machine normalization, and the three rule schemas. It ends with the consequences: choosing whichever formalism makes a closure proof easier, cubic-time parsing via Chomsky normal form, and what is still missing before Lesson 18.

Subject: Theory of Computation · 112 slides · symbolic lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Equivalence of PDAs and CFGs

Title

Theory of Computation · Lesson 17

The stack holds the unfinished derivation. That one observation converts a grammar into a machine — and, with more work, back again.

2. What you will be able to do

Objectives

Lesson 13 gave a generator and Lesson 16 a recognizer. This lesson proves they describe the same class. By the end you can:

  1. State the theorem in both directions and say which is the easier one.
  2. Convert any grammar into a machine that simulates leftmost derivations on its stack.
  1. Prove that construction correct, by relating stack contents to sentential forms.
  2. Explain the triple-variable idea behind the reverse construction.
  1. Carry out the machine-to-grammar conversion on a small machine.
  2. Use the equivalence to settle questions that are awkward in one formalism and easy in the other.

3. What survived from Pushdown Automata?

Warm-up

Discussion prompt

Before we open Equivalence of PDAs & CFGs: without looking back, what was the main idea of Pushdown Automata, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Lesson 16 adds one unbounded stack to a finite automaton. It explains why a stack is exactly the right addition and what its last-in-first-out discipline still forbids, then covers transitions that read, pop, and push, with each part optional, and the bottom-marker trick for testing emptiness. It gives the formal seven-tuple and the six-tuple variant, configurations and the computation relation, and the final-state and empty-stack acceptance conventions with conversions in both directions. Designs follow for matched counts, palindromes, two kinds of bracket, and strict inequalities. It closes with the result that deterministic pushdown automata are strictly weaker, giving the standard witness language and explaining why the subset construction cannot be transferred.

4. The Theorem

Section

Section 1

5. The statement

Concept

The pairing that makes the chapter cohere: the grammars and the machines describe exactly the same languages.

the equivalence theorem — A language is generated by some context-free grammar if and only if it is recognized by some pushdown automaton.

\[ \exists G : L = L(G) \quad \iff \quad \exists P : L = L(P) \]

This is the analogue of Kleene's theorem from Lessons 8 and 9, one level up the hierarchy. Both directions need a construction, and as before one is much easier than the other.

6. Break it if you can: The statement

Counterexample

Discussion prompt

The pairing that makes the chapter cohere: the grammars and the machines describe exactly the same languages.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Why the two directions differ in difficulty

Intuition

A grammar has inductive structure and a machine does not, exactly as in the regular case — but here the asymmetry lands differently.

DirectionDifficultyReason
grammar to machineeasythe stack can literally hold the derivation
machine to grammarharda grammar must summarize every stack behaviour

The easy direction needs only three transitions regardless of the grammar's size. The hard direction introduces a variable for every triple of state, state and stack symbol, which is where all the complexity lives.

8. Fill in: Difficulty for Why the two directions differ in difficulty

Comparison

Comparison matrix

From Why the two directions differ in difficulty: refill the Difficulty column from what you know. The rest of the table is as it appeared.

DirectionDifficultyReason
grammar to machineeasythe stack can literally hold the derivation
machine to grammarharda grammar must summarize every stack behaviour

9. The key observation

Concept

Everything in the easy direction follows from noticing what a leftmost derivation looks like at each moment.

In a leftmost derivation, the sentential form splits cleanly: a prefix of terminals already settled, then the leftmost variable, then the rest.

\[ S \Longrightarrow^{*} \underbrace{uAv}_{\text{sentential form}}, \quad u \in \Sigma^{*} \]

The terminal prefix is exactly what the machine has already matched against its input. The remainder — the variable and everything after it — is what the machine still owes, and that is what goes on the stack.

\[ \text{stack} = A\,v \quad\text{— the unmatched part} \]

10. By analogy: The key observation

Analogy

Discussion prompt

Explain The key observation by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Everything in the easy direction follows from noticing what a leftmost derivation looks like at each moment.

11. What has to happen first: Watch a derivation and a stack move together

Ranking

Put in order

Put the moves of Watch a derivation and a stack move together into the order they have to happen.

  1. Apply the recursive rule
  2. Match a terminal
  3. Continue to the end
  4. Verify the invariant held throughout

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The derivation replaces the variable by the rule's right-hand side; the stack pops the variable and pushes that same right-hand side.

12. Watch a derivation and a stack move together

Worked example

Before any construction, run the two side by side on a small grammar and see the correspondence.

\[ S \to aSb \;\mid\; \varepsilon \]

Start

Why: The derivation begins at the start variable; the stack holds exactly that variable.

\[ S \;\;\longleftrightarrow\;\; \text{stack } S \]

Apply the recursive rule

Why: The derivation replaces the variable by the rule's right-hand side; the stack pops the variable and pushes that same right-hand side.

\[ aSb \;\;\longleftrightarrow\;\; \text{stack } aSb \]

Match a terminal

Why: The machine reads an a from the input and pops the a from the top of the stack. The derivation's terminal prefix grows by one.

\[ \text{matched } a \;\;\longleftrightarrow\;\; \text{stack } Sb \]

Continue to the end

Why: Applying the empty rule pops the variable and pushes nothing; matching the b empties the stack.

\[ \text{matched } ab \;\;\longleftrightarrow\;\; \text{stack empty} \]

Verify the invariant held throughout

Why: At every stage, the input matched so far plus the stack contents spelled exactly the current sentential form. That correspondence is the construction, and the rest of this section only formalizes it.

\[ \text{matched} \;\cdot\; \text{stack} = \text{sentential form} \ \checkmark \]

13. Watch a derivation and a stack move together — line by line

Picture it

Animation

Shows: Each line of the worked example "Watch a derivation and a stack move together", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: At every stage, the input matched so far plus the stack contents spelled exactly the current sentential form. That correspondence is the construction, and the rest of this section only formalizes it.

14. The plan for both directions

Concept

Two constructions, each with a proof, and each following the shape used in the regular chapters.

  1. Grammar to machine: a single-state machine whose stack holds the unmatched part of a leftmost derivation.
  2. Machine to grammar: variables recording that the machine goes from one state to another while net-popping one stack symbol.

Sections 2 and 3 do the first with its proof; Sections 4 and 5 do the second. Section 6 collects what the equivalence buys.

15. Teach it back: The plan for both directions

Explain it

Discussion prompt

Explain The plan for both directions to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Two constructions, each with a proof, and each following the shape used in the regular chapters.

16. Rebuild the recipe: Reading a correspondence proof

Ranking

Put in order

These are the steps of Reading a correspondence proof, scrambled. Put them back in order before the next slide shows you.

  1. Identify the invariant relating the two objects at each step.
  2. Prove it holds at the start.
  3. Prove each move of one object can be mirrored by the other.
  4. Prove the mirroring goes both ways, or the languages will only be contained one way.
  5. Read off the language equality from the invariant at the end of a computation.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

17. Reading a correspondence proof

Pattern

Both proofs here relate two very different objects, and they share a shape worth recognizing.

  1. Identify the invariant relating the two objects at each step.
  2. Prove it holds at the start.
  3. Prove each move of one object can be mirrored by the other.
  4. Prove the mirroring goes both ways, or the languages will only be contained one way.
  5. Read off the language equality from the invariant at the end of a computation.

Step four is the one that takes the work. Showing every derivation gives a computation is usually easy; showing every computation gives a derivation needs the invariant stated carefully enough to run backwards.

18. Rule out three: Check yourself: the plan

Elimination

Eliminate the wrong options

In the grammar-to-machine construction, what do the stack contents represent?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The part of the current sentential form not yet matched against the input
  • B. The input symbols read so far
  • C. The rules applied so far, in order
  • D. The remaining input

Survives elimination: A

Why: A leftmost sentential form is a settled terminal prefix followed by the rest. The prefix is what the machine has already matched, so the stack holds precisely the remainder — the leftmost variable and everything after it.

19. Check yourself: the plan

Check

Think about what the stack is holding at each moment.

Check your understanding

In the grammar-to-machine construction, what do the stack contents represent?

  • A. The part of the current sentential form not yet matched against the input (correct)
  • B. The input symbols read so far
  • C. The rules applied so far, in order
  • D. The remaining input

Answer: A

Why: A leftmost sentential form is a settled terminal prefix followed by the rest. The prefix is what the machine has already matched, so the stack holds precisely the remainder — the leftmost variable and everything after it.

Why B tempts people
Those symbols have been consumed and are no longer needed. Keeping them would serve no purpose and would make the stack grow without bound.
Why C tempts people
The machine never needs to know which rules it used. It needs to know what obligations remain, which is what the stack records.
Why D tempts people
The remaining input is held by the input tape, not the stack. The stack holds what the derivation still owes, which is generally a different string.

20. This is what a parser actually does

Intuition

The construction is not a theoretical curiosity — it is the top-down parsing algorithm, written as a machine.

A recursive-descent parser keeps its pending obligations on the call stack, expanding the leftmost unfinished nonterminal and matching terminals as it goes. That is exactly the machine built in the next section.

The difference is only that a real parser must choose which rule to expand by, using lookahead, while the machine guesses nondeterministically. Making that choice deterministic is what parser generators do, and why they need the restrictions of Lesson 16.

\[ \text{nondeterministic guess} \;\longleftrightarrow\; \text{lookahead in a real parser} \]

21. Grammar to Machine

Section

Section 2

22. A machine with one state

Concept

Remarkably, the construction needs no states at all beyond the bookkeeping ones. All the work happens on the stack.

  1. One state does the whole simulation.
  2. Two extra states set up the initial stack and check the end.
  3. The stack alphabet is the grammar's variables together with its terminals.

The machine accepts by empty stack in the cleanest presentation, and by final state after the standard conversion of Lesson 16.

\[ \Gamma = V \cup \Sigma \cup \{Z_0\} \]

23. The three kinds of transition

Concept

Every move belongs to one of three families, and the whole construction is these three lines.

SituationMove
a variable on toppop it and push some right-hand side of one of its rules
a terminal on topread that same symbol from the input and pop it
the bottom marker on topaccept, if the input is exhausted

\[ \varepsilon, A \to \alpha \;\;\text{ for each rule } A \to \alpha; \qquad a, a \to \varepsilon \;\;\text{ for each terminal} \]

The first family consumes no input, which is why the machine can expand a whole derivation without reading anything. The second is where input is actually matched.

24. What each one costs: The three kinds of transition

Trade off

Comparison matrix

From The three kinds of transition: every row here is a choice with a cost. Fill the Move column, then say which row you would actually pick and what you give up for it.

SituationMove
a variable on toppop it and push some right-hand side of one of its rules
a terminal on topread that same symbol from the input and pop it
the bottom marker on topaccept, if the input is exhausted

25. Plan first: Convert a grammar to a machine

Step zero

Discussion prompt

Convert a grammar to a machine — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Set up the stack

Answer:

  1. Set up the stack
  2. Add one expansion transition per rule
  3. Add one matching transition per terminal
  4. Add the acceptance check
  5. Verify by tracing a member and a non-member

26. Convert a grammar to a machine

Worked example

Apply the construction to the matched-counts grammar.

\[ S \to aSb \;\mid\; \varepsilon \]

Set up the stack

Why: A free move from the start state pushes the start variable above the bottom marker, then enters the simulation state.

\[ \varepsilon, \varepsilon \to S Z_0 \]

Add one expansion transition per rule

Why: Two rules, so two transitions. Each pops the variable and pushes the rule's right-hand side.

\[ \varepsilon, S \to aSb; \qquad \varepsilon, S \to \varepsilon \]

Add one matching transition per terminal

Why: Two terminals, so two transitions. Each reads a symbol and pops the same symbol.

\[ a, a \to \varepsilon; \qquad b, b \to \varepsilon \]

Add the acceptance check

Why: A free move popping the bottom marker enters the accepting state, which is reachable only when everything above has been consumed.

\[ \varepsilon, Z_0 \to \varepsilon \]

Verify by tracing a member and a non-member

Why: On aabb the machine expands twice, matches four symbols and empties the stack — accept. On aab, every expansion sequence leaves either a b unmatched on the stack or an unmatched input symbol, so no computation accepts. Both verdicts match the grammar.

\[ aabb \in L(P) \qquad aab \notin L(P) \ \checkmark \]

27. Convert a grammar to a machine — line by line

Picture it

Animation

Shows: Each line of the worked example "Convert a grammar to a machine", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: On aabb the machine expands twice, matches four symbols and empties the stack — accept. On aab, every expansion sequence leaves either a b unmatched on the stack or an unmatched input symbol, so no computation accepts. Both verdicts match the grammar.

28. Guess the shape of the answer: Trace the constructed machine in full

Estimation

Predict first

Follow every configuration, so the correspondence with the derivation is visible.

Commit before you compute: what does Trace the constructed machine in full come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify acceptance by matching the remaining terminals

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Two matching moves consume both b's and empty the stack down to the marker, which the acceptance move then pops.

29. Trace the constructed machine in full

Worked example

Follow every configuration, so the correspondence with the derivation is visible.

\[ \text{input } aabb \]

Initialize

Why: The stack holds the start variable above the marker; nothing has been read.

\[ (q, \; aabb, \; SZ_0) \]

Expand, then match

Why: A free move replaces the variable by the recursive right-hand side; then the leading a is matched and popped.

\[ (q, \; aabb, \; aSbZ_0) \vdash (q, \; abb, \; SbZ_0) \]

Expand and match again

Why: The same pair of moves consumes the second a.

\[ (q, \; abb, \; aSbbZ_0) \vdash (q, \; bb, \; SbbZ_0) \]

Apply the empty rule

Why: The variable is popped and nothing is pushed, leaving only terminals on the stack.

\[ (q, \; bb, \; bbZ_0) \]

Verify acceptance by matching the remaining terminals

Why: Two matching moves consume both b's and empty the stack down to the marker, which the acceptance move then pops. The input is exhausted at exactly that moment, so the string is accepted — and the derivation it mirrors is the leftmost one from the grammar.

\[ (q, \; \varepsilon, \; Z_0) \vdash (q_{\text{acc}}, \; \varepsilon, \; \varepsilon) \ \checkmark \]

30. Trace the constructed machine in full — line by line

Picture it

Animation

Shows: Each line of the worked example "Trace the constructed machine in full", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Two matching moves consume both b's and empty the stack down to the marker, which the acceptance move then pops. The input is exhausted at exactly that moment, so the string is accepted — and the derivation it mirrors is the leftmost one from the grammar.

31. Something is wrong here: pushing the right-hand side in the wrong order

Anomaly

Predict first

A student writes this, and it looks reasonable:

Add the expansion transition for a rule whose right-hand side has several symbols.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Pop the variable, then push the first symbol, then the second, then the third.

Add the expansion transition for a rule whose right-hand side has several symbols.

Why: Pop the variable, then push the first symbol, then the second, then the third.

32. Trap: pushing the right-hand side in the wrong order

Trap

The trap

Add the expansion transition for a rule whose right-hand side has several symbols.

Push the symbols one at a time, left to right

Why: Pop the variable, then push the first symbol, then the second, then the third.

\[ A \to XYZ \;\Longrightarrow\; \text{push } X, \text{ then } Y, \text{ then } Z \]

Inspect the stack

Why: The last symbol pushed sits on top, so the stack reads the right-hand side backwards. The machine will then try to match the last symbol first.

\[ \text{stack top} = Z \quad\text{but the input expects } X \]

The fix

Add the expansion transition for a rule whose right-hand side has several symbols.

Push the whole right-hand side in one move, leftmost on top

Why: The transition function returns a string to push, and the convention is that its leftmost symbol ends up on top.

\[ \varepsilon, A \to XYZ \]

Inspect the stack

Why: The first symbol of the right-hand side is now on top, so it is the first thing matched — which is what a leftmost derivation requires.

\[ \text{stack top} = X \;\checkmark \]

Note why the convention exists

Why: Lesson 16 defined the pushed component as a string precisely so a whole right-hand side goes on in one move, in the order that makes this construction work.

33. Decode the notation: Trap: pushing the right-hand side in the wrong order

Notation

Annotate

From Trap: pushing the right-hand side in the wrong order — read this one piece at a time. What is each part doing?

On: \( A \to XYZ \;\Longrightarrow\; \text{push } X, \text{ then } Y, \text{ then } Z \)

  • Pop the variable, then push the first symbol, then the second, then the third.
  • The last symbol pushed sits on top, so the stack reads the right-hand side backwards. The machine will then try to match the last symbol first.
  • The transition function returns a string to push, and the convention is that its leftmost symbol ends up on top.

34. What has to be given first: Convert a grammar with two variables

Missing information

Discussion prompt

A second conversion, where the stack holds several pending obligations at once.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The start variable is pushed, then replaced by two variables in one move — the first of them on top.

35. Convert a grammar with two variables

Worked example

A second conversion, where the stack holds several pending obligations at once.

\[ S \to AB, \qquad A \to a, \qquad B \to b \]

Set up and expand the start variable

Why: The start variable is pushed, then replaced by two variables in one move — the first of them on top.

\[ (q, ab, SZ_0) \vdash (q, ab, ABZ_0) \]

Expand the leftmost variable

Why: Only the top symbol is visible, so the first variable is the one expanded. Its rule replaces it with a terminal.

\[ (q, ab, aBZ_0) \]

Match, then expand again

Why: The terminal on top is matched against the input; the second variable then becomes visible and is expanded.

\[ (q, b, BZ_0) \vdash (q, b, bZ_0) \]

Match and accept

Why: The last terminal matches, the marker is popped, and the input is exhausted at the same moment.

Verify the pending obligation was held correctly

Why: Between the first expansion and the second, the second variable sat on the stack below the first — invisible but preserved. That is exactly the role of a stack here: holding obligations in the order they will be needed.

\[ ab \in L(P) = L(G) \ \checkmark \]

36. Convert a grammar with two variables — line by line

Picture it

Animation

Shows: Each line of the worked example "Convert a grammar with two variables", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Between the first expansion and the second, the second variable sat on the stack below the first — invisible but preserved. That is exactly the role of a stack here: holding obligations in the order they will be needed.

37. Why one state suffices

Intuition

It is surprising that a machine of three states can simulate a grammar of any size. The reason is worth stating.

A finite automaton keeps everything it knows in its state, so more to remember means more states. This machine keeps everything on the stack, and the stack is unbounded — so the state count never has to grow.

The grammar's size shows up instead in the transition count: one expansion transition per rule. Complexity moved from the states to the transitions and the stack alphabet.

\[ |Q| = 3, \qquad |\delta| = |R| + |\Sigma| + 2 \]

38. The grammar-to-machine recipe

Pattern

Four lines, applicable to any grammar without modification.

  1. Push the start variable above a bottom marker, from a fresh start state.
  2. For each rule, add a free transition popping its variable and pushing its right-hand side, leftmost on top.
  3. For each terminal, add a transition reading it and popping the same symbol.
  4. Pop the bottom marker into an accepting state.
  5. The result has three states regardless of the grammar's size.

Note what the construction does not need: no normal form, no removal of epsilon-rules, no analysis of the grammar at all. It works on any grammar as written.

39. Answer it before you see the options: Check yourself: the construction

Prediction

Predict first

A grammar has 12 variables, 40 rules and 5 terminals. How many states does the constructed machine have?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Three, regardless of the grammar's size

Why: All the grammar's structure lives on the stack, not in the states. The construction uses one state to run the simulation plus two for setup and acceptance, whatever the grammar looks like.

40. Check yourself: the construction

Check

Count what the construction produces.

Check your understanding

A grammar has 12 variables, 40 rules and 5 terminals. How many states does the constructed machine have?

  • A. Three, regardless of the grammar's size (correct)
  • B. Twelve — one per variable
  • C. Forty — one per rule
  • D. Fifty-seven — one per variable, rule and terminal

Answer: A

Why: All the grammar's structure lives on the stack, not in the states. The construction uses one state to run the simulation plus two for setup and acceptance, whatever the grammar looks like.

Why B tempts people
Variables become stack symbols, not states. Putting them in the state set would bound the derivation depth, which must be unbounded.
Why C tempts people
Rules become transitions on the single simulation state. Each is a self-loop, so no new state is needed.
Why D tempts people
This adds up the wrong things entirely. The counts affect the transition and stack alphabets, never the state count.

41. Why the First Construction Works

Section

Section 3

42. The invariant

Concept

The proof rests on a single claim relating the machine's configuration to the grammar's sentential forms.

At any point in a computation, the input already matched, followed by the stack contents read top to bottom, is exactly a sentential form of a leftmost derivation from the start variable.

\[ S \;\Longrightarrow^{*}_{\mathrm{lm}}\; u\,\gamma \quad\text{where } u \text{ is matched and } \gamma \text{ is the stack} \]

Prove that, and both containments follow immediately by looking at the two ends of a computation.

43. Plan first: Prove the invariant

Step zero

Discussion prompt

Prove the invariant — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Expansion case

Answer:

  1. Expansion case
  2. Matching case
  3. Note the leftmost condition is maintained
  4. Verify both containments follow

44. Prove the invariant

Worked example

Induction on the number of moves, with two cases matching the two transition families.

Base case

Why: Before any move, nothing is matched and the stack holds the start variable alone. Their concatenation is the start variable, which is a sentential form in zero steps.

\[ u = \varepsilon, \; \gamma = S \;\Rightarrow\; u\gamma = S \]

Expansion case

Why: A free move pops a variable and pushes a right-hand side. The matched prefix is unchanged, and the concatenation changes by exactly one rule application to the leftmost variable — which is a leftmost derivation step.

\[ uA\beta \;\Longrightarrow_{\mathrm{lm}}\; u\alpha\beta \]

Matching case

Why: A reading move consumes one terminal and pops the same symbol. The matched prefix grows by that symbol and the stack loses it, so the concatenation is unchanged.

\[ u\,a\,\beta \;: \; \text{matched grows}, \; \text{stack shrinks}, \; \text{product fixed} \]

Note the leftmost condition is maintained

Why: Expansion always acts on the top of the stack, which is the leftmost symbol of the unmatched part. Since everything to its left is already terminal, it is the leftmost variable.

Verify both containments follow

Why: An accepting computation ends with an empty stack and the whole input matched, so the invariant gives a leftmost derivation of that input. Conversely, any leftmost derivation can be replayed as a computation, expanding when the top is a variable and matching when it is a terminal. Both directions hold, so the languages are equal.

\[ L(P) = L(G) \ \checkmark \]

45. Prove the invariant — line by line

Picture it

Animation

Shows: Each line of the worked example "Prove the invariant", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: An accepting computation ends with an empty stack and the whole input matched, so the invariant gives a leftmost derivation of that input. Conversely, any leftmost derivation can be replayed as a computation, expanding when the top is a variable and matching when it is a terminal. Both directions hold, so the languages are equal.

46. Why leftmost is the right discipline

Intuition

The proof used leftmost derivations specifically, and no other discipline would have worked.

A stack exposes only its top symbol. The top of the stack is the leftmost unmatched symbol, so the only variable the machine can act on is the leftmost one — which forces the derivation to be leftmost.

A rightmost derivation would need to expand the last variable, which sits at the bottom of the stack and is invisible. So the construction and the discipline are not independent choices: the stack picks the discipline.

\[ \text{stack top} \;=\; \text{leftmost unmatched symbol} \]

47. Where does each piece belong: Equivalence of PDAs & CFGs

Sorting

Sort into buckets

These are the pieces of Equivalence of PDAs & CFGs, out of order. Put each one back under the part of the lesson it belongs to.

The Theorem
The statement; Why the two directions differ in difficulty; The key observation
Grammar to Machine
A machine with one state; The three kinds of transition; Convert a grammar to a machine
Why the First Construction Works
The invariant; Prove the invariant; Why leftmost is the right discipline
s1
The Theorem is where Equivalence of PDAs & CFGs puts The statement, Why the two directions differ in difficulty, The key observation. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Grammar to Machine is where Equivalence of PDAs & CFGs puts A machine with one state, The three kinds of transition, Convert a grammar to a machine. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Why the First Construction Works is where Equivalence of PDAs & CFGs puts The invariant, Prove the invariant, Why leftmost is the right discipline. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

48. Where the nondeterminism is used

Concept

The constructed machine is heavily nondeterministic, and it is worth being precise about where.

Every variable with several alternatives gives several applicable expansion transitions in the same configuration. The machine guesses which rule the derivation used.

\[ A \to \alpha_1 \;\mid\; \alpha_2 \;\Rightarrow\; \text{two moves available} \]

Matching transitions involve no choice at all — the input symbol and the stack top must agree, or the branch dies. So all the guessing is about which rule, never about what to match.

That is exactly the choice a real parser resolves with lookahead, and exactly the choice that cannot always be resolved, which is why Lesson 16's deterministic machines are weaker.

49. State the rule before it runs: Watch a wrong guess die

Hypothesis

Predict first

Watch a wrong guess die is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Guess the empty rule immediately

Why: The machine pops the start variable and pushes nothing, leaving only the bottom marker.

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

50. Watch a wrong guess die

Worked example

Run the constructed machine down a branch that guesses badly, to see the failure mode.

\[ S \to aSb \;\mid\; \varepsilon, \qquad \text{input } ab \]

Guess the empty rule immediately

Why: The machine pops the start variable and pushes nothing, leaving only the bottom marker.

\[ (q, ab, SZ_0) \vdash (q, ab, Z_0) \]

Try to continue

Why: The only remaining move pops the marker into the accepting state — but the input still holds two unread symbols.

See the branch fail

Why: Acceptance requires the input to be exhausted, and it is not. This computation does not accept, and no move extends it usefully.

\[ (q_{\text{acc}}, ab, \varepsilon) \;: \text{ input not exhausted} \]

Find the branch that succeeds

Why: Guessing the recursive rule first pushes the three-symbol right-hand side, after which the a matches, the empty rule fires, and the b matches.

Verify acceptance depends on the good branch only

Why: One accepting computation exists, so the string is accepted — the failing branch places no constraint, exactly as Lesson 5 established. The machine's nondeterminism is doing precisely the work of searching for the right derivation.

\[ ab \in L(P) \ \checkmark \]

51. Watch a wrong guess die — line by line

Picture it

Animation

Shows: Each line of the worked example "Watch a wrong guess die", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: One accepting computation exists, so the string is accepted — the failing branch places no constraint, exactly as Lesson 5 established. The machine's nondeterminism is doing precisely the work of searching for the right derivation.

52. What the correspondence preserves

Concept

The construction is faithful in a stronger sense than mere language equality, and that strengthening is what makes it useful for parsing.

Grammar objectMachine object
a leftmost derivationan accepting computation
a rule applicationan expansion transition
a terminal in the yielda matching transition
a parse treethe whole computation, up to scheduling

So a machine that records which expansion it chose can output the parse tree, not just a verdict. That is exactly what a recursive-descent parser returns.

53. What has to happen first: Recover a parse tree from a computation

Ranking

Put in order

Put the moves of Recover a parse tree from a computation into the order they have to happen.

  1. Record each expansion
  2. Note the order
  3. Rebuild the tree
  4. Check the matching moves need no logging
  5. Verify on the earlier trace

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Every free move that pops a variable corresponds to one rule application.

54. Recover a parse tree from a computation

Worked example

Use the correspondence in the direction that matters in practice.

Record each expansion

Why: Every free move that pops a variable corresponds to one rule application. Log the rule as it happens.

Note the order

Why: Because the derivation is leftmost, the logged rules arrive in leftmost order — which determines the tree uniquely.

Rebuild the tree

Why: Start with the start variable as the root and apply the logged rules in order, always to the leftmost unexpanded variable.

Check the matching moves need no logging

Why: They consume terminals and create leaves, but the leaves are determined by the rules already logged. So only expansions carry information.

Verify on the earlier trace

Why: The trace of the four-symbol string logged the recursive rule twice and the empty rule once, in that order. Replaying them leftmost from the start variable rebuilds the tree whose yield is that string, confirming nothing was lost.

\[ \text{expansions logged} \;\longrightarrow\; \text{unique leftmost derivation} \;\longrightarrow\; \text{tree} \ \checkmark \]

55. Decode the notation: Recover a parse tree from a computation

Notation

Annotate

From Recover a parse tree from a computation — read this one piece at a time. What is each part doing?

On: \( \text{expansions logged} \;\longrightarrow\; \text{unique leftmost derivation} \;\longrightarrow\; \text{tree} \ \checkmark \)

  • Every free move that pops a variable corresponds to one rule application. Log the rule as it happens.
  • Because the derivation is leftmost, the logged rules arrive in leftmost order — which determines the tree uniquely.
  • Start with the start variable as the root and apply the logged rules in order, always to the leftmost unexpanded variable.

56. Check yourself: the correctness proof

Check

Recall which part of the invariant each transition family changes.

Check your understanding

In the correctness proof, what does a matching transition do to the concatenation of the matched input and the stack?

  • A. Leaves it unchanged — the symbol moves from the stack to the matched prefix (correct)
  • B. Applies one rule of the grammar to it
  • C. Shortens it by one symbol
  • D. Reverses it

Answer: A

Why: The transition consumes a terminal from the input and pops the identical symbol from the stack, so the symbol simply crosses from one part of the concatenation to the other. That is why matching moves preserve the sentential form while expansions advance it.

Why B tempts people
Rule application is what the expansion transitions do. Matching transitions correspond to no grammar step at all.
Why C tempts people
The stack shortens but the matched prefix lengthens by the same symbol, so the total is preserved.
Why D tempts people
Nothing in the construction reverses anything. The stack is maintained leftmost-on-top precisely to avoid that.

57. Machine to Grammar: The Idea

Section

Section 4

58. What a grammar must capture

Concept

The reverse direction is harder because a grammar has no stack, so it must express stack behaviour through its variables.

The insight is to focus on the net effect of a stretch of computation, rather than on individual moves. Specifically: what does it take for the machine to go from one state to another while removing exactly one symbol from the stack?

If that question can be answered by a variable, then the whole computation decomposes, because every accepting computation empties the stack one symbol at a time.

\[ A_{pXq} \;: \; \text{from } p \text{ to } q, \text{ net-popping } X \]

59. The triple variables

Concept

triple variable — A grammar variable indexed by a start state, a stack symbol and an end state, generating exactly the input strings that take the machine between those states while net-removing that symbol.

There is one variable per triple, so the grammar has as many variables as the number of states squared times the size of the stack alphabet. That is the source of the construction's size.

\[ |V| = |Q|^{2} \cdot |\Gamma| \]

The start variable is the one for going from the start state to an accepting state while removing the initial stack symbol — which is exactly acceptance by empty stack.

60. Why 'net-popping one symbol' is the right unit

Intuition

The choice of unit is what makes the decomposition work, and it is worth seeing why a cruder unit fails.

During the stretch, the stack may grow enormously — but it must return to exactly one symbol lower than it started, and it may never dip below that level in between. That makes the stretch self-contained: nothing below the symbol is touched.

Self-containment is what allows two stretches to be concatenated without interfering. A unit that allowed the stack to dip below its starting level would let stretches interact, and no context-free rule could describe that.

\[ \text{never dips below} \;\Rightarrow\; \text{stretches compose independently} \]

61. The two rule families

Concept

The grammar's rules come from decomposing a stretch in the two possible ways.

  1. The stretch's first move pops the symbol outright. Then the whole stretch is one move, and the rule sends the variable to that move's input symbol.
  2. The stretch's first move pushes something. Then the pushed symbols must each be removed later, and the rule chains one triple variable per pushed symbol.

The second family is where the branching comes from, and it is why the construction is usually presented after converting the machine to a restricted form.

62. The restricted form

Concept

The construction is far simpler if the machine is first normalized, much as grammars were in Lesson 15.

  1. It has a single accepting state.
  2. It empties its stack before accepting.
  3. Every transition either pushes exactly one symbol or pops exactly one symbol — never both, never neither.

Every machine can be put in this form by adding states and splitting transitions, and the language is unchanged. With it, the second rule family reduces to a single shape.

\[ A_{pXq} \to a \; A_{rYs} \; b \]

63. Plan first: Normalize a machine

Step zero

Discussion prompt

Normalize a machine — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Force a single accepting state

Answer:

  1. Force a single accepting state
  2. Force the stack to empty before accepting
  3. Split transitions that neither push nor pop
  4. Split transitions that both pop and push several
  5. Verify the language is unchanged

64. Normalize a machine

Worked example

Apply the three conditions to a machine that violates all of them.

Force a single accepting state

Why: Add a fresh accepting state, reachable by a free move from every old accepting state, and demote the old ones.

Force the stack to empty before accepting

Why: Add a draining state that pops any symbol without consuming input, placed between the old accepting states and the new one.

\[ \varepsilon, X \to \varepsilon \;\text{ for every } X \]

Split transitions that neither push nor pop

Why: Replace each by two moves through a fresh state: one pushing a dummy symbol, one popping it.

\[ a, \varepsilon \to \varepsilon \;\Longrightarrow\; a, \varepsilon \to D \;\text{ then }\; \varepsilon, D \to \varepsilon \]

Split transitions that both pop and push several

Why: Replace each by a chain: one pop, then one push per symbol of the pushed string, each through a fresh state.

Verify the language is unchanged

Why: Every added move consumes no input and every split preserves the net stack effect, so each original computation corresponds to exactly one normalized computation with the same input and the same verdict.

\[ L(P) = L(P') \ \checkmark \]

65. Normalize a machine — line by line

Picture it

Animation

Shows: Each line of the worked example "Normalize a machine", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Every added move consumes no input and every split preserves the net stack effect, so each original computation corresponds to exactly one normalized computation with the same input and the same verdict.

66. How sure are you: Check yourself: the triple idea

Commit first

Predict first

The variable indexed by states p and q and stack symbol X generates which strings?

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: Those taking the machine from p to q while net-removing X and never dipping below it

Why: The stack may grow arbitrarily during the stretch, but it must end exactly one symbol lower and must never go below that level in between. That self-containment is what lets stretches be concatenated without interference.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

67. Check yourself: the triple idea

Check

Think about what the stack does during a stretch.

Check your understanding

The variable indexed by states p and q and stack symbol X generates which strings?

  • A. Those taking the machine from p to q while net-removing X and never dipping below it (correct)
  • B. Those taking the machine from p to q with X on top throughout
  • C. Those taking the machine from p to q while pushing X
  • D. Those the machine accepts starting in state p

Answer: A

Why: The stack may grow arbitrarily during the stretch, but it must end exactly one symbol lower and must never go below that level in between. That self-containment is what lets stretches be concatenated without interference.

Why B tempts people
Requiring X to stay on top would forbid the stack from growing at all, which would reduce the machine to a finite automaton.
Why C tempts people
The variable is about removing X, not pushing it. Removal is what makes the decomposition terminate.
Why D tempts people
That would be the start variable's job for one particular p, not a general triple variable.

68. Why this direction needs so many variables

Intuition

The grammar-to-machine construction used three states; this one uses a variable per triple. The asymmetry has a clear cause.

A stack is a single unbounded object, and a machine can consult only its top. A grammar has no such object, so every fact about stack behaviour must be encoded in the finitely many variable names.

The triples are the smallest encoding that suffices: which state the stretch starts in, which symbol it removes, and which state it ends in. Dropping any component makes the decomposition unsound.

\[ |Q|^{2}|\Gamma| \text{ variables} \;\longleftrightarrow\; \text{one stack} \]

69. Machine to Grammar: The Construction

Section

Section 5

70. The size of the resulting grammar

Concept

The reverse construction is expensive, and knowing the cost explains why it is rarely run in practice.

QuantitySize
variablesstates squared, times the stack alphabet
rules from the push-pop schemaone per matched push-pop pair, times two states
rules from the splitting schemavariables times the state count
useless variables producedtypically most of them

A ten-state machine with a five-symbol stack alphabet already gives five hundred variables, and the great majority generate nothing. Running the Lesson 15 cleanup afterwards is not optional in practice.

\[ |V| = |Q|^{2}|\Gamma| \]

71. Guess the shape of the answer: Count the variables for a small machine

Estimation

Predict first

Make the cost concrete before carrying out a conversion.

Commit before you compute: what does Count the variables for a small machine come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the cleanup is worth running

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Applying the generating and reachable closures of Lesson 15 removes every variable with no stretch, typically leaving a handful.

72. Count the variables for a small machine

Worked example

Make the cost concrete before carrying out a conversion.

Take a machine

Why: Three states, and a stack alphabet of two symbols plus the bottom marker.

\[ |Q| = 3, \qquad |\Gamma| = 3 \]

Apply the formula

Why: One variable per ordered pair of states and each stack symbol.

\[ |V| = 3^{2} \cdot 3 = 27 \]

Estimate how many are useful

Why: A variable is useful only if the machine can actually get from the first state to the second while removing that symbol. For most triples no such stretch exists.

Note the practical consequence

Why: The generated grammar is unreadable and mostly dead. Its purpose is to establish the theorem, not to be used.

Verify the cleanup is worth running

Why: Applying the generating and reachable closures of Lesson 15 removes every variable with no stretch, typically leaving a handful. That is the grammar worth looking at, and it is usually recognizable as one you could have written directly.

\[ 27 \text{ variables} \;\longrightarrow\; \text{a handful, after cleanup} \ \checkmark \]

73. Count the variables for a small machine — line by line

Picture it

Animation

Shows: Each line of the worked example "Count the variables for a small machine", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Applying the generating and reachable closures of Lesson 15 removes every variable with no stretch, typically leaving a handful. That is the grammar worth looking at, and it is usually recognizable as one you could have written directly.

74. The three rule schemas

Concept

With the machine normalized, the grammar is generated by three schemas.

SchemaWhen
push then later popa push move and a matching pop move exist
split a stretcha stretch passes through an intermediate state
empty stretchthe start and end states coincide

\[ A_{pXq} \to a\,A_{rYs}\,b, \qquad A_{pXq} \to A_{pXr}A_{rXq}, \qquad A_{pXp} \to \varepsilon \]

The third schema is where derivations terminate, and the first is where input is actually generated.

75. Fill in: When for The three rule schemas

Comparison

Comparison matrix

From The three rule schemas: refill the When column from what you know. The rest of the table is as it appeared.

SchemaWhen
push then later popa push move and a matching pop move exist
split a stretcha stretch passes through an intermediate state
empty stretchthe start and end states coincide

76. Guess the shape of the answer: Build the grammar from a machine

Estimation

Predict first

Convert a small normalized machine, one schema at a time.

Commit before you compute: what does Build the grammar from a machine come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the grammar generates the machine's language

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The rules produce a's on the left and b's on the right in matching numbers, terminating with the empty string — exactly the matched-counts language the machine recognizes.

77. Build the grammar from a machine

Worked example

Convert a small normalized machine, one schema at a time.

\[ p: \; a, \varepsilon \to X \qquad q: \; b, X \to \varepsilon \]

List the variables

Why: Two states and one stack symbol beyond the marker, so there are four triple variables in principle, though not all will be useful.

\[ A_{pXp}, \; A_{pXq}, \; A_{qXp}, \; A_{qXq} \]

Apply the push-then-pop schema

Why: The push move reads an a and puts X on the stack; the pop move reads a b and removes it. Together they bracket a stretch.

\[ A_{pXq} \to a \; A_{pXq} \; b \]

Apply the empty-stretch schema

Why: Where the start and end states coincide, the empty string is generated.

\[ A_{pXp} \to \varepsilon \]

Identify the start variable

Why: It records going from the start state to the accepting state while removing the initial marker.

Verify the grammar generates the machine's language

Why: The rules produce a's on the left and b's on the right in matching numbers, terminating with the empty string — exactly the matched-counts language the machine recognizes. Deriving aabb and failing to derive aab confirms it.

\[ L(G) = \{a^{n}b^{n}\} = L(P) \ \checkmark \]

78. Build the grammar from a machine — line by line

Picture it

Animation

Shows: Each line of the worked example "Build the grammar from a machine", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The rules produce a's on the left and b's on the right in matching numbers, terminating with the empty string — exactly the matched-counts language the machine recognizes. Deriving aabb and failing to derive aab confirms it.

79. Why the construction is correct

Concept

The proof is two inductions, mirroring the two directions of the language equality.

  1. Every derivation gives a computation. Induction on derivation length, using the schemas to build the corresponding stretch.
  2. Every computation gives a derivation. Induction on computation length, decomposing at the first point the stack returns to its starting level.

The second induction is the delicate one, and the never-dips-below condition is exactly what makes the decomposition point well defined.

\[ \text{first return to the starting level} \;=\; \text{the split point} \]

80. Plan first: Decompose a computation to find its derivation

Step zero

Discussion prompt

Decompose a computation to find its derivation — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Take an accepting computation

Answer:

  1. Take an accepting computation
  2. Find the first move
  3. In the pushing case, find the matching pop
  4. Split into two stretches
  5. Verify the induction terminates

81. Decompose a computation to find its derivation

Worked example

Run the harder induction on a concrete computation, to see the decomposition happen.

Take an accepting computation

Why: It starts with the marker on the stack and ends with the stack empty, so it net-removes exactly that one symbol.

Find the first move

Why: In the normalized machine it either pops the marker outright, or pushes a symbol that must be removed later.

In the pushing case, find the matching pop

Why: Follow the computation until the stack first returns to its starting level. That move is the matching pop, and it is unique.

Split into two stretches

Why: The part strictly between the push and its matching pop is one self-contained stretch; the part after the pop is another. Each is shorter than the whole.

\[ A_{pXq} \to a \; A_{rYs} \; b \quad\text{applies here} \]

Verify the induction terminates

Why: Both stretches are strictly shorter than the original computation, so the induction is well founded and bottoms out at the empty-stretch schema. Every accepting computation therefore yields a derivation, completing the harder containment.

\[ \text{both pieces shorter} \;\Rightarrow\; \text{induction well founded} \ \checkmark \]

82. Decompose a computation to find its derivation — line by line

Picture it

Animation

Shows: Each line of the worked example "Decompose a computation to find its derivation", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Both stretches are strictly shorter than the original computation, so the induction is well founded and bottoms out at the empty-stretch schema. Every accepting computation therefore yields a derivation, completing the harder containment.

83. Something is wrong here: forgetting that the stack may not dip below

Anomaly

Predict first

A student writes this, and it looks reasonable:

Decompose a stretch that goes from one state to another while removing one symbol.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Choose any convenient moment when the depth matches, and cut the stretch there.

Decompose a stretch that goes from one state to another while removing one symbol.

Why: Choose any convenient moment when the depth matches, and cut the stretch there.

84. Trap: forgetting that the stack may not dip below

Trap

The trap

Decompose a stretch that goes from one state to another while removing one symbol.

Split at any point where the stack returns to its starting height

Why: Choose any convenient moment when the depth matches, and cut the stretch there.

See what goes wrong

Why: If the stack dipped below the starting level in between, the symbol being removed was already gone, so the two pieces are not independent — one of them touched what lay beneath.

\[ \text{dips below} \;\Rightarrow\; \text{pieces interact} \]

The fix

Decompose a stretch that goes from one state to another while removing one symbol.

Split at the FIRST return to the starting level

Why: Taking the earliest such moment guarantees the stack never went below it beforehand, so the symbol was untouched throughout the first piece.

Confirm both pieces are self-contained

Why: The first piece never sees below the symbol it is bracketed by, and the second starts fresh at the lower level. Neither can interfere with the other.

\[ \text{first return} \;\Rightarrow\; \text{both pieces self-contained} \ \checkmark \]

85. Which of these survive contact with Equivalence of PDAs & CFGs?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
The pairing that makes the chapter cohere: the grammars and the machines describe exactly the same languages.; A grammar has inductive structure and a machine does not, exactly as in the regular case — but here the asymmetry lands differently.; Everything in the easy direction follows from noticing what a leftmost derivation looks like at each moment.
Breaks
Add the expansion transition for a rule whose right-hand side has several symbols.; Decompose a stretch that goes from one state to another while removing one symbol.
sound
These are stated as this lesson states them — each one survives the edge cases Equivalence of PDAs & CFGs puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

86. Without one step: The machine-to-grammar recipe

Constraint

Discussion prompt

Run The machine-to-grammar recipe with this step confiscated:

Add a rule for each matched push-pop pair, bracketing an intermediate stretch.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Normalize the machine: one accepting state, empty stack before accepting, every move a single push or a single pop.
  2. Create one variable per triple of state, stack symbol and state.
  3. Add a rule for each matched push-pop pair, bracketing an intermediate stretch.
  4. Add a splitting rule for each intermediate state.
  5. Add an empty rule wherever the start and end states coincide.

87. The machine-to-grammar recipe

Pattern

Collected, though it is rarely carried out by hand on anything large.

  1. Normalize the machine: one accepting state, empty stack before accepting, every move a single push or a single pop.
  2. Create one variable per triple of state, stack symbol and state.
  3. Add a rule for each matched push-pop pair, bracketing an intermediate stretch.
  4. Add a splitting rule for each intermediate state.
  5. Add an empty rule wherever the start and end states coincide.

The grammar produced is large and unreadable, and usually contains many useless variables. Running the cleanup pass of Lesson 15 afterwards is normal.

88. Where does it stop working: The machine-to-grammar recipe

Edge cases

Discussion prompt

The machine-to-grammar recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

Collected, though it is rarely carried out by hand on anything large.

89. Answer it before you see the options: Check yourself: the reverse construction

Prediction

Predict first

Why must a stretch be split at the FIRST moment the stack returns to its starting level?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: A later split could allow the stack to have dipped below, so the pieces would not be independent

Why: The triple variables describe stretches that never go below their starting level, which is what makes them composable. Splitting at a later return would permit a dip in between, and the piece would no longer describe a self-contained stretch.

90. Check yourself: the reverse construction

Check

Think about where the decomposition must cut.

Check your understanding

Why must a stretch be split at the FIRST moment the stack returns to its starting level?

  • A. A later split could allow the stack to have dipped below, so the pieces would not be independent (correct)
  • B. The first return is the only one that exists
  • C. Later returns would make the pieces longer than the original
  • D. The grammar has no rule for later splits

Answer: A

Why: The triple variables describe stretches that never go below their starting level, which is what makes them composable. Splitting at a later return would permit a dip in between, and the piece would no longer describe a self-contained stretch.

Why B tempts people
A computation may return to a given level many times. The construction chooses the first specifically to rule out the dips.
Why C tempts people
Any split produces two pieces shorter than the whole, so termination is not the issue.
Why D tempts people
The splitting schema applies at any intermediate state. The constraint is about stack depth, not about which rules exist.

91. Consequences

Section

Section 6

92. What the equivalence licenses

Concept

With both constructions in hand, either formalism may be used for any argument about the class.

\[ \text{CFG} \;\equiv\; \text{PDA} \]

93. What has to happen first: Use the equivalence to prove a closure property

Ranking

Put in order

Put the moves of Use the equivalence to prove a closure property into the order they have to happen.

  1. Choose the formalism
  2. Take grammars for both languages
  3. Add a fresh start variable
  4. Read off the language
  5. Verify the machine side would have been harder

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Grammars, because a union is one extra rule and a machine construction would need care about disjoint stacks.

94. Use the equivalence to prove a closure property

Worked example

Show the context-free languages are closed under union, using whichever side is easier.

Choose the formalism

Why: Grammars, because a union is one extra rule and a machine construction would need care about disjoint stacks.

Take grammars for both languages

Why: By the theorem, each language has a grammar. Rename variables so the two sets are disjoint.

Add a fresh start variable

Why: One alternative per grammar, so a derivation commits to one branch immediately.

\[ S \to S_1 \;\mid\; S_2 \]

Read off the language

Why: A derivation goes entirely through one branch, so the generated strings are exactly those of one grammar or the other.

Verify the machine side would have been harder

Why: A machine construction would need a fresh start state with free moves into both machines, plus care that neither machine's stack symbols interfered with the other's — three conditions instead of one rule. The equivalence let the easier side be chosen, which is the whole point.

\[ L_1 \cup L_2 \text{ context-free} \ \checkmark \]

95. Use the equivalence to prove a closure property — line by line

Picture it

Animation

Shows: Each line of the worked example "Use the equivalence to prove a closure property", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: A machine construction would need a fresh start state with free moves into both machines, plus care that neither machine's stack symbols interfered with the other's — three conditions instead of one rule. The equivalence let the easier side be chosen, which is the whole point.

96. The hierarchy, now with recognizers on both rows

Concept

Two classes are fully described, each by a generator and a recognizer proved equivalent.

ClassGeneratorRecognizerEquivalence proved in
regularregular expressionfinite automatonLessons 8 and 9
context-freecontext-free grammarpushdown automatonthis lesson

The pattern is deliberate, and it repeats once more: Lesson 20 introduces a machine, and Lesson 23 argues that it captures the informal notion of an algorithm — the analogue, at the top of the hierarchy, of these equivalence theorems.

97. Teach it back: The hierarchy, now with recognizers on both rows

Explain it

Discussion prompt

Explain The hierarchy, now with recognizers on both rows to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Two classes are fully described, each by a generator and a recognizer proved equivalent.

98. The parsing problem

Concept

The equivalence also settles a practical question: given a grammar and a string, can membership be decided?

Yes. Convert the grammar to Chomsky normal form by Lesson 15, then fill a table indexed by substrings, combining pairs using the binary rules. The algorithm runs in time cubic in the string's length.

\[ O\big(n^{3} \cdot |G|\big) \]

The normal form is essential: the table combines exactly two pieces per entry, which needs every rule to have exactly two variables on the right. That is why Lesson 15 came before this one.

99. By analogy: The parsing problem

Analogy

Discussion prompt

Explain The parsing problem by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The equivalence also settles a practical question: given a grammar and a string, can membership be decided?

100. Where each formalism wins

Intuition

Two equivalent descriptions, and the choice is never arbitrary in practice.

TaskFormalism
specifying a languagegrammar
proving closure under union, concatenation, stargrammar
proving a language is deterministicmachine
implementing a recognizermachine, derived from the grammar
recovering structure from a stringgrammar, via the parse tree

The third row is the one that needs the machine: determinism is a property of computations, and grammars have no notion of one.

101. What each one costs: Where each formalism wins

Trade off

Comparison matrix

From Where each formalism wins: every row here is a choice with a cost. Fill the Formalism column, then say which row you would actually pick and what you give up for it.

TaskFormalism
specifying a languagegrammar
proving closure under union, concatenation, stargrammar
proving a language is deterministicmachine
implementing a recognizermachine, derived from the grammar
recovering structure from a stringgrammar, via the parse tree

102. What is still missing

Concept

The class is now described from the inside by two equivalent formalisms. Nothing yet describes it from the outside.

Every technique here proves membership by exhibiting an object. To prove a language is not context-free, none of it helps — exactly the situation Lesson 9 left the regular languages in.

Lesson 18 supplies the missing tool, and its shape is already visible: normal form bounds a parse tree's height by the string's length, a tall tree must repeat a variable on some path, and a repeated variable can be pumped.

\[ \text{long string} \;\Rightarrow\; \text{tall tree} \;\Rightarrow\; \text{repeated variable} \;\Rightarrow\; \text{pumping} \]

103. Break it if you can: What is still missing

Counterexample

Discussion prompt

The class is now described from the inside by two equivalent formalisms. Nothing yet describes it from the outside.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Every technique here proves membership by exhibiting an object. To prove a language is not context-free, none of it helps — exactly the situation Lesson 9 left the regular languages in.

104. Rule out three: Check yourself: consequences

Elimination

Eliminate the wrong options

You want to prove the context-free languages are closed under concatenation. Which formalism gives the shorter proof?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Grammars — a fresh start variable with one rule joining the two old start variables
  • B. Machines — run one and then the other
  • C. Neither — the class is not closed under concatenation
  • D. Both are equally short

Survives elimination: A

Why: One rule sending a fresh start variable to the two old ones in sequence does it, and the correctness argument is a single sentence about derivations splitting. The machine construction needs the two stacks kept from interfering, which takes real care.

105. Check yourself: consequences

Check

Think about which side each property is naturally stated on.

Check your understanding

You want to prove the context-free languages are closed under concatenation. Which formalism gives the shorter proof?

  • A. Grammars — a fresh start variable with one rule joining the two old start variables (correct)
  • B. Machines — run one and then the other
  • C. Neither — the class is not closed under concatenation
  • D. Both are equally short

Answer: A

Why: One rule sending a fresh start variable to the two old ones in sequence does it, and the correctness argument is a single sentence about derivations splitting. The machine construction needs the two stacks kept from interfering, which takes real care.

Why B tempts people
Running one machine then another requires knowing when the first has finished, which is not marked in the input, and needs the stacks separated. It is doable but much longer.
Why C tempts people
The class is closed under concatenation, and Lesson 19 proves it — by exactly the grammar construction described here.
Why D tempts people
The grammar proof is one rule and one sentence; the machine proof needs normalization and disjointness arguments. They are not comparable in length.

106. Two equivalence theorems, one pattern

Intuition

Comparing this lesson with Lessons 8 and 9 shows the same argument shape twice, which makes both easier to remember.

Regular classContext-free class
easy directionexpression to machine, by gadgetsgrammar to machine, by stack simulation
what makes it easythe expression is inductivethe derivation is inductive
hard directionmachine to expression, by state eliminationmachine to grammar, by triple variables
what makes it harda graph has no inductive structurea stack must be summarized finitely

In both cases the generator has structure to recurse on and the recognizer does not. That asymmetry, not any accident of the models, is what makes one direction short and the other long.

107. Fill in: Context-free class for Two equivalence theorems, one pattern

Comparison

Comparison matrix

From Two equivalence theorems, one pattern: refill the Context-free class column from what you know. The rest of the table is as it appeared.

Regular classContext-free class
easy directionexpression to machine, by gadgetsgrammar to machine, by stack simulation
what makes it easythe expression is inductivethe derivation is inductive
hard directionmachine to expression, by state eliminationmachine to grammar, by triple variables
what makes it harda graph has no inductive structurea stack must be summarized finitely

108. Rebuild the recipe: Choosing a direction to convert

Ranking

Put in order

These are the steps of Choosing a direction to convert, scrambled. Put them back in order before the next slide shows you.

  1. Grammar to machine: cheap, three states, done constantly — it is what a parser is.
  2. Machine to grammar: expensive, quadratic in states, done mainly to prove the theorem.
  3. If you have a machine and want a grammar, first ask whether you can describe the language directly.
  4. If you have a grammar and want a recognizer, convert — and determinize if the language allows it.
  5. Always run the Lesson 15 cleanup on a grammar produced by the reverse construction.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

109. Choosing a direction to convert

Pattern

Both conversions exist, but only one is routinely worth doing.

  1. Grammar to machine: cheap, three states, done constantly — it is what a parser is.
  2. Machine to grammar: expensive, quadratic in states, done mainly to prove the theorem.
  3. If you have a machine and want a grammar, first ask whether you can describe the language directly.
  4. If you have a grammar and want a recognizer, convert — and determinize if the language allows it.
  5. Always run the Lesson 15 cleanup on a grammar produced by the reverse construction.

The third line is honest advice. The reverse construction is proof machinery, and a grammar written from an understanding of the language is invariably clearer than one generated from a machine.

110. Where this shows up: Equivalence of PDAs & CFGs

Real world

Discussion prompt

Outside this lesson: where does Equivalence of PDAs & CFGs actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of Choosing a direction to convert is doing the work in it.

Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.

Answer:

Lesson 17 proves that the generator and the recognizer describe the same class. It starts from the key observation that a leftmost sentential form splits into matched input and stack contents, then gives the three-state grammar-to-machine construction with its expansion and matching transitions and the push-order trap, the correctness proof by invariant, and why the stack forces leftmost derivations. It shows how to recover a parse tree from a computation, then gives the triple-variable encoding for the reverse direction with its never-dips-below condition, machine normalization, and the three rule schemas. It ends with the consequences: choosing whichever formalism makes a closure proof easier, cubic-time parsing via Chomsky normal form, and what is still missing before Lesson 18.

111. Connect it up: Equivalence of PDAs & CFGs

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The Theorem · Grammar to Machine · Why the First Construction Works · Machine to Grammar: The Idea · Machine to Grammar: The Construction · Consequences. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

112. What you can do now

Recap

The generator and the recognizer are now proved equivalent, completing the pairing this chapter set out to establish.

SituationMove
a grammar to runthree states, expand on variables, match on terminals
a right-hand side to pushleftmost symbol on top, in one move
a machine to describe as a grammarnormalize first, then one variable per triple
a closure property to provegrammars, unless the property is about computations
a language to show is not context-freeneither — that is Lesson 18

Lesson 18 supplies the pumping lemma for this class, built on the tree-height bound that Chomsky normal form made available.

Sources

  1. Sipser, Introduction to the Theory of Computation, 3rd ed., Ch. 2.2, Theorem 2.20 (Equivalence of CFGs and PDAs) — Cengage, 2013.
  2. Hopcroft, Motwani & Ullman, Introduction to Automata Theory, Languages, and Computation, 3rd ed., Ch. 6.3 (Equivalence of PDAs and CFGs) — Pearson, 2007.
  3. Chomsky, 'Context-free grammars and pushdown storage', MIT Research Laboratory of Electronics Quarterly Progress Report 65 — MIT, 1962.
  4. Evey, 'Application of pushdown store machines', Proceedings of the Fall Joint Computer Conference — AFIPS, 1963.
  5. Every construction, trace and decomposition in this deck was checked by hand against the corresponding derivation or computation. — Verified 2026-08-08.

Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108