Chomsky Normal Form

Lesson 15 makes grammars uniform. It gives the two permitted rule shapes and the single exception for the start variable, and the derivation-length property that follows. It then removes useless variables, computing generating before reachable; removes epsilon-rules by adding one alternative per subset of nullable positions, and explains why every subset is needed; removes unit rules by closing over unit pairs, cycles included; and lifts terminals and binarizes long rules. The full four-pass pipeline follows, with a justification for each ordering constraint. It ends with what the form enables: cubic-time parsing, the pumping lemma's height bound, and the decidability results of Lesson 24.

Subject: Theory of Computation · 112 slides · symbolic lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Chomsky Normal Form

Title

Theory of Computation · Lesson 15

Every grammar can be rewritten so that each rule has one of exactly two shapes. Four cleanup passes, in a fixed order that cannot be shuffled.

2. What you will be able to do

Objectives

Lessons 13 and 14 wrote grammars in whatever shape was natural. This lesson makes them uniform, which every later algorithm depends on. By the end you can:

  1. State the two rule shapes allowed in Chomsky normal form, and the one exception.
  2. Remove useless variables, in the right order.
  1. Remove epsilon-rules by generating every subset of nullable positions.
  2. Remove unit rules by following unit chains to their ends.
  1. Binarize long rules and lift terminals, completing the conversion.
  2. Run the whole pipeline in the correct order, and say why each ordering constraint exists.

3. What survived from Grammar Design & Ambiguity?

Warm-up

Discussion prompt

Before we open Chomsky Normal Form: without looking back, what was the main idea of Grammar Design & Ambiguity, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Lesson 14 takes up what happens when one string has two structures. It defines ambiguity on parse trees rather than on derivations and explains why the distinction matters, then works the arithmetic grammar and its two separate defects, repairing precedence by layering variables and associativity by one-sided recursion. It covers the dangling-else problem with all three standard resolutions and the matched-unmatched grammatical repair, and the habits that keep you from introducing ambiguity by accident. It then treats inherent ambiguity with the standard overlapping-union witness, the undecidability of the ambiguity question, and how to read a parser generator's conflict report, closing with a diagnostic checklist that maps each symptom to its repair.

4. Why a Normal Form

Section

Section 1

5. Grammars vary too much for uniform proofs

Concept

Nothing in Lesson 13 constrained the shape of a rule. A right-hand side may be empty, a single variable, a single terminal, or a long mixture.

That freedom is convenient for writing grammars and awkward for everything else. A proof by induction on a derivation must handle every shape, and an algorithm must branch on all of them.

A normal form fixes the shapes in advance, so proofs and algorithms have a small fixed number of cases. The language generated is unchanged; only the presentation is standardized.

\[ L(G) = L(G') \quad\text{with}\quad G' \text{ in normal form} \]

6. Break it if you can: Grammars vary too much for uniform proofs

Counterexample

Discussion prompt

Nothing in Lesson 13 constrained the shape of a rule. A right-hand side may be empty, a single variable, a single terminal, or a long mixture.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

A normal form fixes the shapes in advance, so proofs and algorithms have a small fixed number of cases. The language generated is unchanged; only the presentation is standardized.

7. The two allowed shapes

Concept

Chomsky normal form — A grammar in which every rule sends a variable either to exactly two variables, or to exactly one terminal.

\[ A \to BC \qquad\text{or}\qquad A \to a \]

Neither shape may produce the empty string, and neither may mix variables with terminals. So a grammar in this form cannot generate the empty string at all.

One exception is allowed, and only one: the start variable may go to the empty string, provided the start variable never appears on any right-hand side.

\[ S \to \varepsilon \quad\text{permitted, if } S \text{ appears on no right-hand side} \]

8. By analogy: The two allowed shapes

Analogy

Discussion prompt

Explain The two allowed shapes by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Neither shape may produce the empty string, and neither may mix variables with terminals. So a grammar in this form cannot generate the empty string at all.

9. What the two shapes buy

Intuition

The restriction looks arbitrary until you see what becomes true once it holds.

The second is the one that makes algorithms possible: the parsing algorithm of Lesson 17 and the pumping lemma of Lesson 18 both depend on knowing a tree's height from the string's length.

10. Teach it back: What the two shapes buy

Explain it

Discussion prompt

Explain What the two shapes buy to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The restriction looks arbitrary until you see what becomes true once it holds.

11. What has to happen first: Count the steps in a normal-form derivation

Ranking

Put in order

Put the moves of Count the steps in a normal-form derivation into the order they have to happen.

  1. Count what each rule shape contributes
  2. Track the variable count
  3. Count the terminal rules
  4. Solve for the binary rules
  5. Verify the total on a short string

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. A binary rule replaces one variable by two, raising the variable count by one and producing no terminals.

12. Count the steps in a normal-form derivation

Worked example

Establish the length property, since so much later work rests on it.

\[ A \to BC \quad\text{or}\quad A \to a \]

Count what each rule shape contributes

Why: A binary rule replaces one variable by two, raising the variable count by one and producing no terminals. A terminal rule removes one variable and produces one terminal.

Track the variable count

Why: The derivation starts with one variable and must end with none. Each binary rule adds one, and each terminal rule removes one.

Count the terminal rules

Why: Exactly one terminal rule is used per symbol of the final string, so a string of length n uses n of them.

\[ \#\text{terminal rules} = n \]

Solve for the binary rules

Why: Starting at one variable and ending at zero, the additions must balance the removals: one plus the binary count equals the terminal count.

\[ 1 + \#\text{binary rules} = n \;\Rightarrow\; \#\text{binary rules} = n-1 \]

Verify the total on a short string

Why: A string of length three needs three terminal rules and two binary ones, so five steps in total. Deriving any three-symbol string in a normal-form grammar confirms this — the count never varies, which is exactly the property the later algorithms exploit.

\[ \text{total steps} = 2n - 1 \ \checkmark \]

13. Count the steps in a normal-form derivation — line by line

Picture it

Animation

Shows: Each line of the worked example "Count the steps in a normal-form derivation", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: A string of length three needs three terminal rules and two binary ones, so five steps in total. Deriving any three-symbol string in a normal-form grammar confirms this — the count never varies, which is exactly the property the later algorithms exploit.

14. The conversion pipeline

Concept

Getting an arbitrary grammar into this form takes four passes, and the order is not negotiable.

  1. Remove useless variables — those that cannot be reached, or cannot produce terminals.
  2. Remove epsilon-rules, except possibly one for a fresh start variable.
  3. Remove unit rules, which send one variable straight to another.
  4. Binarize long rules and lift terminals out of mixed right-hand sides.

Sections 2 to 5 take these one at a time, and Section 6 explains why each ordering constraint exists. Running them in the wrong order does not merely waste effort — it can leave rules the form forbids.

15. Rebuild the recipe: Preparing a grammar for conversion

Ranking

Put in order

These are the steps of Preparing a grammar for conversion, scrambled. Put them back in order before the next slide shows you.

  1. Add a fresh start variable with a single rule sending it to the old start.
  2. Check that no variable name is reused for two purposes.
  3. Write every barred line as separate rules, so nothing is miscounted.
  4. Note which variables can derive the empty string, before removing anything.
  5. Keep the original grammar beside you, so the language can be checked at each stage.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

16. Preparing a grammar for conversion

Pattern

Two preliminaries make the four passes go smoothly.

  1. Add a fresh start variable with a single rule sending it to the old start.
  2. Check that no variable name is reused for two purposes.
  3. Write every barred line as separate rules, so nothing is miscounted.
  4. Note which variables can derive the empty string, before removing anything.
  5. Keep the original grammar beside you, so the language can be checked at each stage.

The fresh start variable is what makes the empty-string exception usable. Because nothing points at it, giving it an empty alternative cannot affect any other rule.

17. Rule out three: Check yourself: the form

Elimination

Eliminate the wrong options

Which rule is allowed in Chomsky normal form?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. A → BC
  • B. A → aB
  • C. A → B
  • D. A → aBc

Survives elimination: A

Why: Two variables on the right is one of the two permitted shapes. The other is a single terminal, and nothing else is allowed except one empty rule for an unreferenced start variable.

18. Check yourself: the form

Check

Recall exactly which two shapes are allowed.

Check your understanding

Which rule is allowed in Chomsky normal form?

  • A. A → BC (correct)
  • B. A → aB
  • C. A → B
  • D. A → aBc

Answer: A

Why: Two variables on the right is one of the two permitted shapes. The other is a single terminal, and nothing else is allowed except one empty rule for an unreferenced start variable.

Why B tempts people
This mixes a terminal with a variable, which the form forbids. The terminal must be lifted into its own variable, as Section 5 does.
Why C tempts people
A single variable on the right is a unit rule, removed in Section 4. It carries no structure and would break the derivation-length property.
Why D tempts people
Three symbols on the right, and mixed at that. Binarization in Section 5 splits long right-hand sides into pairs.

19. Normal form is a tool, not a better grammar

Intuition

A converted grammar is almost always larger and much harder to read than the one it came from.

Variables multiply, meaningful names disappear, and the structure that made the original grammar readable is flattened into binary pieces. Nobody writes grammars in this form by hand.

It exists so that proofs and algorithms have a fixed shape to work with. Convert, use the result, and keep the original for human consumption.

\[ \text{readable grammar} \;\longrightarrow\; \text{normal form} \;\longrightarrow\; \text{algorithm} \]

20. Removing Useless Variables

Section

Section 2

21. Why the empty string needs an exception

Concept

The two permitted shapes both produce at least one symbol, so a grammar obeying them strictly cannot generate the empty string at all.

That would make the normal form unable to express languages containing the empty string, which is too severe a restriction. Hence the single exception.

\[ A \to BC \;\text{ and }\; A \to a \;\Rightarrow\; \text{every derived string is nonempty} \]

The exception is kept harmless by requiring the start variable to appear on no right-hand side. Then its empty alternative can only fire as the very first and only step, so it adds exactly the empty string and affects nothing else.

22. Plan first: Check the exception cannot leak

Step zero

Discussion prompt

Check the exception cannot leak — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Suppose the start variable appeared on a right-hand side

Answer:

  1. Suppose the start variable appeared on a right-hand side
  2. See what that would break
  3. Apply the condition
  4. Conclude the empty alternative fires at most once
  5. Verify the derivation-length property survives

23. Check the exception cannot leak

Worked example

Confirm that the start-variable condition really does confine the empty alternative.

Suppose the start variable appeared on a right-hand side

Why: Then some rule could place it in the middle of a sentential form, where its empty alternative would delete it mid-derivation.

See what that would break

Why: A variable vanishing mid-derivation is exactly an epsilon-rule in effect, so the derivation-length property would fail and the parsing algorithms would break.

\[ A \to SB, \; S \to \varepsilon \;\Rightarrow\; A \Longrightarrow^{*} B \]

Apply the condition

Why: Because the start variable appears nowhere on the right, no rule can introduce it after the first step.

Conclude the empty alternative fires at most once

Why: It can only apply to the initial sentential form, which is the start variable alone. Applying it finishes the derivation immediately.

Verify the derivation-length property survives

Why: Every nonempty string is derived without ever using the empty alternative, so the count of twice the length minus one still holds for them. The empty string is the single special case, derived in one step.

\[ |w| > 0 \;\Rightarrow\; 2|w|-1 \text{ steps}; \quad |w| = 0 \;\Rightarrow\; 1 \text{ step} \ \checkmark \]

24. Check the exception cannot leak — line by line

Picture it

Animation

Shows: Each line of the worked example "Check the exception cannot leak", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Every nonempty string is derived without ever using the empty alternative, so the count of twice the length minus one still holds for them. The empty string is the single special case, derived in one step.

25. Two ways a variable can be useless

Concept

A variable earns its place only if it can appear in the derivation of some finished string. Two independent things can go wrong.

  1. Non-generating: the variable cannot derive any string of terminals at all.
  2. Unreachable: no derivation from the start variable ever produces it.

A variable is useful exactly when it is both generating and reachable. Removing the others changes no derivation of any terminal string, so the language is untouched.

\[ \text{useful} \iff \text{generating} \;\text{and}\; \text{reachable} \]

26. Guess the shape of the answer: Find the generating variables

Estimation

Predict first

Compute the generating set by working upward from the terminals.

Commit before you compute: what does Find the generating variables come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the language is unchanged

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The deleted rules could never contribute to a finished string, because any derivation using them would still contain the non-generating variable forever.

27. Find the generating variables

Worked example

Compute the generating set by working upward from the terminals.

\[ S \to AB \;\mid\; a, \qquad A \to b, \qquad B \to BB \]

Start with the variables that go straight to terminals

Why: Any variable with an alternative consisting only of terminals is immediately generating.

\[ \text{generating} \supseteq \{S, A\} \]

Close the set upward

Why: Add any variable having an alternative all of whose symbols are terminals or already-generating variables. Repeat until nothing new appears.

Check the remaining variable

Why: Its only alternative mentions itself, and it is not yet known to be generating, so nothing is added. A second pass adds nothing either.

\[ B \to BB \;: \; \text{never bottoms out} \]

Remove the non-generating variable and every rule mentioning it

Why: Deleting it removes the rule that produced it and also the start rule that used it.

\[ S \to a, \qquad A \to b \]

Verify the language is unchanged

Why: The deleted rules could never contribute to a finished string, because any derivation using them would still contain the non-generating variable forever. The original grammar generates only the single symbol a, and so does the reduced one.

\[ L(G) = \{a\} = L(G') \ \checkmark \]

28. Find the generating variables — line by line

Picture it

Animation

Shows: Each line of the worked example "Find the generating variables", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Its only alternative mentions itself, and it is not yet known to be generating, so nothing is added. A second pass adds nothing either.

29. State the rule before it runs: Find the reachable variables

Hypothesis

Predict first

Find the reachable variables is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Start with the start variable

Why: It is reachable by definition, in zero steps.

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

30. Find the reachable variables

Worked example

Compute the reachable set by working downward from the start variable.

\[ S \to AB, \qquad A \to a, \qquad B \to b, \qquad C \to c \]

Start with the start variable

Why: It is reachable by definition, in zero steps.

\[ \text{reachable} \supseteq \{S\} \]

Close the set downward

Why: For each reachable variable, add every variable appearing on the right-hand side of any of its rules. Repeat until nothing new appears.

\[ \text{reachable} = \{S, A, B\} \]

Identify what is left out

Why: The fourth variable appears on no reachable right-hand side, so no derivation from the start variable ever produces it.

Remove it and its rules

Why: Deleting it removes only rules that no derivation could ever have used.

Verify the language is unchanged

Why: Every derivation from the start variable in the original grammar avoids the removed variable entirely, so every terminal string derivable before is derivable after — and no new ones appear, since rules were only deleted.

\[ L(G) = \{ab\} = L(G') \ \checkmark \]

31. Find the reachable variables — line by line

Picture it

Animation

Shows: Each line of the worked example "Find the reachable variables", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Every derivation from the start variable in the original grammar avoids the removed variable entirely, so every terminal string derivable before is derivable after — and no new ones appear, since rules were only deleted.

32. Something is wrong here: removing unreachable variables first

Anomaly

Predict first

A student writes this, and it looks reasonable:

Remove the useless variables from a grammar where both defects are present.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Start from the start variable and collect everything reachable.

Remove the useless variables from a grammar where both defects are present.

Why: Start from the start variable and collect everything reachable. Every variable here is reachable, so nothing is removed.

33. Trap: removing unreachable variables first

Trap

The trap

Remove the useless variables from a grammar where both defects are present.

Compute reachability first

Why: Start from the start variable and collect everything reachable. Every variable here is reachable, so nothing is removed.

\[ S \to AB, \qquad A \to a, \qquad B \to BB \]

Now remove the non-generating variable

Why: The looping variable cannot produce terminals, so delete it and the rules mentioning it.

\[ A \to a \]

Notice what is left over

Why: The start variable's only rule mentioned the deleted variable, so the start variable now has no rules — and the surviving variable is no longer reachable from it. Useless variables remain.

The fix

Remove the useless variables from a grammar where both defects are present.

Compute generating first

Why: The looping variable cannot bottom out, so it is non-generating. Delete it and every rule mentioning it, including the start rule.

\[ A \to a \qquad\text{(start variable now has no rules)} \]

Now compute reachability, on the reduced grammar

Why: With the start rule gone, nothing is reachable but the start variable itself, so the remaining variable is removed too.

Read the result correctly

Why: The grammar generates nothing at all, which is right: the original could never finish a derivation either. Doing generating first exposes that; doing reachability first hides it.

\[ L(G) = \varnothing \ \checkmark \]

34. Decode the notation: Trap: removing unreachable variables first

Notation

Annotate

From Trap: removing unreachable variables first — read this one piece at a time. What is each part doing?

On: \( S \to AB, \qquad A \to a, \qquad B \to BB \)

  • Start from the start variable and collect everything reachable. Every variable here is reachable, so nothing is removed.
  • The looping variable cannot produce terminals, so delete it and the rules mentioning it.
  • The start variable's only rule mentioned the deleted variable, so the start variable now has no rules — and the surviving variable is no longer reachable from it. Useless variables remain.

35. Why generating must come first

Concept

The ordering constraint is not a convention. Removing non-generating variables can make other variables unreachable, but not the other way round.

Deleting a non-generating variable also deletes the rules that mentioned it, which can cut the only path to some other variable. So reachability must be recomputed afterwards.

Deleting an unreachable variable removes only rules no derivation used, so it cannot change whether any surviving variable is generating. One order needs a second pass and the other does not.

\[ \text{generating, then reachable} \;\Rightarrow\; \text{one pass each} \]

36. Without one step: The useless-variable pass

Constraint

Discussion prompt

Run The useless-variable pass with this step confiscated:

Compute the reachable set on what remains: start from the start variable and close downward.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Compute the generating set: start from variables with all-terminal alternatives, and close upward.
  2. Delete every non-generating variable and every rule mentioning one.
  3. Compute the reachable set on what remains: start from the start variable and close downward.
  4. Delete every unreachable variable and its rules.
  5. If the start variable was deleted or has no rules, the language is empty.

37. The useless-variable pass

Pattern

Two closures, in a fixed order.

  1. Compute the generating set: start from variables with all-terminal alternatives, and close upward.
  2. Delete every non-generating variable and every rule mentioning one.
  3. Compute the reachable set on what remains: start from the start variable and close downward.
  4. Delete every unreachable variable and its rules.
  5. If the start variable was deleted or has no rules, the language is empty.

Both closures terminate because each adds at least one variable per round and there are finitely many variables.

38. Where does it stop working: The useless-variable pass

Edge cases

Discussion prompt

The useless-variable pass works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

Both closures terminate because each adds at least one variable per round and there are finitely many variables.

39. Answer it before you see the options: Check yourself: useless variables

Prediction

Predict first

Why must non-generating variables be removed before unreachable ones?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Deleting non-generating variables can make other variables unreachable

Why: Removing a non-generating variable also removes the rules mentioning it, which can sever the only route to some other variable. Computing reachability afterwards catches those; computing it first would miss them.

40. Check yourself: useless variables

Check

Think about which deletion can invalidate the other computation.

Check your understanding

Why must non-generating variables be removed before unreachable ones?

  • A. Deleting non-generating variables can make other variables unreachable (correct)
  • B. Unreachable variables are always also non-generating
  • C. Reachability cannot be computed until the grammar is smaller
  • D. The order does not matter; either works

Answer: A

Why: Removing a non-generating variable also removes the rules mentioning it, which can sever the only route to some other variable. Computing reachability afterwards catches those; computing it first would miss them.

Why B tempts people
The two defects are independent. A variable can be perfectly able to produce terminals and yet never be reached from the start.
Why C tempts people
Both closures are straightforward on a grammar of any size, and neither depends on the other having run for computability reasons.
Why D tempts people
The trap on the previous slides exhibits a grammar where the wrong order leaves useless variables behind, so the order genuinely matters.

41. Removing Epsilon-Rules

Section

Section 3

42. Nullable variables

Concept

nullable variable — A variable that can derive the empty string, in any number of steps.

A variable is nullable if it has an empty alternative outright, or an alternative consisting entirely of nullable variables. That is a closure computation, exactly like the generating one.

\[ A \to \varepsilon \quad\text{or}\quad A \to B_1\cdots B_k \text{ with every } B_i \text{ nullable} \]

Note that nullability is about can, not must. A variable with several alternatives is nullable if even one route reaches the empty string.

43. Plan first: Compute the nullable set

Step zero

Discussion prompt

Compute the nullable set — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Start with the outright empty alternatives

Answer:

  1. Start with the outright empty alternatives
  2. Close upward
  3. Check no further rounds add anything
  4. Note the consequence for the language
  5. Verify nullability by deriving the empty string

44. Compute the nullable set

Worked example

Find every variable that can vanish.

\[ S \to AB, \qquad A \to aA \;\mid\; \varepsilon, \qquad B \to bB \;\mid\; \varepsilon \]

Start with the outright empty alternatives

Why: Two variables have an empty alternative, so both are immediately nullable.

\[ \text{nullable} \supseteq \{A, B\} \]

Close upward

Why: The start variable's only alternative consists of two variables, both now known nullable, so it is nullable too.

\[ \text{nullable} = \{S, A, B\} \]

Check no further rounds add anything

Why: Every variable is already in the set, so the closure has finished.

Note the consequence for the language

Why: The start variable is nullable, so the empty string is in the language and the conversion will need the start-variable exception.

\[ \varepsilon \in L(G) \]

Verify nullability by deriving the empty string

Why: Applying the start rule then both empty alternatives gives the empty string in three steps, confirming the closure computed the right set rather than over-approximating.

\[ S \Longrightarrow AB \Longrightarrow B \Longrightarrow \varepsilon \ \checkmark \]

45. Compute the nullable set — line by line

Picture it

Animation

Shows: Each line of the worked example "Compute the nullable set", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Every variable is already in the set, so the closure has finished.

46. Nullable is not the same as useless

Intuition

Both notions were computed by upward closure, and it is easy to conflate them. They are unrelated.

PropertyMeansConsequence
nullablecan derive the empty stringmust be anticipated in every rule mentioning it
non-generatingcannot derive any terminal stringdelete it
unreachablenever appears from the startdelete it

A nullable variable is perfectly useful — it just has the empty string among its options. Deleting it would lose every string that used one of its other alternatives.

47. Fill in: Means for Nullable is not the same as useless

Comparison

Comparison matrix

From Nullable is not the same as useless: refill the Means column from what you know. The rest of the table is as it appeared.

PropertyMeansConsequence
nullablecan derive the empty stringmust be anticipated in every rule mentioning it
non-generatingcannot derive any terminal stringdelete it
unreachablenever appears from the startdelete it

48. What has to be given first: Distinguish nullable from non-generating

Missing information

Discussion prompt

A grammar with one of each, to keep the two computations apart.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The middle variable has an empty alternative, so it is nullable. The last has no route to the empty string, and neither does the start variable, since the last one blocks it.

49. Distinguish nullable from non-generating

Worked example

A grammar with one of each, to keep the two computations apart.

\[ S \to AB, \qquad A \to a \;\mid\; \varepsilon, \qquad B \to BB \]

Compute the nullable set

Why: The middle variable has an empty alternative, so it is nullable. The last has no route to the empty string, and neither does the start variable, since the last one blocks it.

\[ \text{nullable} = \{A\} \]

Compute the generating set

Why: The middle variable produces a terminal outright. The last cannot bottom out, so it is non-generating, and the start variable depends on it and is non-generating too.

\[ \text{generating} = \{A\} \]

Note that one variable is both

Why: The middle variable is nullable and generating. Those are independent properties and nothing prevents both holding.

Apply the right treatment to each

Why: The non-generating variable is deleted along with every rule mentioning it. The nullable one is kept and its nullability is anticipated in the epsilon-removal pass.

Verify the outcome

Why: After deleting the non-generating variable, the start rule goes with it and the grammar generates nothing. The nullable variable survives but is now unreachable and is removed in the second half of the pass — for a completely different reason than the first deletion.

\[ L(G) = \varnothing \ \checkmark \]

50. Distinguish nullable from non-generating — line by line

Picture it

Animation

Shows: Each line of the worked example "Distinguish nullable from non-generating", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: After deleting the non-generating variable, the start rule goes with it and the grammar generates nothing. The nullable variable survives but is now unreachable and is removed in the second half of the pass — for a completely different reason than the first deletion.

51. The replacement rule

Concept

Epsilon-rules are removed by anticipating every way a nullable variable might vanish, rather than letting it vanish later.

For each rule, and each subset of the nullable variables on its right-hand side, add a new alternative with exactly that subset deleted. Then delete every empty alternative.

\[ A \to BC \;\;\text{with both nullable} \;\;\Longrightarrow\;\; A \to BC \;\mid\; B \;\mid\; C \]

The one alternative that must not be added is the fully deleted one when it would leave an empty right-hand side — that is the epsilon-rule being removed.

52. What has to happen first: Remove the epsilon-rules

Ranking

Put in order

Put the moves of Remove the epsilon-rules into the order they have to happen.

  1. Expand the start rule
  2. Expand the recursive rules
  3. Delete the empty alternatives
  4. Restore the empty string via a fresh start variable
  5. Verify the language is unchanged on three strings

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Both variables on the right are nullable, so add alternatives with each deleted, and with both deleted — the last of which would be empty and is therefore dropped.

53. Remove the epsilon-rules

Worked example

Apply the replacement to the grammar whose nullable set was just computed.

\[ S \to AB, \qquad A \to aA \;\mid\; \varepsilon, \qquad B \to bB \;\mid\; \varepsilon \]

Expand the start rule

Why: Both variables on the right are nullable, so add alternatives with each deleted, and with both deleted — the last of which would be empty and is therefore dropped.

\[ S \to AB \;\mid\; A \;\mid\; B \]

Expand the recursive rules

Why: In each, the single variable on the right is nullable, so add an alternative with it deleted.

\[ A \to aA \;\mid\; a, \qquad B \to bB \;\mid\; b \]

Delete the empty alternatives

Why: Both original empty alternatives are removed. No variable can now vanish.

Restore the empty string via a fresh start variable

Why: The old start variable was nullable, so the empty string was in the language and must be put back — on a fresh start variable that appears on no right-hand side.

\[ S_0 \to S \;\mid\; \varepsilon \]

Verify the language is unchanged on three strings

Why: The empty string comes from the fresh start's empty alternative. The string a comes from the start rule's single-variable alternative followed by the new terminal alternative. The string ab comes from the two-variable alternative. All three were derivable before, and no new strings appeared.

\[ \varepsilon,\ a,\ b,\ ab \in L(G') = L(G) \ \checkmark \]

54. Remove the epsilon-rules — line by line

Picture it

Animation

Shows: Each line of the worked example "Remove the epsilon-rules", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The empty string comes from the fresh start's empty alternative. The string a comes from the start rule's single-variable alternative followed by the new terminal alternative. The string ab comes from the two-variable alternative. All three were derivable before, and no new strings appeared.

55. Why every subset is needed

Intuition

It is tempting to add just one alternative per nullable variable rather than one per subset. That loses strings.

With two nullable variables on a right-hand side, a derivation may vanish either one, both, or neither. Four possibilities, so four alternatives are needed — minus the empty one.

\[ 2^{k} \text{ subsets for } k \text{ nullable positions} \]

Adding fewer would make some previously derivable strings underivable. The exponential is real, and it is the reason this pass can blow up the grammar's size.

56. The cost of this pass

Concept

Each pass in the pipeline has a size cost, and this one is the worst.

PassSize effect
useless variablesshrinks the grammar
epsilon-rulescan multiply rules exponentially in the nullable count per rule
unit rulescan multiply rules by the number of variables
binarize and liftgrows linearly

In practice the exponential is mild, because right-hand sides are short and few positions are nullable. But it is genuine, and it is why binarizing before this pass is sometimes preferred in implementations.

57. Check yourself: epsilon-removal

Check

Count the subsets of nullable positions.

Check your understanding

A rule sends A to BCD, and all three of B, C and D are nullable. How many alternatives replace it?

  • A. Seven (correct)
  • B. Eight
  • C. Four
  • D. Three

Answer: A

Why: Each of the three positions may be kept or deleted, giving eight subsets. The subset deleting all three would leave an empty right-hand side, which is exactly the epsilon-rule being eliminated, so it is dropped — leaving seven.

Why B tempts people
Eight counts every subset including the fully deleted one, which would reintroduce an epsilon-rule and defeat the purpose of the pass.
Why C tempts people
Four would be right for two nullable positions. With three the count doubles again.
Why D tempts people
Three would be one alternative per nullable variable, which loses every string where two or three vanish together.

58. Removing Unit Rules

Section

Section 4

59. Unit rules carry no structure

Concept

unit rule — A rule whose right-hand side is a single variable.

Such a rule consumes no input and produces no branching. It exists only to rename, and in a parse tree it produces a chain of nodes each with a single child.

\[ A \to B \]

Normal form forbids them because they break the derivation-length property: a chain of unit rules adds steps without adding symbols, so a string of length n would no longer have a fixed step count.

60. Term to definition: Chomsky Normal Form

Matching

Match the pairs

Match each term to the definition this lesson gave it — not the one you would guess from the word.

  • t1. Chomsky normal form
  • t2. nullable variable
  • t3. unit rule
  • d1. A grammar in which every rule sends a variable either to exactly two variables, or to exactly one terminal.
  • d2. A variable that can derive the empty string, in any number of steps.
  • d3. A rule whose right-hand side is a single variable.

Why: These are the working definitions of Chomsky normal form, nullable variable, unit rule as Chomsky Normal Form uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.

61. Unit pairs

Concept

Removal works by following each chain of unit rules to its end, then short-circuiting.

Say one variable unit-derives another when the first reaches the second using unit rules alone, in any number of steps including zero.

\[ A \Longrightarrow^{*}_{\mathrm{unit}} B \]

Collect all such pairs by closure. Then for each pair, copy every non-unit alternative of the second variable onto the first, and finally delete every unit rule.

62. Guess the shape of the answer: Remove the unit rules

Estimation

Predict first

Apply the procedure to a layered arithmetic grammar, which is full of unit rules by construction.

Commit before you compute: what does Remove the unit rules come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the language and the structure are preserved

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Deriving the mixed expression still forces the multiplication deeper, because the expression variable's own alternative places the term variable on the right.

63. Remove the unit rules

Worked example

Apply the procedure to a layered arithmetic grammar, which is full of unit rules by construction.

\[ E \to E+T \;\mid\; T, \qquad T \to T\times F \;\mid\; F, \qquad F \to n \]

Find the unit pairs

Why: Each variable unit-derives itself in zero steps. Beyond that, the expression variable reaches the term variable and, through it, the factor variable.

\[ (E,T), \; (E,F), \; (T,F) \]

Copy the non-unit alternatives upward

Why: The expression variable gains the term variable's non-unit alternative and the factor variable's, and the term variable gains the factor variable's.

\[ E \to E+T \;\mid\; T\times F \;\mid\; n \]

Do the same for the middle level

Why: It gains the factor variable's only non-unit alternative.

\[ T \to T\times F \;\mid\; n \]

Delete every unit rule

Why: The alternatives sending one variable straight to another are removed. The bottom variable keeps its terminal alternative.

\[ F \to n \]

Verify the language and the structure are preserved

Why: Deriving the mixed expression still forces the multiplication deeper, because the expression variable's own alternative places the term variable on the right. So precedence survives, and the derived strings are unchanged — only the renaming steps are gone.

\[ n + n \times n \in L(G') = L(G) \ \checkmark \]

64. Remove the unit rules — line by line

Picture it

Animation

Shows: Each line of the worked example "Remove the unit rules", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Deriving the mixed expression still forces the multiplication deeper, because the expression variable's own alternative places the term variable on the right. So precedence survives, and the derived strings are unchanged — only the renaming steps are gone.

65. What is lost, and why it does not matter

Intuition

Removing unit rules flattens the layered structure that Lesson 14 worked to build. That feels like a regression, and in one sense it is.

The converted grammar no longer names the precedence levels, so it reads worse. But the trees still have the right shape, because the copied alternatives only ever appear where the original chain could have reached.

This is the clearest instance of the general point: normal form is for algorithms, and the readable grammar is kept separately. Nobody debugs a precedence bug in the converted grammar.

\[ \text{same trees, worse names} \]

66. Plan first: Handle a cycle of unit rules

Step zero

Discussion prompt

Handle a cycle of unit rules — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Compute the unit pairs

Answer:

  1. Compute the unit pairs
  2. Find the non-unit alternatives
  3. Copy it to every variable
  4. Delete every unit rule
  5. Verify the language is unchanged and the cycle caused no trouble

67. Handle a cycle of unit rules

Worked example

Unit rules can form a loop, and the closure handles it without special-casing.

\[ A \to B, \qquad B \to C, \qquad C \to A \;\mid\; a \]

Compute the unit pairs

Why: Following the chain around the cycle, every variable unit-derives every other, including itself.

\[ \text{all nine ordered pairs} \]

Find the non-unit alternatives

Why: Only one alternative in the whole grammar is not a unit rule: the terminal alternative on the third variable.

Copy it to every variable

Why: Each variable that unit-derives the third one gains its terminal alternative, which is all three of them.

\[ A \to a, \qquad B \to a, \qquad C \to a \]

Delete every unit rule

Why: The whole cycle disappears, since every rule in it was a unit rule.

Verify the language is unchanged and the cycle caused no trouble

Why: The original grammar generated exactly the single symbol a, by going round the cycle any number of times and then taking the terminal alternative. The converted grammar generates it directly. The closure terminated because there are finitely many pairs, so the cycle needed no special handling.

\[ L(G) = \{a\} = L(G') \ \checkmark \]

68. Handle a cycle of unit rules — line by line

Picture it

Animation

Shows: Each line of the worked example "Handle a cycle of unit rules", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The original grammar generated exactly the single symbol a, by going round the cycle any number of times and then taking the terminal alternative. The converted grammar generates it directly. The closure terminated because there are finitely many pairs, so the cycle needed no special handling.

69. Why unit removal comes after epsilon removal

Concept

The second ordering constraint in the pipeline, and it has the same character as the first.

Removing epsilon-rules can create unit rules. A rule sending a variable to two others, one of them nullable, gains an alternative with that one deleted — leaving a single variable on the right.

\[ A \to BC, \; C \text{ nullable} \;\Longrightarrow\; A \to BC \;\mid\; B \]

So unit removal must come second, or the units it creates would survive. The reverse cannot happen: removing unit rules never creates an epsilon-rule, because only non-empty alternatives are copied.

70. How sure are you: Check yourself: unit rules

Commit first

Predict first

Why must epsilon-rules be removed before unit rules?

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: Epsilon-removal can create new unit rules, but unit-removal creates no epsilon-rules

Why: Deleting a nullable variable from a two-variable right-hand side leaves one variable — a unit rule. So epsilon-removal must run first, or those newly created units would never be eliminated. The reverse creation cannot occur, since unit-removal only copies non-empty alternatives.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

71. Check yourself: unit rules

Check

Think about what each pass can create.

Check your understanding

Why must epsilon-rules be removed before unit rules?

  • A. Epsilon-removal can create new unit rules, but unit-removal creates no epsilon-rules (correct)
  • B. Unit rules cannot be found until the grammar has no epsilon-rules
  • C. Unit removal would delete the nullable variables
  • D. The two passes are independent and the order is arbitrary

Answer: A

Why: Deleting a nullable variable from a two-variable right-hand side leaves one variable — a unit rule. So epsilon-removal must run first, or those newly created units would never be eliminated. The reverse creation cannot occur, since unit-removal only copies non-empty alternatives.

Why B tempts people
Unit rules are found by inspection at any stage; the issue is which ones exist by the time the pass runs.
Why C tempts people
Unit removal deletes only unit rules. It never removes a variable, nullable or otherwise.
Why D tempts people
Running them in the wrong order leaves unit rules in the final grammar, which the normal form forbids, so the order is forced.

72. Binarizing and Lifting Terminals

Section

Section 5

73. What remains after three passes

Concept

After the first three passes, every right-hand side is either a single terminal, or a string of two or more symbols with no empty and no unit alternatives left.

Two defects can still remain, and the last pass fixes both.

  1. Right-hand sides longer than two symbols.
  2. Right-hand sides mixing terminals with variables.

Both are handled by introducing fresh variables, and neither can reintroduce a defect the earlier passes removed.

74. Lifting terminals

Concept

For each terminal appearing in a right-hand side of length two or more, introduce a fresh variable whose only rule produces that terminal, and substitute it in.

\[ A \to aB \;\Longrightarrow\; A \to X_a B, \qquad X_a \to a \]

One fresh variable per terminal suffices for the whole grammar — they can be shared. After this pass, every right-hand side of length two or more consists only of variables.

Right-hand sides that were already a single terminal are left alone, since that is one of the two permitted shapes.

75. Binarizing long rules

Concept

A right-hand side of three or more variables is split by introducing fresh variables that peel two symbols at a time.

\[ A \to B_1B_2B_3B_4 \;\Longrightarrow\; A \to B_1Y_1, \; Y_1 \to B_2Y_2, \; Y_2 \to B_3B_4 \]

Each fresh variable is used by exactly one rule, so no ambiguity is introduced. A right-hand side of length k needs k minus two fresh variables.

After this pass every rule has exactly two variables or exactly one terminal on the right, which is the normal form.

76. What has to be given first: Complete a conversion

Missing information

Discussion prompt

Take a grammar that has already been through the first three passes and finish the job.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Both right-hand sides have length two or more and contain terminals, so introduce one fresh variable per terminal and substitute.

77. Complete a conversion

Worked example

Take a grammar that has already been through the first three passes and finish the job.

\[ S \to aSb \;\mid\; ab \]

Lift the terminals

Why: Both right-hand sides have length two or more and contain terminals, so introduce one fresh variable per terminal and substitute.

\[ S \to X_a S X_b \;\mid\; X_a X_b, \qquad X_a \to a, \qquad X_b \to b \]

Check what still violates the form

Why: The second alternative now has exactly two variables and is finished. The first has three and must be split.

Binarize the long rule

Why: Peel the first symbol and hand the remaining two to a fresh variable.

\[ S \to X_a Y, \qquad Y \to S X_b \]

Collect the finished grammar

Why: Five rules, each with either two variables or one terminal on the right.

\[ S \to X_a Y \;\mid\; X_a X_b, \; Y \to S X_b, \; X_a \to a, \; X_b \to b \]

Verify the language and the step count

Why: Deriving the four-symbol string aabb uses the converted rules and yields the same string as before. Its derivation takes seven steps, matching the predicted count of twice the length minus one — confirming both the language and the normal-form property.

\[ aabb \in L(G'), \quad 2 \cdot 4 - 1 = 7 \text{ steps} \ \checkmark \]

78. Complete a conversion — line by line

Picture it

Animation

Shows: Each line of the worked example "Complete a conversion", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The second alternative now has exactly two variables and is finished. The first has three and must be split.

79. Something is wrong here: lifting terminals from single-terminal rules

Anomaly

Predict first

A student writes this, and it looks reasonable:

Convert a grammar containing a rule sending a variable directly to one terminal.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Apply the lifting rule everywhere a terminal appears, including in the single-terminal rule, for consistency.

Convert a grammar containing a rule sending a variable directly to one terminal.

Why: Apply the lifting rule everywhere a terminal appears, including in the single-terminal rule, for consistency.

80. Trap: lifting terminals from single-terminal rules

Trap

The trap

Convert a grammar containing a rule sending a variable directly to one terminal.

Lift every terminal uniformly

Why: Apply the lifting rule everywhere a terminal appears, including in the single-terminal rule, for consistency.

\[ A \to a \;\Longrightarrow\; A \to X_a, \qquad X_a \to a \]

Inspect the result

Why: The rewritten rule now sends one variable to one variable — a unit rule, which the form forbids and which the earlier pass was meant to have eliminated.

\[ A \to X_a \;\text{ is a unit rule} \]

The fix

Convert a grammar containing a rule sending a variable directly to one terminal.

Lift only from right-hand sides of length two or more

Why: A right-hand side that is a single terminal is already one of the two permitted shapes, so it is left exactly as it is.

\[ A \to a \;\text{ — already legal, leave it} \]

Confirm no unit rule appears

Why: Nothing was rewritten, so nothing became a unit rule and the third pass's work is preserved.

\[ A \to a \;\text{ has the shape } A \to a \ \checkmark \]

81. Which of these survive contact with Chomsky Normal Form?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Nothing in Lesson 13 constrained the shape of a rule. A right-hand side may be empty, a single variable, a single terminal, or a long mixture.; Neither shape may produce the empty string, and neither may mix variables with terminals. So a grammar in this form cannot generate the empty string at all.; The restriction looks arbitrary until you see what becomes true once it holds.
Breaks
Remove the useless variables from a grammar where both defects are present.; Convert a grammar containing a rule sending a variable directly to one terminal.
sound
These are stated as this lesson states them — each one survives the edge cases Chomsky Normal Form puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

82. Guess the shape of the answer: Convert a grammar from scratch

Estimation

Predict first

Run the whole pipeline on a small grammar with every defect present.

Commit before you compute: what does Convert a grammar from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the result is in normal form and generates the same strings

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every surviving rule has two variables or one terminal on the right, plus the permitted empty alternative on the untouched fresh start variable.

83. Convert a grammar from scratch

Worked example

Run the whole pipeline on a small grammar with every defect present.

\[ S \to ASA \;\mid\; aB, \qquad A \to B \;\mid\; S, \qquad B \to b \;\mid\; \varepsilon \]

Add a fresh start variable

Why: Nothing points at the new one, so the empty-string exception will be available if needed.

\[ S_0 \to S \]

Remove epsilon-rules

Why: The last variable is nullable, and so is the middle one through it. Add alternatives with each nullable occurrence deleted, then drop the empty alternatives.

\[ S \to ASA \;\mid\; SA \;\mid\; AS \;\mid\; S \;\mid\; aB \;\mid\; a \]

Remove unit rules

Why: Several appeared, including one the previous pass created. Copy each variable's non-unit alternatives along its unit chains, then delete the units.

\[ A \to b \;\mid\; ASA \;\mid\; SA \;\mid\; AS \;\mid\; aB \;\mid\; a \]

Lift terminals and binarize

Why: Introduce a fresh variable for each terminal appearing in a longer right-hand side, then split every right-hand side longer than two.

\[ X_a \to a, \quad X_b \to b, \quad Y \to S A \]

Verify the result is in normal form and generates the same strings

Why: Every surviving rule has two variables or one terminal on the right, plus the permitted empty alternative on the untouched fresh start variable. Deriving the shortest few strings gives the same set as the original grammar, so the four passes preserved the language throughout.

\[ \text{every rule: } A \to BC \;\text{ or }\; A \to a \ \checkmark \]

84. Convert a grammar from scratch — line by line

Picture it

Animation

Shows: Each line of the worked example "Convert a grammar from scratch", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Every surviving rule has two variables or one terminal on the right, plus the permitted empty alternative on the untouched fresh start variable. Deriving the shortest few strings gives the same set as the original grammar, so the four passes preserved the language throughout.

85. Answer it before you see the options: Check yourself: binarizing

Prediction

Predict first

How many fresh variables does binarizing a right-hand side of five variables require?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Three

Why: Each fresh variable absorbs one symbol from the front, and the last two symbols are handled by the final rule without a new variable. A right-hand side of length k therefore needs k minus two, which is three here.

86. Check yourself: binarizing

Check

Count what a long right-hand side needs.

Check your understanding

How many fresh variables does binarizing a right-hand side of five variables require?

  • A. Three (correct)
  • B. Five
  • C. Four
  • D. Two

Answer: A

Why: Each fresh variable absorbs one symbol from the front, and the last two symbols are handled by the final rule without a new variable. A right-hand side of length k therefore needs k minus two, which is three here.

Why B tempts people
One per symbol would be needed if each symbol required its own variable, but the last rule handles two symbols at once.
Why C tempts people
Four would be one per symbol beyond the first. The final pair still shares a single rule, so one fewer is needed.
Why D tempts people
Two would suffice for a right-hand side of length four. Length five needs one more.

87. The Pipeline and What It Enables

Section

Section 6

88. Conversion preserves the language, not the trees

Concept

Every pass was checked for language preservation. None preserved parse trees, and that distinction matters.

Removing unit rules collapses chains of nodes; binarizing inserts fresh internal nodes; epsilon-removal deletes leaves. A string's tree in the converted grammar generally looks nothing like its tree in the original.

\[ L(G) = L(G') \quad\text{but}\quad T_G(w) \neq T_{G'}(w) \]

So an algorithm that parses with the converted grammar and needs the original structure must map the tree back afterwards. Real parsers keep that mapping, and it is why they do not simply discard the source grammar.

89. Ambiguity is preserved, though

Intuition

One property does survive every pass, and it is worth knowing which.

If the original grammar was unambiguous, the converted one is too, and conversely. Each pass sets up a one-to-one correspondence between trees, even though it changes their shape.

That is what makes the form usable for the ambiguity results of Lesson 14 and the parsing algorithms ahead: converting cannot introduce a second reading of a string, nor silently remove one.

\[ G \text{ ambiguous} \iff G' \text{ ambiguous} \]

90. Rebuild the recipe: The complete conversion algorithm

Ranking

Put in order

These are the steps of The complete conversion algorithm, scrambled. Put them back in order before the next slide shows you.

  1. Add a fresh start variable pointing at the old one.
  2. Remove useless variables: non-generating first, then unreachable.
  3. Remove epsilon-rules, adding one alternative per subset of nullable positions.
  4. Remove unit rules, copying non-unit alternatives along unit chains.
  5. Lift terminals out of long right-hand sides, then binarize.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

91. The complete conversion algorithm

Pattern

All four passes, with the preliminaries, in the only order that works.

  1. Add a fresh start variable pointing at the old one.
  2. Remove useless variables: non-generating first, then unreachable.
  3. Remove epsilon-rules, adding one alternative per subset of nullable positions.
  4. Remove unit rules, copying non-unit alternatives along unit chains.
  5. Lift terminals out of long right-hand sides, then binarize.

If the original grammar generated the empty string, add the single permitted empty alternative to the fresh start variable at the very end.

92. Where this shows up: Chomsky Normal Form

Real world

Discussion prompt

Outside this lesson: where does Chomsky Normal Form actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The complete conversion algorithm is doing the work in it.

Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.

Answer:

Lesson 15 makes grammars uniform. Covers the two permitted rule shapes and the single start-variable exception, the derivation-length property that follows, removing useless variables with generating computed before reachable, removing epsilon-rules by adding one alternative per subset of nullable positions and why every subset is needed, removing unit rules by closing over unit pairs including cycles, lifting terminals and binarizing long rules, and the full four-pass pipeline with a justification for each ordering constraint.

93. Why the order is forced

Concept

Three of the four ordering constraints have been met individually. Collected, they explain the whole pipeline.

ConstraintReason
generating before reachabledeleting non-generating rules can sever reachability
epsilon before unitepsilon-removal creates unit rules
unit before binarizebinarizing a unit rule would leave it a unit rule
useless firstthe later passes waste effort on variables that will be deleted

The last constraint is about efficiency rather than correctness, but it matters: running epsilon-removal on a grammar full of useless variables can multiply rules that are about to be thrown away.

94. Teach it back: Why the order is forced

Explain it

Discussion prompt

Explain Why the order is forced to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Three of the four ordering constraints have been met individually. Collected, they explain the whole pipeline.

95. Each pass can undo a later one's precondition

Intuition

The pipeline is fragile in one direction and robust in the other, and knowing which is which makes the order memorable.

Reading the passes in order, each one can create defects that only earlier passes handle — never later ones. That is why a single forward sweep suffices and no iteration is needed.

Reading them backwards, each pass would need the next to run again. So the order is not a preference; it is the unique order in which one sweep terminates.

\[ \text{one forward sweep} \;\Rightarrow\; \text{no iteration needed} \]

96. By analogy: Each pass can undo a later one's precondition

Analogy

Discussion prompt

Explain Each pass can undo a later one's precondition by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The pipeline is fragile in one direction and robust in the other, and knowing which is which makes the order memorable.

97. What normal form makes possible

Concept

Three later results depend on this form, and none of them would work without it.

Each of these would be far harder to state, and some impossible to prove, against an arbitrary grammar. That is the whole return on four passes of tedium.

98. Break it if you can: What normal form makes possible

Counterexample

Discussion prompt

Each of these would be far harder to state, and some impossible to prove, against an arbitrary grammar. That is the whole return on four passes of tedium.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

99. Plan first: Use the form to bound tree height

Step zero

Discussion prompt

Use the form to bound tree height — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Note the branching factor

Answer:

  1. Note the branching factor
  2. Bound the yield by the height
  3. Invert the bound
  4. Note which direction matters later
  5. Verify the bound on a concrete case

100. Use the form to bound tree height

Worked example

A first payoff: relate the height of a parse tree to the length of the string it yields.

Note the branching factor

Why: Every internal node has at most two children, since every rule has at most two symbols on the right.

Bound the yield by the height

Why: A binary tree of height h has at most two to the h leaves, so the string it yields has at most that many symbols.

\[ |w| \le 2^{h} \]

Invert the bound

Why: Taking logarithms, a string of length n forces the tree to have height at least the logarithm of n.

\[ h \ge \log_2 |w| \]

Note which direction matters later

Why: Lesson 18 uses the contrapositive: a long string forces a tall tree, and a tall tree must repeat a variable on some root-to-leaf path. That repetition is what the pumping lemma pumps.

Verify the bound on a concrete case

Why: A string of length four needs a tree of height at least two, and the converted grammar from Section 5 does produce a tree of height three for the four-symbol string — comfortably above the bound, as required.

\[ |w| = 4 \;\Rightarrow\; h \ge 2 \ \checkmark \]

101. Use the form to bound tree height — line by line

Picture it

Animation

Shows: Each line of the worked example "Use the form to bound tree height", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: A string of length four needs a tree of height at least two, and the converted grammar from Section 5 does produce a tree of height three for the four-symbol string — comfortably above the bound, as required.

102. Rule out three: Check yourself: the pipeline

Elimination

Eliminate the wrong options

You remove unit rules and then remove epsilon-rules. What can go wrong?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Epsilon-removal creates new unit rules, which now survive into the final grammar
  • B. Unit-removal creates new epsilon-rules
  • C. The language changes
  • D. Nothing — both orders work

Survives elimination: A

Why: Deleting a nullable variable from a two-symbol right-hand side leaves one variable, which is a unit rule. With unit-removal already finished, nothing eliminates those, so the result violates the normal form.

103. Check yourself: the pipeline

Check

Recall which pass creates which defect.

Check your understanding

You remove unit rules and then remove epsilon-rules. What can go wrong?

  • A. Epsilon-removal creates new unit rules, which now survive into the final grammar (correct)
  • B. Unit-removal creates new epsilon-rules
  • C. The language changes
  • D. Nothing — both orders work

Answer: A

Why: Deleting a nullable variable from a two-symbol right-hand side leaves one variable, which is a unit rule. With unit-removal already finished, nothing eliminates those, so the result violates the normal form.

Why B tempts people
Unit-removal copies only non-empty alternatives, so it can never produce an empty right-hand side.
Why C tempts people
Both passes preserve the language individually, so the wrong order gives a correct language in the wrong shape — which is exactly why the error is easy to miss.
Why D tempts people
The result would contain unit rules, which the form forbids, so the wrong order genuinely fails.

104. What has to happen first: Bound the size of the converted grammar

Ranking

Put in order

Put the moves of Bound the size of the converted grammar into the order they have to happen.

  1. Count what useless-variable removal does
  2. Count what epsilon-removal does
  3. Count what unit-removal does
  4. Count what binarizing does
  5. Verify the overall bound is polynomial for bounded right-hand sides

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. It only deletes, so the grammar shrinks or stays the same.

105. Bound the size of the converted grammar

Worked example

Knowing roughly how much a conversion grows is useful before running one on a large grammar.

Count what useless-variable removal does

Why: It only deletes, so the grammar shrinks or stays the same.

Count what epsilon-removal does

Why: A rule whose right-hand side has m nullable positions becomes up to two to the m alternatives. With right-hand sides bounded by a constant, this is a constant factor.

\[ |R| \;\longmapsto\; |R| \cdot 2^{m} \]

Count what unit-removal does

Why: Each variable can gain the non-unit alternatives of every variable it unit-derives, so the count can grow by a factor of the number of variables.

\[ |R| \;\longmapsto\; |R| \cdot |V| \]

Count what binarizing does

Why: A right-hand side of length k becomes k minus one rules, so this is linear in the total length of all right-hand sides.

Verify the overall bound is polynomial for bounded right-hand sides

Why: With right-hand side length bounded by a constant, the epsilon pass contributes a constant factor and the unit pass a factor of the variable count, so the final grammar is polynomial in the original. The exponential only appears when right-hand sides are long and heavily nullable.

\[ |R'| = O\big(|R| \cdot |V|\big) \;\text{ for bounded right-hand sides} \ \checkmark \]

106. Bound the size of the converted grammar — line by line

Picture it

Animation

Shows: Each line of the worked example "Bound the size of the converted grammar", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: With right-hand side length bounded by a constant, the epsilon pass contributes a constant factor and the unit pass a factor of the variable count, so the final grammar is polynomial in the original. The exponential only appears when right-hand sides are long and heavily nullable.

107. When to convert and when not to

Intuition

Conversion is a means, not an end, and running it unnecessarily costs readability for nothing.

TaskConvert?
proving a theorem about all context-free languagesyes — the two cases make induction tractable
running a table-based parsing algorithmyes — the algorithm requires binary rules
writing a grammar for a language specno — keep it readable
debugging an ambiguityno — work on the original, where the structure is visible
building a practical parserusually no — parser generators accept general grammars

The last row is worth noting: real parser generators do their own internal transformations and do not require this form. Chomsky normal form is primarily a theoretical instrument.

108. Fill in: Convert? for When to convert and when not to

Comparison

Comparison matrix

From When to convert and when not to: refill the Convert? column from what you know. The rest of the table is as it appeared.

TaskConvert?
proving a theorem about all context-free languagesyes — the two cases make induction tractable
running a table-based parsing algorithmyes — the algorithm requires binary rules
writing a grammar for a language specno — keep it readable
debugging an ambiguityno — work on the original, where the structure is visible
building a practical parserusually no — parser generators accept general grammars

109. Other normal forms

Concept

Chomsky's is the one this course uses, but it is not the only one, and the alternatives are worth recognizing.

FormRule shapeUsed for
Chomskytwo variables, or one terminalparsing tables, pumping lemma
Greibachone terminal followed by variablesconverting to a stack machine
binary onlyat most two symbols, mixed allowedimplementations, avoids lifting

Greibach normal form matters for Lesson 17, where a grammar is turned into a machine: having a terminal at the front of every right-hand side means each rule consumes exactly one input symbol.

110. Fill in: Rule shape for Other normal forms

Comparison

Comparison matrix

From Other normal forms: refill the Rule shape column from what you know. The rest of the table is as it appeared.

FormRule shapeUsed for
Chomskytwo variables, or one terminalparsing tables, pumping lemma
Greibachone terminal followed by variablesconverting to a stack machine
binary onlyat most two symbols, mixed allowedimplementations, avoids lifting

111. Connect it up: Chomsky Normal Form

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why a Normal Form · Removing Useless Variables · Removing Epsilon-Rules · Removing Unit Rules · Binarizing and Lifting Terminals · The Pipeline and What It Enables. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

112. What you can do now

Recap

You can put any grammar into a uniform shape, which every algorithm in the following lessons assumes.

SituationMove
a grammar to feed an algorithmrun the four passes in order
a variable that never bottoms outnon-generating — delete it first
a right-hand side with nullable partsone alternative per subset, minus the empty one
a chain of renaming rulesclose over unit pairs, copy non-unit alternatives
a right-hand side of length kk minus two fresh variables

Lesson 16 introduces the machine that recognizes exactly this class — a finite automaton with a stack — and Lesson 17 proves it equivalent to the grammar.

Sources

  1. Sipser, Introduction to the Theory of Computation, 3rd ed., Ch. 2.1, Theorem 2.9 (Chomsky normal form) — Cengage, 2013.
  2. Hopcroft, Motwani & Ullman, Introduction to Automata Theory, Languages, and Computation, 3rd ed., Ch. 7.1 (Normal forms for context-free grammars) — Pearson, 2007.
  3. Chomsky, 'On certain formal properties of grammars', Information and Control 2(2) — Elsevier, 1959.
  4. Greibach, 'A new normal-form theorem for context-free phrase structure grammars', Journal of the ACM 12(1) — ACM, 1965.
  5. Every closure computation, replacement and conversion in this deck was carried out by hand and checked against the original grammar's language. — Verified 2026-08-08.

Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108