Lesson 14 takes up what happens when one string has two structures. It defines ambiguity on parse trees rather than on derivations and explains why the distinction matters, then works the arithmetic grammar and its two separate defects, repairing precedence by layering variables and associativity by one-sided recursion. It covers the dangling-else problem with all three standard resolutions and the matched-unmatched grammatical repair, and the habits that keep you from introducing ambiguity by accident. It then treats inherent ambiguity with the standard overlapping-union witness, the undecidability of the ambiguity question, and how to read a parser generator's conflict report, closing with a diagnostic checklist that maps each symptom to its repair.
Subject: Theory of Computation · 117 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Theory of Computation · Lesson 14
When one string has two structures, it has two meanings. Precedence, associativity, and the languages where no amount of rewriting can help.
Objectives
Lesson 13 built grammars and drew parse trees. This lesson asks what happens when a string has more than one tree. By the end you can:
Warm-up
Discussion prompt
Before we open Grammar Design & Ambiguity: without looking back, what was the main idea of Context-Free Grammars, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Lesson 13 introduces the first model that generates rather than recognizes. It covers variables, terminals, production rules, and the start variable, then one-step rewriting and why "context-free" means the surroundings are ignored, the language of a grammar, and the distinction between a sentential form and a member. It gives the four-tuple and the bar shorthand, then derivations, including the leftmost and rightmost disciplines and why they agree, and parse trees and yields with the many-derivations-one-tree relationship. A design recipe built on describing shapes recursively follows, with worked grammars for matched counts, palindromes, unions, and equal counts in any order. It closes by proving that every regular language is context-free via right-linear grammars, and showing that the containment is strict.
Section
Section 1
Concept
Lesson 13 showed that one tree usually has several derivations, differing only in scheduling. Ambiguity is a different phenomenon entirely.
ambiguous grammar — A grammar that assigns some string of its language two or more distinct parse trees.
The definition is about trees, not derivations. Counting derivations would count scheduling differences, which carry no information.
\[ G \text{ ambiguous} \iff \exists w \in L(G) \text{ with two distinct parse trees} \]
Counterexample
Discussion prompt
Lesson 13 showed that one tree usually has several derivations, differing only in scheduling. Ambiguity is a different phenomenon entirely.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The definition is about trees, not derivations. Counting derivations would count scheduling differences, which carry no information.
Intuition
It is worth being precise about this, because the two counts differ enormously and only one of them matters.
| Object | Count for a typical string | Carries information? |
|---|---|---|
| derivations | many | no — only scheduling |
| leftmost derivations | one per tree | yes |
| parse trees | one, or several | yes |
Because each tree has exactly one leftmost derivation, ambiguity can equivalently be defined by counting leftmost derivations. That is the form most textbooks use, and the two definitions are interchangeable.
\[ \text{two trees} \iff \text{two leftmost derivations} \]
Comparison
Comparison matrix
From Why trees and not derivations: refill the Count for a typical string column from what you know. The rest of the table is as it appeared.
| Object | Count for a typical string | Carries information? |
|---|---|---|
| derivations | many | no — only scheduling |
| leftmost derivations | one per tree | yes |
| parse trees | one, or several | yes |
Ranking
Put in order
Put the moves of Exhibit an ambiguous grammar into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Apply the doubling rule at the root so the left child covers the first two symbols and the right child covers the third.
Worked example
The shortest ambiguous grammar worth studying, and a string with two trees.
\[ S \to SS \;\mid\; a \]
Take the string of three a's and find two different structures for it.
Build a tree grouping to the left
Why: Apply the doubling rule at the root so the left child covers the first two symbols and the right child covers the third.
\[ (aa)a \]
Build a tree grouping to the right
Why: Apply the same rule so the left child covers the first symbol and the right child covers the last two.
\[ a(aa) \]
Confirm both are legitimate
Why: Each uses only rules of the grammar, each is rooted at the start variable, and each yields the same three-symbol string.
Confirm the two trees differ
Why: The root's left child spans one symbol in one tree and two in the other, so the trees are not the same object.
Verify the grammar is therefore ambiguous by definition
Why: A string of the language has been given two distinct parse trees, which is exactly what the definition requires. Note that the language itself — nonempty runs of a — is perfectly ordinary and has unambiguous grammars; the fault is in this grammar, not the language.
\[ aaa \in L(G), \quad T_1 \neq T_2, \quad \mathrm{yield}(T_1) = \mathrm{yield}(T_2) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Exhibit an ambiguous grammar", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: A string of the language has been given two distinct parse trees, which is exactly what the definition requires. Note that the language itself — nonempty runs of a — is perfectly ordinary and has unambiguous grammars; the fault is in this grammar, not the language.
Concept
The previous example had an ambiguous grammar for an unambiguous language. Keeping the two notions apart is essential.
| Statement | About |
|---|---|
| this grammar is ambiguous | the grammar |
| this language is inherently ambiguous | the language, and every grammar for it |
Most ambiguity met in practice is the first kind and can be repaired by rewriting the grammar. The second kind exists but is rare, and Section 5 gives an example.
\[ \text{ambiguous } G \;\not\Rightarrow\; \text{inherently ambiguous } L(G) \]
Trade off
Comparison matrix
From Ambiguity is a property of the grammar: every row here is a choice with a cost. Fill the About column, then say which row you would actually pick and what you give up for it.
| Statement | About |
|---|---|
| this grammar is ambiguous | the grammar |
| this language is inherently ambiguous | the language, and every grammar for it |
Step zero
Discussion prompt
Repair the ambiguity by rewriting — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Diagnose the cause
Answer:
Worked example
Fix the previous grammar without changing its language.
\[ S \to SS \;\mid\; a \]
Diagnose the cause
Why: The doubling rule lets the split point fall anywhere, so a string of several symbols can be cut in several places. Nothing forces a canonical cut.
Force a direction
Why: Replace the symmetric rule by one that always peels a single symbol from the left, leaving the rest to recurse.
\[ S \to aS \;\mid\; a \]
Check the language is unchanged
Why: Both grammars generate exactly the nonempty runs of a. The new one produces a run by peeling one symbol at a time until a single symbol remains.
Check the new grammar is unambiguous
Why: At each step the only choice is whether to stop, and that is determined by whether symbols remain. So each string has exactly one tree.
Verify on the string that was ambiguous before
Why: The three-symbol string now has a single tree: peel, peel, stop. There is no alternative, because the recursive alternative always leaves exactly one symbol behind it. The repair worked and the language is identical.
\[ L(G_{\text{old}}) = L(G_{\text{new}}) = a^{+}, \quad \text{one tree per string} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Repair the ambiguity by rewriting", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Both grammars generate exactly the nonempty runs of a. The new one produces a run by peeling one symbol at a time until a single symbol remains.
Concept
Only a few rule shapes can produce two trees, and recognizing them by sight prevents most accidental ambiguity.
| Rule shape | Why it can give two trees |
|---|---|
| a variable on both sides of a symbol | the split point is unconstrained |
| two alternatives generating overlapping strings | either route reaches the same string |
| a sequencing rule plus an empty alternative | empty pieces insert anywhere |
| an optional trailing part | it can attach at several depths |
Every ambiguity in this lesson is one of these four. The repairs in Sections 3 and 4 are the corresponding fixes, taken in the same order.
Analogy
Discussion prompt
Explain Where ambiguity comes from by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Only a few rule shapes can produce two trees, and recognizing them by sight prevents most accidental ambiguity.
Estimation
Predict first
Practise diagnosing without deriving anything, using the table above.
Commit before you compute: what does Spot the ambiguity from the rules alone come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the prediction on the same string
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The right operand of a plus must be a single non-recursive variable, so it spans exactly one operand.
Worked example
Practise diagnosing without deriving anything, using the table above.
Inspect a first grammar
Why: The rule sends the variable to itself, a symbol, and itself again. The variable appears on both sides of the symbol.
\[ S \to S\,+\,S \;\mid\; n \]
Predict the ambiguity
Why: A string with two plus signs has two possible split points at the root, so it should have two trees.
Confirm by deriving
Why: The three-operand string groups either left or right, exactly as predicted.
\[ n+n+n \;\longrightarrow\; (n+n)+n \;\text{ or }\; n+(n+n) \]
Inspect a second grammar and predict nothing
Why: Here the recursive occurrence appears only on the left, and the right side descends to a non-recursive variable. No shape from the table is present.
\[ S \to S\,+\,F \;\mid\; F, \qquad F \to n \]
Verify the prediction on the same string
Why: The right operand of a plus must be a single non-recursive variable, so it spans exactly one operand. The last plus is forced to the root and the grouping is unique — matching the prediction made from the rules alone.
\[ n+n+n \;\longrightarrow\; (n+n)+n \;\text{ only} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Spot the ambiguity from the rules alone", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The right operand of a plus must be a single non-recursive variable, so it spans exactly one operand. The last plus is forced to the root and the grouping is unique — matching the prediction made from the rules alone.
Ranking
Put in order
These are the steps of How to show a grammar is ambiguous, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Demonstrating ambiguity is an exhibition, so it is short when done properly.
Showing a grammar is unambiguous is much harder, since it quantifies over all strings. It usually requires an argument that the rules force a unique choice at every step.
Elimination
Eliminate the wrong options
A grammar assigns some string five different derivations but only one parse tree. Is it ambiguous?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Several derivations of one tree differ only in the order independent variables were expanded, which carries no structural information. Ambiguity requires two distinct trees, and one tree was reported here.
Check
Recall which objects the definition counts.
Check your understanding
A grammar assigns some string five different derivations but only one parse tree. Is it ambiguous?
Answer: A
Why: Several derivations of one tree differ only in the order independent variables were expanded, which carries no structural information. Ambiguity requires two distinct trees, and one tree was reported here.
Intuition
For a pure membership question ambiguity is harmless: the string is in the language whether it has one tree or five.
It matters because the tree is usually the output. A compiler, an interpreter or a query engine consumes the tree, so two trees mean two behaviours for one input — and the choice between them is unspecified.
It also matters for efficiency. Ambiguous grammars make parsing harder, and the deterministic parsing techniques used in practice require unambiguous grammars to begin with.
\[ \text{tree} \;\longrightarrow\; \text{meaning} \;\Rightarrow\; \text{two trees} \;\longrightarrow\; \text{two meanings} \]
Explain it
Discussion prompt
Explain Why ambiguity matters in practice to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
For a pure membership question ambiguity is harmless: the string is in the language whether it has one tree or five.
Section
Section 2
Concept
Write down the obvious grammar for arithmetic expressions and it is ambiguous immediately.
\[ E \to E + E \;\mid\; E \times E \;\mid\; ( E ) \;\mid\; n \]
Each binary rule has the variable on both sides, so a string with two operators can be grouped either way. The grammar says nothing about which operator binds tighter, or how a chain of equal operators associates.
This is the canonical example because both defects appear at once, and both have standard repairs.
Missing information
Discussion prompt
Show the ambiguity concretely, and read off the two meanings.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Apply the addition rule first, so the left child is the first number and the right child is the multiplication.
Worked example
Show the ambiguity concretely, and read off the two meanings.
\[ w = n + n \times n \]
Group with the addition at the root
Why: Apply the addition rule first, so the left child is the first number and the right child is the multiplication.
\[ n + (n \times n) \]
Group with the multiplication at the root
Why: Apply the multiplication rule first, so the left child is the addition and the right child is the last number.
\[ (n + n) \times n \]
Read off the two meanings
Why: With numbers substituted, the first grouping evaluates the product before the sum, and the second evaluates the sum first. The results differ.
\[ 2 + (3 \times 4) = 14 \qquad (2 + 3) \times 4 = 20 \]
Note that both trees are legal
Why: Neither violates any rule of the grammar. The grammar simply does not express a preference, so both structures are permitted.
Verify this is genuine ambiguity rather than two spellings
Why: The two trees have different roots and different child spans, and both yield the identical string. Since they also evaluate differently, the ambiguity has real consequences — this is not a bookkeeping artefact.
\[ \mathrm{yield}(T_1) = \mathrm{yield}(T_2) = n + n \times n, \quad T_1 \neq T_2 \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Two trees for one expression", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The two trees have different roots and different child spans, and both yield the identical string. Since they also evaluate differently, the ambiguity has real consequences — this is not a bookkeeping artefact.
Concept
The grammar has two independent problems, and they need different repairs. Diagnosing which is which is the first step.
| Defect | Symptom | Repair |
|---|---|---|
| no precedence | different operators group either way | one variable per precedence level |
| no associativity | the same operator chains group either way | make each rule recursive on one side only |
Fixing precedence alone leaves chains of equal operators ambiguous, and fixing associativity alone leaves mixed operators ambiguous. Both repairs are needed.
Comparison
Comparison matrix
From Two separate defects: refill the Repair column from what you know. The rest of the table is as it appeared.
| Defect | Symptom | Repair |
|---|---|---|
| no precedence | different operators group either way | one variable per precedence level |
| no associativity | the same operator chains group either way | make each rule recursive on one side only |
Worked example
Isolate the second defect by using only one operator.
\[ E \to E + E \;\mid\; n \]
Take a chain of two additions
Why: Three numbers joined by two plus signs.
\[ w = n + n + n \]
Group to the left
Why: The root's left child covers the first sum and its right child the last number.
\[ (n + n) + n \]
Group to the right
Why: The root's left child covers the first number and its right child the second sum.
\[ n + (n + n) \]
Note that here the values agree
Why: Addition is associative, so both groupings evaluate the same. The ambiguity is still real — two trees exist — but its consequences are invisible for this operator.
Verify why the defect must still be repaired
Why: Replace addition by subtraction and the two groupings give different values, since subtraction is not associative. So the grammar is wrong in the same way for both operators, and only the visibility of the error differs.
\[ (5 - 3) - 1 = 1 \quad\text{but}\quad 5 - (3 - 1) = 3 \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Show the associativity defect separately", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Replace addition by subtraction and the two groupings give different values, since subtraction is not associative. So the grammar is wrong in the same way for both operators, and only the visibility of the error differs.
Intuition
Both notions are usually taught as rules about evaluation. In grammar terms they are constraints on which trees are allowed.
So both are enforced structurally, by controlling where each variable may appear in a rule. No evaluation rules are needed, and none are available: a grammar has no notion of computing a value.
\[ \text{lower precedence} \;\Rightarrow\; \text{nearer the root} \]
Prediction
Predict first
In an unambiguous arithmetic grammar with the usual precedence, where does the addition sit in the tree for the expression n + n × n?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: At the root, with the multiplication in its right subtree
Why: Multiplication binds tighter, so it groups first and therefore sits deeper. Addition groups last and sits at the root, with the product forming its right operand — which is exactly the grouping that evaluates the product before the sum.
Check
Think about which operator ends up at the root.
Check your understanding
In an unambiguous arithmetic grammar with the usual precedence, where does the addition sit in the tree for the expression n + n × n?
Answer: A
Why: Multiplication binds tighter, so it groups first and therefore sits deeper. Addition groups last and sits at the root, with the product forming its right operand — which is exactly the grouping that evaluates the product before the sum.
Section
Section 3
Concept
The layered grammar of the next slides fixes an order, so something must be able to override it. That is the job of the bracket rule.
The bracket alternative sits at the tightest level, and its contents return to the loosest. So inside brackets the whole hierarchy starts again, and any grouping can be forced.
\[ F \to ( E ) \;\mid\; n \]
Without that rule the grammar could express only the conventional grouping. With it, every grouping is expressible — but each string still has exactly one tree, because the brackets appear in the string and pin the structure down.
Step zero
Discussion prompt
Show brackets restore the lost grouping — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Recall what the unbracketed string forces
Answer:
Worked example
Check that the layered grammar can still express the non-conventional reading.
\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to (E) \;\mid\; n \]
Recall what the unbracketed string forces
Why: Without brackets, the multiplication is forced deeper and the sum is evaluated last.
\[ n + n \times n \;\longrightarrow\; n + (n \times n) \]
Write the other grouping explicitly
Why: Put brackets around the sum, so the multiplication must take the bracketed group as an operand.
\[ (n + n) \times n \]
Derive it
Why: The top level descends to a term, the term uses its recursive alternative, and its left operand is a factor that is a bracketed expression.
\[ E \Longrightarrow T \Longrightarrow T \times F \Longrightarrow F \times F \Longrightarrow (E) \times n \]
Note the brackets are part of the string
Why: The two readings are now different strings, not two trees for one string. That is exactly how a grammar expresses a choice without becoming ambiguous.
Verify each of the two strings has one tree
Why: The unbracketed string forces the multiplication deeper; the bracketed one forces the sum deeper because a bracketed group is a factor. Each derivation is determined at every step, so both strings are unambiguous.
\[ n+n\times n \;\text{ and }\; (n+n)\times n \;: \; \text{one tree each} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Show brackets restore the lost grouping", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The unbracketed string forces the multiplication deeper; the bracketed one forces the sum deeper because a bracketed group is a factor. Each derivation is determined at every step, so both strings are unambiguous.
Concept
The standard repair introduces a layer of variables, one for each precedence level, arranged from loosest to tightest.
| Variable | Level | Handles |
|---|---|---|
| expression | loosest | addition |
| term | middle | multiplication |
| factor | tightest | numbers and brackets |
Each level may only descend to the next, which forces tighter operators deeper into the tree. The layering is what encodes precedence.
\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to (E) \;\mid\; n \]
Sorting
Sort into buckets
These are the pieces of Grammar Design & Ambiguity, out of order. Put each one back under the part of the lesson it belongs to.
Concept
Look at where each recursive variable sits. That single choice fixes the associativity.
In the grammar above, the expression rule recurses on the left and descends on the right, so addition is left-associative. Swapping the sides would make it right-associative instead.
\[ E \to E + T \;\;\text{left} \qquad E \to T + E \;\;\text{right} \]
Worked example
Check that the repaired grammar assigns exactly one tree to the earlier problem string.
\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to (E) \;\mid\; n \]
Start at the top level
Why: The string contains a plus sign at the top level, so the expression rule must use its recursive alternative. Nothing else can produce a plus sign here.
\[ E \Longrightarrow E + T \]
Force the left operand
Why: Everything before the plus is a single number, so the left expression must descend all the way to a factor.
\[ E \Longrightarrow T \Longrightarrow F \Longrightarrow n \]
Force the right operand
Why: Everything after the plus contains a multiplication sign, so the term must use its recursive alternative.
\[ T \Longrightarrow T \times F \Longrightarrow F \times F \Longrightarrow n \times n \]
Observe there was no choice at any step
Why: At each point the symbols present determined which alternative could apply. A different choice would have produced a different string.
Verify the unique tree has the intended shape
Why: The addition sits at the root and the multiplication in its right subtree, which is the grouping that evaluates the product first. The grammar now encodes the conventional precedence structurally.
\[ n + (n \times n) \quad \text{— the only tree} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Verify the layered grammar is unambiguous", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The addition sits at the root and the multiplication in its right subtree, which is the grouping that evaluates the product first. The grammar now encodes the conventional precedence structurally.
Estimation
Predict first
Confirm the second repair took effect, on a chain of equal operators.
Commit before you compute: what does Check left-associativity on a chain come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the choice matters for a non-associative operator
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Replacing addition by subtraction, the left-leaning grammar gives the conventional reading and the right-leaning one does not.
Worked example
Confirm the second repair took effect, on a chain of equal operators.
\[ w = n + n + n \]
Apply the expression rule
Why: Its recursive occurrence is on the left, so the left child is an expression and the right child is a term.
\[ E \Longrightarrow E + T \]
Ask what the right child can span
Why: A term cannot contain a plus sign at its top level, since no term rule produces one. So the right child must be the final number alone.
Conclude the split is forced
Why: The last plus sign must be the one at the root, so the left child spans the first two numbers and their plus. The grouping leans left, and no other split is possible.
\[ (n + n) + n \]
Contrast with the right-recursive variant
Why: Had the rule been written with the recursion on the right, the same argument would force the first plus to the root, giving the opposite grouping.
Verify the choice matters for a non-associative operator
Why: Replacing addition by subtraction, the left-leaning grammar gives the conventional reading and the right-leaning one does not. So the side of the recursion is a real design decision, not a stylistic one.
\[ (5-3)-1 = 1 \;\text{(left)} \qquad 5-(3-1) = 3 \;\text{(right)} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Check left-associativity on a chain", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Replacing addition by subtraction, the left-leaning grammar gives the conventional reading and the right-leaning one does not. So the side of the recursion is a real design decision, not a stylistic one.
Pattern
Turning an ambiguous operator grammar into an unambiguous one is mechanical once the conventions are chosen.
The bracket rule returning to the loosest level is what lets parentheses override precedence — inside them, the whole hierarchy starts again.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Repair the arithmetic grammar so it is unambiguous.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Split the single variable into an expression level for addition and a term level for multiplication, so the operators no longer compete.
Repair the arithmetic grammar so it is unambiguous.
Why: Split the single variable into an expression level for addition and a term level for multiplication, so the operators no longer compete.
Trap
Repair the arithmetic grammar so it is unambiguous.
Add a level per operator
Why: Split the single variable into an expression level for addition and a term level for multiplication, so the operators no longer compete.
\[ E \to E + E \;\mid\; T, \qquad T \to T \times T \;\mid\; n \]
Declare the grammar fixed
Why: Mixed expressions now have a unique grouping, since the levels force the multiplication deeper. Precedence is enforced.
\[ n + n \times n \;\longrightarrow\; n + (n \times n) \]
Find the ambiguity that remains
Why: A chain of equal operators still groups either way, because each level's rule is recursive on both sides.
\[ n + n + n \;\longrightarrow\; (n+n)+n \;\text{ or }\; n+(n+n) \]
Repair the arithmetic grammar so it is unambiguous.
Add a level per operator, and recurse on one side only
Why: Each level keeps its recursive occurrence on the left and descends to the next level on the right. That fixes precedence and associativity together.
\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to n \]
Check the mixed case
Why: The multiplication is forced into the term level and therefore deeper, so precedence holds.
\[ n + n \times n \;\longrightarrow\; n + (n \times n) \]
Check the chain case
Why: A term cannot contain a plus sign, so the rightmost operand of an addition is always a single term. The grouping is forced to lean left, and the grammar is now unambiguous.
\[ n + n + n \;\longrightarrow\; (n+n)+n \;\text{ only} \ \checkmark \]
Notation
Annotate
From Trap: fixing precedence but forgetting associativity — read this one piece at a time. What is each part doing?
On: \( E \to E + E \;\mid\; T, \qquad T \to T \times T \;\mid\; n \)
Check
Look at which side each recursive occurrence sits on.
Check your understanding
A grammar has the rule E → T ^ E | T, where the caret is exponentiation. What does this enforce?
Answer: A
Why: The recursive occurrence sits on the right of the operator, so the right operand may itself contain another exponentiation while the left may not. That forces a chain to group rightwards, which matches the usual mathematical convention.
Section
Section 4
Ranking
Put in order
Put the moves of Add a third precedence level into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. It binds tighter than multiplication, so it sits between the term level and the atoms.
Worked example
Extend the layered grammar with exponentiation, which binds tighter than multiplication and groups to the right.
Decide where the new level goes
Why: It binds tighter than multiplication, so it sits between the term level and the atoms.
\[ E \;>\; T \;>\; P \;>\; F \]
Write the new level with right recursion
Why: Exponentiation groups to the right, so the recursive occurrence goes on the right and the descent on the left.
\[ P \to F \;\hat{\;}\; P \;\mid\; F \]
Rewire the level above it
Why: Multiplication now descends to the new level rather than straight to the atoms.
\[ T \to T \times P \;\mid\; P \]
Check a mixed expression
Why: In a string with a product and a power, the power must sit deeper because the term level can only reach it through the new variable.
Verify both conventions hold at once
Why: A chain of powers groups rightwards, because the recursion is on the right; and a power binds tighter than a product, because its level is lower. Testing a string with both confirms the intended structure is the only one derivable.
\[ n \times n \hat{\;} n \hat{\;} n \;\longrightarrow\; n \times \big(n \hat{\;} (n \hat{\;} n)\big) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Add a third precedence level", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: In a string with a product and a power, the power must sit deeper because the term level can only reach it through the new variable.
Intuition
Layering only works because precedence is a linear ordering: every pair of operators has a definite winner.
If two operators were declared to have equal precedence but different associativity, no layering could express it — they would have to share a level, and a shared level cannot recurse on two different sides at once.
Real language specifications therefore assign every operator a distinct precedence, or group equal-precedence operators that share an associativity into one level. Anything else is not expressible by this technique.
\[ \text{equal precedence} \;\Rightarrow\; \text{same level} \;\Rightarrow\; \text{same associativity} \]
Concept
The most famous ambiguity in programming languages, and it appears in a three-rule grammar.
\[ S \to \mathrm{if}\;E\;S \;\mid\; \mathrm{if}\;E\;S\;\mathrm{else}\;S \;\mid\; s \]
A statement with two conditionals and one else clause can attach the else to either conditional. Both attachments use only these rules, so the grammar permits both.
The consequence is a genuine behavioural difference: the else branch runs under different conditions depending on which conditional it belongs to.
Fill the middle
Fill in the blanks
From Exhibit the dangling-else ambiguity — finish the line. Write what belongs on the right of the equals sign before you look.
T_1 \neq T_2, \quad \mathrm\mathrm{yield}(T_2) \ \checkmark(T_1) = ___
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. The outer conditional uses the else-less alternative, and its body is a complete conditional with an else.
Worked example
Find the two trees for the shortest problematic statement.
\[ \mathrm{if}\;E\;\mathrm{if}\;E\;s\;\mathrm{else}\;s \]
Attach the else to the inner conditional
Why: The outer conditional uses the else-less alternative, and its body is a complete conditional with an else.
\[ \mathrm{if}\;E\;\big(\mathrm{if}\;E\;s\;\mathrm{else}\;s\big) \]
Attach the else to the outer conditional
Why: The outer conditional uses the with-else alternative, and its then-branch is an else-less conditional.
\[ \big(\mathrm{if}\;E\;(\mathrm{if}\;E\;s)\;\mathrm{else}\;s\big) \]
Read off the behavioural difference
Why: In the first, the else runs when the outer test succeeds and the inner fails. In the second, it runs when the outer test fails — regardless of the inner test.
Note both are structurally legal
Why: Each tree uses only the three rules and yields the same token sequence, so the grammar genuinely permits both readings.
Verify the ambiguity is not merely theoretical
Why: The two readings disagree on the case where the outer test fails: one runs the else branch and the other runs nothing. A language leaving that unspecified would have programs whose behaviour depends on the parser, so real languages must resolve it.
\[ T_1 \neq T_2, \quad \mathrm{yield}(T_1) = \mathrm{yield}(T_2) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Exhibit the dangling-else ambiguity", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The two readings disagree on the case where the outer test fails: one runs the else branch and the other runs nothing. A language leaving that unspecified would have programs whose behaviour depends on the parser, so real languages must resolve it.
Concept
Three resolutions are used in practice, and they differ in what they change.
| Resolution | Changes | Used by |
|---|---|---|
| attach to the nearest conditional | the parser, by a stated rule | C, Java, most languages |
| split into matched and unmatched statements | the grammar | textbook presentations |
| require an explicit terminator | the language itself | Ada, and many modern languages |
Only the second and third make the grammar unambiguous. The first leaves the grammar ambiguous and resolves the conflict outside it, which is why language specifications must state the rule explicitly.
Trade off
Comparison matrix
From Resolving the dangling else: every row here is a choice with a cost. Fill the Changes column, then say which row you would actually pick and what you give up for it.
| Resolution | Changes | Used by |
|---|---|---|
| attach to the nearest conditional | the parser, by a stated rule | C, Java, most languages |
| split into matched and unmatched statements | the grammar | textbook presentations |
| require an explicit terminator | the language itself | Ada, and many modern languages |
Hypothesis
Predict first
Repair the grammar by splitting the statement class is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Identify the invariant to enforce
Why: An else must attach to the nearest unmatched conditional. So a statement appearing before an else must itself be fully matched.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
The grammatical fix, which removes the ambiguity rather than deferring it.
Identify the invariant to enforce
Why: An else must attach to the nearest unmatched conditional. So a statement appearing before an else must itself be fully matched.
Split the statement variable in two
Why: One variable for statements whose conditionals all have else clauses, and one for statements that may end with an unmatched conditional.
Write the matched rules
Why: A matched statement is a plain statement, or a conditional both of whose branches are matched.
\[ M \to \mathrm{if}\;E\;M\;\mathrm{else}\;M \;\mid\; s \]
Write the unmatched rules
Why: An unmatched statement is a conditional with no else, or one whose else branch is unmatched — and whose then branch is matched.
\[ U \to \mathrm{if}\;E\;S \;\mid\; \mathrm{if}\;E\;M\;\mathrm{else}\;U, \qquad S \to M \;\mid\; U \]
Verify the problem string now has one tree
Why: The then-branch of a conditional with an else must be matched, and a bare conditional is not matched. So the else cannot attach to the outer conditional, and only the nearest-attachment tree survives — which is the intended reading.
\[ \mathrm{if}\;E\;\big(\mathrm{if}\;E\;s\;\mathrm{else}\;s\big) \quad\text{— the only tree} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Repair the grammar by splitting the statement class", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The then-branch of a conditional with an else must be matched, and a bare conditional is not matched. So the else cannot attach to the outer conditional, and only the nearest-attachment tree survives — which is the intended reading.
Intuition
Both repairs in this lesson used the same move, and it is worth naming.
When a grammar permits an unwanted grouping, split the variable into two, one for the forms allowed in the constrained position and one for the rest. Then use the constrained variable exactly where the unwanted grouping would have occurred.
The layered arithmetic grammar did this by precedence level; the dangling-else repair did it by matched and unmatched. In both cases the number of variables grew and the number of trees fell to one.
\[ \text{unwanted grouping} \;\Rightarrow\; \text{split the variable, constrain the position} \]
Step zero
Discussion prompt
Design a grammar for nested structures with two bracket kinds — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Write the shapes a member can have
Answer:
Worked example
A design where ambiguity is easy to introduce accidentally.
\[ L = \text{properly nested strings over } \{\,(,\,),\,[,\,]\,\} \]
Write the shapes a member can have
Why: Empty; a member wrapped in round brackets; a member wrapped in square brackets; or two members side by side.
\[ S \to (S) \;\mid\; [S] \;\mid\; SS \;\mid\; \varepsilon \]
Notice this grammar is ambiguous
Why: The side-by-side alternative together with the empty alternative lets a single member be split into pieces in many ways, including empty pieces.
Remove the ambiguity by forcing a direction
Why: Use a second variable for a single bracketed group, and let the top level be a sequence of them built one at a time from the left.
\[ S \to A\,S \;\mid\; \varepsilon, \qquad A \to (S) \;\mid\; [S] \]
Check the language is unchanged
Why: Every properly nested string is a sequence of bracketed groups, and each group contains a properly nested string. Both grammars generate exactly those.
Verify uniqueness on a string with two groups
Why: The string with a round pair followed by a square pair has one tree: the first group is peeled, then the second, then the sequence ends. The empty alternative can only be used once, at the very end, so no alternative split exists.
\[ ()[] \in L(G), \quad \text{one tree} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Design a grammar for nested structures with two bracket kinds", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Every properly nested string is a sequence of bracketed groups, and each group contains a properly nested string. Both grammars generate exactly those.
Constraint
Discussion prompt
Run Designing without introducing ambiguity with this step confiscated:
Give each precedence level its own variable, descending in one direction only.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Ambiguity is usually introduced by accident. These habits avoid most of it.
The second bullet is the one that catches people. A variable that can both split and vanish gives every string unboundedly many trees, since empty pieces can be inserted anywhere.
Edge cases
Discussion prompt
Designing without introducing ambiguity works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Ambiguity is usually introduced by accident. These habits avoid most of it.
Commit first
Predict first
Why is the grammar S → SS | a | ε especially badly ambiguous?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: Empty pieces can be inserted anywhere, so each string has unboundedly many trees
Why: The doubling rule can split off a piece that then derives the empty string, and that can be done any number of times at any position. So a single symbol already has infinitely many parse trees, not merely two.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Look for the combination that creates unbounded trees.
Check your understanding
Why is the grammar S → SS | a | ε especially badly ambiguous?
Answer: A
Why: The doubling rule can split off a piece that then derives the empty string, and that can be done any number of times at any position. So a single symbol already has infinitely many parse trees, not merely two.
Section
Section 5
Worked example
A pattern that appears in every real syntax, and one where the empty case is easy to get wrong.
\[ L = \text{zero or more items separated by commas} \]
Write the wrong version first
Why: The obvious attempt combines a sequencing rule with an empty alternative — the third entry in the ambiguity table.
\[ L \to L\,,\,L \;\mid\; i \;\mid\; \varepsilon \]
See the problem
Why: A single item can be derived directly, or as two pieces with one empty, so the grammar is badly ambiguous.
Separate the empty case from the nonempty one
Why: Use one variable for the possibly-empty list and another for the nonempty list, so the empty alternative can fire only once, at the top.
\[ L \to N \;\mid\; \varepsilon, \qquad N \to i\,,\,N \;\mid\; i \]
Check the separator count
Why: A nonempty list of k items has exactly k minus one commas, because each use of the recursive alternative contributes one comma and one item, and the base contributes one item and no comma.
Verify on the three shortest lists
Why: The empty list uses the empty alternative alone; a single item uses the base of the nonempty variable; two items use one recursive step then the base. Each has exactly one derivation, and no trailing comma is derivable.
\[ \varepsilon,\ i,\ i\,,\,i \in L(G) \qquad i\,, \notin L(G) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Design a grammar for a list with separators", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: A nonempty list of k items has exactly k minus one commas, because each use of the recursive alternative contributes one comma and one item, and the base contributes one item and no comma.
Concept
That repair used a move worth stating on its own, because it recurs in every list-like grammar.
Put the empty alternative on a variable that is used once, at the outermost position, and let a separate variable handle the nonempty recursion. Then emptiness is decided exactly once rather than at every level.
\[ L \to N \;\mid\; \varepsilon \qquad\text{rather than}\qquad N \to \cdots \;\mid\; \varepsilon \]
The same shape handles optional clauses, optional trailing separators, and optional whole sections. Whenever something may be absent, decide its absence at one place only.
Concept
Every ambiguity so far was the grammar's fault and could be rewritten away. That is not always possible.
inherently ambiguous language — A context-free language for which every context-free grammar is ambiguous.
The claim quantifies over all grammars, so exhibiting one bad grammar proves nothing. Establishing inherent ambiguity requires ruling out every grammar at once, much as non-regularity did in Lesson 10.
\[ \forall G \; \big(L(G) = L \;\Rightarrow\; G \text{ ambiguous}\big) \]
Concept
The usual witness is a union of two languages that overlap in an awkward way.
\[ L = \{\, a^{i}b^{j}c^{k} : i = j \;\text{ or }\; j = k \,\} \]
Each half is easy to generate: one grammar matches the a's against the b's and lets the c's run free, and the other matches the b's against the c's. Their union is generated by combining them with a fresh start variable.
The difficulty is the overlap. Strings where both conditions hold belong to both halves, and any grammar must generate them somehow — which turns out to force two trees.
\[ a^{n}b^{n}c^{n} \in \text{both halves} \]
Missing information
Discussion prompt
The full proof is beyond this course, but the obstruction is visible and worth seeing.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
One variable per half, joined by a fresh start variable with two alternatives.
Worked example
The full proof is beyond this course, but the obstruction is visible and worth seeing.
Write the obvious grammar
Why: One variable per half, joined by a fresh start variable with two alternatives.
\[ S \to S_1 \;\mid\; S_2 \]
Generate a string in the overlap
Why: A string with equal counts of all three symbols satisfies both conditions, so it can be derived through either alternative.
Observe the two trees
Why: One tree begins by choosing the first alternative and matches the a's against the b's; the other chooses the second and matches the b's against the c's. The trees differ at the root.
\[ T_1 \text{ via } S_1, \qquad T_2 \text{ via } S_2 \]
Try to repair by removing the overlap
Why: The natural fix is to make the second alternative generate only strings not already generated by the first. But that set is not context-free, so no grammar can express the restriction.
Verify why the obvious repairs all fail
Why: Every repair must either drop strings of the overlap, changing the language, or keep both routes to them, keeping the ambiguity. Since no context-free grammar can generate the difference of the two halves, no third option exists — which is the heart of the full proof.
\[ \text{overlap not removable} \;\Rightarrow\; \text{every grammar ambiguous} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "See why the overlap forces two trees", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Every repair must either drop strings of the overlap, changing the language, or keep both routes to them, keeping the ambiguity. Since no context-free grammar can generate the difference of the two halves, no third option exists — which is the heart of the full proof.
Intuition
Inherent ambiguity is a genuine phenomenon but almost never an engineering problem, and it is worth knowing why.
The witness languages are unions of two structurally different descriptions that overlap. Programming languages and data formats are not designed that way — their constructs are made distinguishable on purpose, usually with keywords or delimiters.
So when a real grammar is ambiguous, the overwhelmingly likely cause is a repairable defect: missing precedence, both-sided recursion, or an unconstrained optional part. Reach for the repairs of Section 3 before suspecting the language.
\[ \text{ambiguous grammar in practice} \;\Rightarrow\; \text{almost always repairable} \]
Concept
A natural next question: given a grammar, can a program tell whether it is ambiguous? The answer is no.
The problem of deciding whether an arbitrary context-free grammar is ambiguous is undecidable. Lesson 29 proves this by reduction from the Post Correspondence Problem.
Deciding whether a language is inherently ambiguous is undecidable too. So both questions in this section are, in general, beyond any algorithm — which is why parser generators report conflicts rather than proving ambiguity.
\[ \text{'is } G \text{ ambiguous?'} \;\text{ is undecidable} \]
Explain it
Discussion prompt
Explain Deciding ambiguity is impossible to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A natural next question: given a grammar, can a program tell whether it is ambiguous? The answer is no.
Estimation
Predict first
What a tool can and cannot tell you, given the undecidability just stated.
Commit before you compute: what does Interpret a parser generator's conflict report come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the asymmetry is forced by undecidability
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A tool that answered the ambiguity question in both directions would decide an undecidable problem.
Worked example
What a tool can and cannot tell you, given the undecidability just stated.
Understand what the tool checks
Why: A parser generator tries to build a deterministic parser for a restricted class of grammars. It reports a conflict when its construction cannot proceed.
Note what a conflict does not prove
Why: A conflict means the grammar is outside the tool's class. It does not prove the grammar is ambiguous — many unambiguous grammars are outside the class.
Note what the absence of a conflict does prove
Why: If the construction succeeds, the grammar is unambiguous, because a deterministic parser assigns at most one tree per string.
\[ \text{no conflict} \;\Rightarrow\; \text{unambiguous} \]
Read the asymmetry correctly
Why: The tool gives a one-way guarantee, exactly like the pumping lemma did. Success is conclusive; failure is a prompt to investigate.
Verify the asymmetry is forced by undecidability
Why: A tool that answered the ambiguity question in both directions would decide an undecidable problem. So the one-way guarantee is not a limitation of any particular tool but a mathematical necessity.
\[ \text{two-way answer} \;\Rightarrow\; \text{decides the undecidable} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Interpret a parser generator's conflict report", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: A tool that answered the ambiguity question in both directions would decide an undecidable problem. So the one-way guarantee is not a limitation of any particular tool but a mathematical necessity.
Prediction
Predict first
You write an ambiguous grammar for a language. What have you shown?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Nothing about the language — only that this grammar is ambiguous
Why: Inherent ambiguity quantifies over every grammar for the language, so a single ambiguous grammar is no evidence at all. The three-a example earlier in this lesson had an ambiguous grammar for a language with an easy unambiguous one.
Check
Recall what the definition quantifies over.
Check your understanding
You write an ambiguous grammar for a language. What have you shown?
Answer: A
Why: Inherent ambiguity quantifies over every grammar for the language, so a single ambiguous grammar is no evidence at all. The three-a example earlier in this lesson had an ambiguous grammar for a language with an easy unambiguous one.
Section
Section 6
Ranking
Put in order
These are the steps of Diagnosing an ambiguous grammar, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
When a grammar turns out ambiguous, the cause is usually one of four things. Check them in this order.
The fifth line is genuinely last. Inherent ambiguity is rare enough that it should be suspected only after the first four have been ruled out.
Real world
Discussion prompt
Outside this lesson: where does Grammar Design & Ambiguity actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of Diagnosing an ambiguous grammar is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
Lesson 14 takes up what happens when one string has two structures. It defines ambiguity on parse trees rather than on derivations and explains why the distinction matters, then works the arithmetic grammar and its two separate defects, repairing precedence by layering variables and associativity by one-sided recursion. It covers the dangling-else problem with all three standard resolutions and the matched-unmatched grammatical repair, and the habits that keep you from introducing ambiguity by accident. It then treats inherent ambiguity with the standard overlapping-union witness, the undecidability of the ambiguity question, and how to read a parser generator's conflict report, closing with a diagnostic checklist that maps each symptom to its repair.
Step zero
Discussion prompt
Diagnose and repair an unfamiliar grammar — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Test the shortest suspicious string
Answer:
Worked example
Apply the checklist to a grammar you have not seen.
\[ S \to S\,;\,S \;\mid\; \mathrm{stmt} \;\mid\; \varepsilon \]
Test the shortest suspicious string
Why: A single statement can be derived directly, or by splitting into two pieces one of which is empty. Two trees already.
Identify the cause from the checklist
Why: This is the third item: a sequencing rule combined with an empty alternative. The empty piece can be inserted anywhere, unboundedly often.
Apply the corresponding repair
Why: Peel one item at a time, and let a separate variable handle the empty case so it can be used only once.
\[ S \to \mathrm{stmt}\,;\,S \;\mid\; \mathrm{stmt} \;\mid\; \varepsilon \]
Re-test the problem string
Why: A single statement now matches the second alternative and nothing else, since the first requires a semicolon and the third requires no symbols at all.
Verify the language is unchanged and uniqueness now holds
Why: Both grammars generate any number of statements separated by semicolons, including none. In the repaired grammar each string determines its derivation at every step — the number of semicolons decides which alternative applies — so each has exactly one tree.
\[ L(G_{\text{old}}) = L(G_{\text{new}}), \quad \text{one tree per string} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Diagnose and repair an unfamiliar grammar", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Both grammars generate any number of statements separated by semicolons, including none. In the repaired grammar each string determines its derivation at every step — the number of semicolons decides which alternative applies — so each has exactly one tree.
Ranking
Put in order
Put the moves of Check a repaired grammar systematically into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Two operators, two list items, or two nested groups — whichever the grammar builds.
Worked example
After any repair, run this check rather than trusting the fix.
Test the shortest string with two of the same construct
Why: Two operators, two list items, or two nested groups — whichever the grammar builds. If it has one tree, the associativity defect is gone.
Test the shortest string with two different constructs
Why: One of each operator, or a group inside a sequence. If it has one tree, the precedence defect is gone.
Test the empty case
Why: Derive the empty string, if it belongs, and confirm exactly one derivation reaches it. Empty alternatives are the usual source of unbounded ambiguity.
Confirm the language did not change
Why: List the shortest few strings of the old grammar and check each is still derivable, and that no new short string appeared.
Verify the four checks are sufficient in practice
Why: Each check targets one row of the ambiguity table from Section 1, and the fourth guards against a repair that removed ambiguity by removing strings. Together they catch every defect this lesson has covered.
\[ \text{4 checks} \;\longleftrightarrow\; \text{4 causes} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Check a repaired grammar systematically", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Each check targets one row of the ambiguity table from Section 1, and the fourth guards against a repair that removed ambiguity by removing strings. Together they catch every defect this lesson has covered.
Intuition
Every repair in this lesson was checked twice: once for uniqueness and once for the language being unchanged. The second check is the one people skip.
It is easy to remove ambiguity by accidentally removing strings. Forcing one-sided recursion can drop the empty case; splitting a variable can leave one branch unreachable; adding a level can make a construct underivable.
So a repair is only correct when both halves hold: one tree per string, and the same set of strings. Reporting the first without the second is the commonest error in grammar exercises.
\[ \text{repair correct} \iff \text{unambiguous} \;\text{and}\; L \text{ unchanged} \]
Analogy
Discussion prompt
Explain Repairs must preserve the language by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Every repair in this lesson was checked twice: once for uniqueness and once for the language being unchanged. The second check is the one people skip.
Concept
Ambiguous grammars are not always wrong. It is worth knowing when to tolerate one.
| Situation | Ambiguity acceptable? |
|---|---|
| you only need membership | yes — the tree is discarded anyway |
| you need a unique meaning | no — resolve it |
| you use a general parser that returns all trees | sometimes — if downstream can choose |
| you use a deterministic parser generator | no — it will refuse to build |
Natural-language grammars are deliberately ambiguous, because natural language is; the parser returns several trees and later stages pick. Programming-language grammars are deliberately not.
Comparison
Comparison matrix
From What ambiguity costs and what it buys: refill the Ambiguity acceptable? column from what you know. The rest of the table is as it appeared.
| Situation | Ambiguity acceptable? |
|---|---|
| you only need membership | yes — the tree is discarded anyway |
| you need a unique meaning | no — resolve it |
| you use a general parser that returns all trees | sometimes — if downstream can choose |
| you use a deterministic parser generator | no — it will refuse to build |
Intuition
The question raised here reappears in a stronger form twice more before the course ends.
So this lesson's practical question turns out to have a deep answer, and the tools to give it are still four chapters away.
Counterexample
Discussion prompt
The question raised here reappears in a stronger form twice more before the course ends.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Elimination
Eliminate the wrong options
A grammar gives the string of three operands joined by two identical operators exactly two parse trees. Which repair applies?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A chain of one operator grouping two ways is the associativity defect, caused by the recursive variable appearing on both sides. Restricting the recursion to one side forces a consistent lean and leaves exactly one tree.
Check
Match the symptom to the standard cause.
Check your understanding
A grammar gives the string of three operands joined by two identical operators exactly two parse trees. Which repair applies?
Answer: A
Why: A chain of one operator grouping two ways is the associativity defect, caused by the recursive variable appearing on both sides. Restricting the recursion to one side forces a consistent lean and leaves exactly one tree.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — What Ambiguity Is · The Arithmetic Example · Removing Ambiguity · Designing Real Grammars · Inherent Ambiguity · Diagnosis and Practice. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can tell a defective grammar from a defective language, and repair the first kind.
| Situation | Move |
|---|---|
| operators grouping either way | recurse on one side only |
| different operators competing | one variable per precedence level |
| unboundedly many trees | look for sequencing plus an empty alternative |
| an optional part attaching in two places | split into matched and unmatched |
| all four repairs fail | consider inherent ambiguity — but check again first |
Lesson 15 puts grammars into a normal form, which makes the algorithms and proofs of the following lessons uniform.
Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.