Grammar Design & Ambiguity

Lesson 14 takes up what happens when one string has two structures. It defines ambiguity on parse trees rather than on derivations and explains why the distinction matters, then works the arithmetic grammar and its two separate defects, repairing precedence by layering variables and associativity by one-sided recursion. It covers the dangling-else problem with all three standard resolutions and the matched-unmatched grammatical repair, and the habits that keep you from introducing ambiguity by accident. It then treats inherent ambiguity with the standard overlapping-union witness, the undecidability of the ambiguity question, and how to read a parser generator's conflict report, closing with a diagnostic checklist that maps each symptom to its repair.

Subject: Theory of Computation · 117 slides · symbolic lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Grammar Design and Ambiguity

Title

Theory of Computation · Lesson 14

When one string has two structures, it has two meanings. Precedence, associativity, and the languages where no amount of rewriting can help.

2. What you will be able to do

Objectives

Lesson 13 built grammars and drew parse trees. This lesson asks what happens when a string has more than one tree. By the end you can:

  1. Define ambiguity precisely, in terms of trees rather than derivations.
  2. Exhibit two trees for a string, and explain why the two structures mean different things.
  1. Rewrite an arithmetic grammar to enforce precedence and associativity.
  2. Design grammars for nested and interleaved patterns without introducing ambiguity.
  1. Recognize the dangling-else problem and the standard ways of resolving it.
  2. State what inherent ambiguity is, and why it cannot be repaired.

3. What survived from Context-Free Grammars?

Warm-up

Discussion prompt

Before we open Grammar Design & Ambiguity: without looking back, what was the main idea of Context-Free Grammars, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Lesson 13 introduces the first model that generates rather than recognizes. It covers variables, terminals, production rules, and the start variable, then one-step rewriting and why "context-free" means the surroundings are ignored, the language of a grammar, and the distinction between a sentential form and a member. It gives the four-tuple and the bar shorthand, then derivations, including the leftmost and rightmost disciplines and why they agree, and parse trees and yields with the many-derivations-one-tree relationship. A design recipe built on describing shapes recursively follows, with worked grammars for matched counts, palindromes, unions, and equal counts in any order. It closes by proving that every regular language is context-free via right-linear grammars, and showing that the containment is strict.

4. What Ambiguity Is

Section

Section 1

5. The definition

Concept

Lesson 13 showed that one tree usually has several derivations, differing only in scheduling. Ambiguity is a different phenomenon entirely.

ambiguous grammar — A grammar that assigns some string of its language two or more distinct parse trees.

The definition is about trees, not derivations. Counting derivations would count scheduling differences, which carry no information.

\[ G \text{ ambiguous} \iff \exists w \in L(G) \text{ with two distinct parse trees} \]

6. Break it if you can: The definition

Counterexample

Discussion prompt

Lesson 13 showed that one tree usually has several derivations, differing only in scheduling. Ambiguity is a different phenomenon entirely.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The definition is about trees, not derivations. Counting derivations would count scheduling differences, which carry no information.

7. Why trees and not derivations

Intuition

It is worth being precise about this, because the two counts differ enormously and only one of them matters.

ObjectCount for a typical stringCarries information?
derivationsmanyno — only scheduling
leftmost derivationsone per treeyes
parse treesone, or severalyes

Because each tree has exactly one leftmost derivation, ambiguity can equivalently be defined by counting leftmost derivations. That is the form most textbooks use, and the two definitions are interchangeable.

\[ \text{two trees} \iff \text{two leftmost derivations} \]

8. Fill in: Count for a typical string for Why trees and not derivations

Comparison

Comparison matrix

From Why trees and not derivations: refill the Count for a typical string column from what you know. The rest of the table is as it appeared.

ObjectCount for a typical stringCarries information?
derivationsmanyno — only scheduling
leftmost derivationsone per treeyes
parse treesone, or severalyes

9. What has to happen first: Exhibit an ambiguous grammar

Ranking

Put in order

Put the moves of Exhibit an ambiguous grammar into the order they have to happen.

  1. Build a tree grouping to the left
  2. Build a tree grouping to the right
  3. Confirm both are legitimate
  4. Confirm the two trees differ
  5. Verify the grammar is therefore ambiguous by definition

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Apply the doubling rule at the root so the left child covers the first two symbols and the right child covers the third.

10. Exhibit an ambiguous grammar

Worked example

The shortest ambiguous grammar worth studying, and a string with two trees.

\[ S \to SS \;\mid\; a \]

Take the string of three a's and find two different structures for it.

Build a tree grouping to the left

Why: Apply the doubling rule at the root so the left child covers the first two symbols and the right child covers the third.

\[ (aa)a \]

Build a tree grouping to the right

Why: Apply the same rule so the left child covers the first symbol and the right child covers the last two.

\[ a(aa) \]

Confirm both are legitimate

Why: Each uses only rules of the grammar, each is rooted at the start variable, and each yields the same three-symbol string.

Confirm the two trees differ

Why: The root's left child spans one symbol in one tree and two in the other, so the trees are not the same object.

Verify the grammar is therefore ambiguous by definition

Why: A string of the language has been given two distinct parse trees, which is exactly what the definition requires. Note that the language itself — nonempty runs of a — is perfectly ordinary and has unambiguous grammars; the fault is in this grammar, not the language.

\[ aaa \in L(G), \quad T_1 \neq T_2, \quad \mathrm{yield}(T_1) = \mathrm{yield}(T_2) \ \checkmark \]

11. Exhibit an ambiguous grammar — line by line

Picture it

Animation

Shows: Each line of the worked example "Exhibit an ambiguous grammar", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: A string of the language has been given two distinct parse trees, which is exactly what the definition requires. Note that the language itself — nonempty runs of a — is perfectly ordinary and has unambiguous grammars; the fault is in this grammar, not the language.

12. Ambiguity is a property of the grammar

Concept

The previous example had an ambiguous grammar for an unambiguous language. Keeping the two notions apart is essential.

StatementAbout
this grammar is ambiguousthe grammar
this language is inherently ambiguousthe language, and every grammar for it

Most ambiguity met in practice is the first kind and can be repaired by rewriting the grammar. The second kind exists but is rare, and Section 5 gives an example.

\[ \text{ambiguous } G \;\not\Rightarrow\; \text{inherently ambiguous } L(G) \]

13. What each one costs: Ambiguity is a property of the grammar

Trade off

Comparison matrix

From Ambiguity is a property of the grammar: every row here is a choice with a cost. Fill the About column, then say which row you would actually pick and what you give up for it.

StatementAbout
this grammar is ambiguousthe grammar
this language is inherently ambiguousthe language, and every grammar for it

14. Plan first: Repair the ambiguity by rewriting

Step zero

Discussion prompt

Repair the ambiguity by rewriting — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Diagnose the cause

Answer:

  1. Diagnose the cause
  2. Force a direction
  3. Check the language is unchanged
  4. Check the new grammar is unambiguous
  5. Verify on the string that was ambiguous before

15. Repair the ambiguity by rewriting

Worked example

Fix the previous grammar without changing its language.

\[ S \to SS \;\mid\; a \]

Diagnose the cause

Why: The doubling rule lets the split point fall anywhere, so a string of several symbols can be cut in several places. Nothing forces a canonical cut.

Force a direction

Why: Replace the symmetric rule by one that always peels a single symbol from the left, leaving the rest to recurse.

\[ S \to aS \;\mid\; a \]

Check the language is unchanged

Why: Both grammars generate exactly the nonempty runs of a. The new one produces a run by peeling one symbol at a time until a single symbol remains.

Check the new grammar is unambiguous

Why: At each step the only choice is whether to stop, and that is determined by whether symbols remain. So each string has exactly one tree.

Verify on the string that was ambiguous before

Why: The three-symbol string now has a single tree: peel, peel, stop. There is no alternative, because the recursive alternative always leaves exactly one symbol behind it. The repair worked and the language is identical.

\[ L(G_{\text{old}}) = L(G_{\text{new}}) = a^{+}, \quad \text{one tree per string} \ \checkmark \]

16. Repair the ambiguity by rewriting — line by line

Picture it

Animation

Shows: Each line of the worked example "Repair the ambiguity by rewriting", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Both grammars generate exactly the nonempty runs of a. The new one produces a run by peeling one symbol at a time until a single symbol remains.

17. Where ambiguity comes from

Concept

Only a few rule shapes can produce two trees, and recognizing them by sight prevents most accidental ambiguity.

Rule shapeWhy it can give two trees
a variable on both sides of a symbolthe split point is unconstrained
two alternatives generating overlapping stringseither route reaches the same string
a sequencing rule plus an empty alternativeempty pieces insert anywhere
an optional trailing partit can attach at several depths

Every ambiguity in this lesson is one of these four. The repairs in Sections 3 and 4 are the corresponding fixes, taken in the same order.

18. By analogy: Where ambiguity comes from

Analogy

Discussion prompt

Explain Where ambiguity comes from by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Only a few rule shapes can produce two trees, and recognizing them by sight prevents most accidental ambiguity.

19. Guess the shape of the answer: Spot the ambiguity from the rules alone

Estimation

Predict first

Practise diagnosing without deriving anything, using the table above.

Commit before you compute: what does Spot the ambiguity from the rules alone come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the prediction on the same string

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The right operand of a plus must be a single non-recursive variable, so it spans exactly one operand.

20. Spot the ambiguity from the rules alone

Worked example

Practise diagnosing without deriving anything, using the table above.

Inspect a first grammar

Why: The rule sends the variable to itself, a symbol, and itself again. The variable appears on both sides of the symbol.

\[ S \to S\,+\,S \;\mid\; n \]

Predict the ambiguity

Why: A string with two plus signs has two possible split points at the root, so it should have two trees.

Confirm by deriving

Why: The three-operand string groups either left or right, exactly as predicted.

\[ n+n+n \;\longrightarrow\; (n+n)+n \;\text{ or }\; n+(n+n) \]

Inspect a second grammar and predict nothing

Why: Here the recursive occurrence appears only on the left, and the right side descends to a non-recursive variable. No shape from the table is present.

\[ S \to S\,+\,F \;\mid\; F, \qquad F \to n \]

Verify the prediction on the same string

Why: The right operand of a plus must be a single non-recursive variable, so it spans exactly one operand. The last plus is forced to the root and the grouping is unique — matching the prediction made from the rules alone.

\[ n+n+n \;\longrightarrow\; (n+n)+n \;\text{ only} \ \checkmark \]

21. Spot the ambiguity from the rules alone — line by line

Picture it

Animation

Shows: Each line of the worked example "Spot the ambiguity from the rules alone", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The right operand of a plus must be a single non-recursive variable, so it spans exactly one operand. The last plus is forced to the root and the grouping is unique — matching the prediction made from the rules alone.

22. Rebuild the recipe: How to show a grammar is ambiguous

Ranking

Put in order

These are the steps of How to show a grammar is ambiguous, scrambled. Put them back in order before the next slide shows you.

  1. Find a string with at least two independent opportunities to apply the same rule.
  2. Build one parse tree, grouping one way.
  3. Build a second, grouping the other way.
  4. Check both yield the same string and both use only the grammar's rules.
  5. Point out one structural difference — usually the span of a child of the root.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

23. How to show a grammar is ambiguous

Pattern

Demonstrating ambiguity is an exhibition, so it is short when done properly.

  1. Find a string with at least two independent opportunities to apply the same rule.
  2. Build one parse tree, grouping one way.
  3. Build a second, grouping the other way.
  4. Check both yield the same string and both use only the grammar's rules.
  5. Point out one structural difference — usually the span of a child of the root.

Showing a grammar is unambiguous is much harder, since it quantifies over all strings. It usually requires an argument that the rules force a unique choice at every step.

24. Rule out three: Check yourself: the definition

Elimination

Eliminate the wrong options

A grammar assigns some string five different derivations but only one parse tree. Is it ambiguous?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Not on the evidence given — ambiguity counts trees, not derivations
  • B. Yes, because five is more than one
  • C. Yes, because only ambiguous grammars have several derivations
  • D. No, because a string can never have two trees

Survives elimination: A

Why: Several derivations of one tree differ only in the order independent variables were expanded, which carries no structural information. Ambiguity requires two distinct trees, and one tree was reported here.

25. Check yourself: the definition

Check

Recall which objects the definition counts.

Check your understanding

A grammar assigns some string five different derivations but only one parse tree. Is it ambiguous?

  • A. Not on the evidence given — ambiguity counts trees, not derivations (correct)
  • B. Yes, because five is more than one
  • C. Yes, because only ambiguous grammars have several derivations
  • D. No, because a string can never have two trees

Answer: A

Why: Several derivations of one tree differ only in the order independent variables were expanded, which carries no structural information. Ambiguity requires two distinct trees, and one tree was reported here.

Why B tempts people
Counting derivations counts scheduling. Almost every string in almost every grammar has several derivations, so that count would make nearly all grammars ambiguous.
Why C tempts people
Unambiguous grammars routinely give strings many derivations. The tree is what must be unique, not the sequence.
Why D tempts people
Strings certainly can have two trees — the worked example above exhibits one. That is precisely what ambiguity means.

26. Why ambiguity matters in practice

Intuition

For a pure membership question ambiguity is harmless: the string is in the language whether it has one tree or five.

It matters because the tree is usually the output. A compiler, an interpreter or a query engine consumes the tree, so two trees mean two behaviours for one input — and the choice between them is unspecified.

It also matters for efficiency. Ambiguous grammars make parsing harder, and the deterministic parsing techniques used in practice require unambiguous grammars to begin with.

\[ \text{tree} \;\longrightarrow\; \text{meaning} \;\Rightarrow\; \text{two trees} \;\longrightarrow\; \text{two meanings} \]

27. Teach it back: Why ambiguity matters in practice

Explain it

Discussion prompt

Explain Why ambiguity matters in practice to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

For a pure membership question ambiguity is harmless: the string is in the language whether it has one tree or five.

28. The Arithmetic Example

Section

Section 2

29. The natural grammar is ambiguous

Concept

Write down the obvious grammar for arithmetic expressions and it is ambiguous immediately.

\[ E \to E + E \;\mid\; E \times E \;\mid\; ( E ) \;\mid\; n \]

Each binary rule has the variable on both sides, so a string with two operators can be grouped either way. The grammar says nothing about which operator binds tighter, or how a chain of equal operators associates.

This is the canonical example because both defects appear at once, and both have standard repairs.

30. What has to be given first: Two trees for one expression

Missing information

Discussion prompt

Show the ambiguity concretely, and read off the two meanings.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Apply the addition rule first, so the left child is the first number and the right child is the multiplication.

31. Two trees for one expression

Worked example

Show the ambiguity concretely, and read off the two meanings.

\[ w = n + n \times n \]

Group with the addition at the root

Why: Apply the addition rule first, so the left child is the first number and the right child is the multiplication.

\[ n + (n \times n) \]

Group with the multiplication at the root

Why: Apply the multiplication rule first, so the left child is the addition and the right child is the last number.

\[ (n + n) \times n \]

Read off the two meanings

Why: With numbers substituted, the first grouping evaluates the product before the sum, and the second evaluates the sum first. The results differ.

\[ 2 + (3 \times 4) = 14 \qquad (2 + 3) \times 4 = 20 \]

Note that both trees are legal

Why: Neither violates any rule of the grammar. The grammar simply does not express a preference, so both structures are permitted.

Verify this is genuine ambiguity rather than two spellings

Why: The two trees have different roots and different child spans, and both yield the identical string. Since they also evaluate differently, the ambiguity has real consequences — this is not a bookkeeping artefact.

\[ \mathrm{yield}(T_1) = \mathrm{yield}(T_2) = n + n \times n, \quad T_1 \neq T_2 \ \checkmark \]

32. Two trees for one expression — line by line

Picture it

Animation

Shows: Each line of the worked example "Two trees for one expression", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The two trees have different roots and different child spans, and both yield the identical string. Since they also evaluate differently, the ambiguity has real consequences — this is not a bookkeeping artefact.

33. Two separate defects

Concept

The grammar has two independent problems, and they need different repairs. Diagnosing which is which is the first step.

DefectSymptomRepair
no precedencedifferent operators group either wayone variable per precedence level
no associativitythe same operator chains group either waymake each rule recursive on one side only

Fixing precedence alone leaves chains of equal operators ambiguous, and fixing associativity alone leaves mixed operators ambiguous. Both repairs are needed.

34. Fill in: Repair for Two separate defects

Comparison

Comparison matrix

From Two separate defects: refill the Repair column from what you know. The rest of the table is as it appeared.

DefectSymptomRepair
no precedencedifferent operators group either wayone variable per precedence level
no associativitythe same operator chains group either waymake each rule recursive on one side only

35. Show the associativity defect separately

Worked example

Isolate the second defect by using only one operator.

\[ E \to E + E \;\mid\; n \]

Take a chain of two additions

Why: Three numbers joined by two plus signs.

\[ w = n + n + n \]

Group to the left

Why: The root's left child covers the first sum and its right child the last number.

\[ (n + n) + n \]

Group to the right

Why: The root's left child covers the first number and its right child the second sum.

\[ n + (n + n) \]

Note that here the values agree

Why: Addition is associative, so both groupings evaluate the same. The ambiguity is still real — two trees exist — but its consequences are invisible for this operator.

Verify why the defect must still be repaired

Why: Replace addition by subtraction and the two groupings give different values, since subtraction is not associative. So the grammar is wrong in the same way for both operators, and only the visibility of the error differs.

\[ (5 - 3) - 1 = 1 \quad\text{but}\quad 5 - (3 - 1) = 3 \ \checkmark \]

36. Show the associativity defect separately — line by line

Picture it

Animation

Shows: Each line of the worked example "Show the associativity defect separately", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Replace addition by subtraction and the two groupings give different values, since subtraction is not associative. So the grammar is wrong in the same way for both operators, and only the visibility of the error differs.

37. Precedence and associativity are questions about the tree

Intuition

Both notions are usually taught as rules about evaluation. In grammar terms they are constraints on which trees are allowed.

So both are enforced structurally, by controlling where each variable may appear in a rule. No evaluation rules are needed, and none are available: a grammar has no notion of computing a value.

\[ \text{lower precedence} \;\Rightarrow\; \text{nearer the root} \]

38. Answer it before you see the options: Check yourself: reading the trees

Prediction

Predict first

In an unambiguous arithmetic grammar with the usual precedence, where does the addition sit in the tree for the expression n + n × n?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: At the root, with the multiplication in its right subtree

Why: Multiplication binds tighter, so it groups first and therefore sits deeper. Addition groups last and sits at the root, with the product forming its right operand — which is exactly the grouping that evaluates the product before the sum.

39. Check yourself: reading the trees

Check

Think about which operator ends up at the root.

Check your understanding

In an unambiguous arithmetic grammar with the usual precedence, where does the addition sit in the tree for the expression n + n × n?

  • A. At the root, with the multiplication in its right subtree (correct)
  • B. At the root, with the multiplication in its left subtree
  • C. Below the multiplication, since addition is evaluated first
  • D. At the same level as the multiplication

Answer: A

Why: Multiplication binds tighter, so it groups first and therefore sits deeper. Addition groups last and sits at the root, with the product forming its right operand — which is exactly the grouping that evaluates the product before the sum.

Why B tempts people
The multiplication is written to the right of the plus sign, so it must occupy the right subtree. Placing it left would change which symbols each subtree spans.
Why C tempts people
Being evaluated first means grouping first, which puts an operator deeper in the tree, not shallower. Multiplication is the one evaluated first here.
Why D tempts people
A tree has one root, and the two operators are applied at different times, so they cannot occupy the same level.

40. Removing Ambiguity

Section

Section 3

41. Brackets are what let precedence be overridden

Concept

The layered grammar of the next slides fixes an order, so something must be able to override it. That is the job of the bracket rule.

The bracket alternative sits at the tightest level, and its contents return to the loosest. So inside brackets the whole hierarchy starts again, and any grouping can be forced.

\[ F \to ( E ) \;\mid\; n \]

Without that rule the grammar could express only the conventional grouping. With it, every grouping is expressible — but each string still has exactly one tree, because the brackets appear in the string and pin the structure down.

42. Plan first: Show brackets restore the lost grouping

Step zero

Discussion prompt

Show brackets restore the lost grouping — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Recall what the unbracketed string forces

Answer:

  1. Recall what the unbracketed string forces
  2. Write the other grouping explicitly
  3. Note the brackets are part of the string
  4. Verify each of the two strings has one tree

43. Show brackets restore the lost grouping

Worked example

Check that the layered grammar can still express the non-conventional reading.

\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to (E) \;\mid\; n \]

Recall what the unbracketed string forces

Why: Without brackets, the multiplication is forced deeper and the sum is evaluated last.

\[ n + n \times n \;\longrightarrow\; n + (n \times n) \]

Write the other grouping explicitly

Why: Put brackets around the sum, so the multiplication must take the bracketed group as an operand.

\[ (n + n) \times n \]

Derive it

Why: The top level descends to a term, the term uses its recursive alternative, and its left operand is a factor that is a bracketed expression.

\[ E \Longrightarrow T \Longrightarrow T \times F \Longrightarrow F \times F \Longrightarrow (E) \times n \]

Note the brackets are part of the string

Why: The two readings are now different strings, not two trees for one string. That is exactly how a grammar expresses a choice without becoming ambiguous.

Verify each of the two strings has one tree

Why: The unbracketed string forces the multiplication deeper; the bracketed one forces the sum deeper because a bracketed group is a factor. Each derivation is determined at every step, so both strings are unambiguous.

\[ n+n\times n \;\text{ and }\; (n+n)\times n \;: \; \text{one tree each} \ \checkmark \]

44. Show brackets restore the lost grouping — line by line

Picture it

Animation

Shows: Each line of the worked example "Show brackets restore the lost grouping", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The unbracketed string forces the multiplication deeper; the bracketed one forces the sum deeper because a bracketed group is a factor. Each derivation is determined at every step, so both strings are unambiguous.

45. One variable per precedence level

Concept

The standard repair introduces a layer of variables, one for each precedence level, arranged from loosest to tightest.

VariableLevelHandles
expressionloosestaddition
termmiddlemultiplication
factortightestnumbers and brackets

Each level may only descend to the next, which forces tighter operators deeper into the tree. The layering is what encodes precedence.

\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to (E) \;\mid\; n \]

46. Where does each piece belong: Grammar Design & Ambiguity

Sorting

Sort into buckets

These are the pieces of Grammar Design & Ambiguity, out of order. Put each one back under the part of the lesson it belongs to.

What Ambiguity Is
The definition; Why trees and not derivations; Exhibit an ambiguous grammar
The Arithmetic Example
The natural grammar is ambiguous; Two trees for one expression; Two separate defects
Removing Ambiguity
Brackets are what let precedence be overridden; Show brackets restore the lost grouping; One variable per precedence level
s1
What Ambiguity Is is where Grammar Design & Ambiguity puts The definition, Why trees and not derivations, Exhibit an ambiguous grammar. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The Arithmetic Example is where Grammar Design & Ambiguity puts The natural grammar is ambiguous, Two trees for one expression, Two separate defects. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Removing Ambiguity is where Grammar Design & Ambiguity puts Brackets are what let precedence be overridden, Show brackets restore the lost grouping, One variable per precedence level. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

47. Recursion on one side gives associativity

Concept

Look at where each recursive variable sits. That single choice fixes the associativity.

  1. The recursive occurrence on the left makes the operator group to the left.
  2. The recursive occurrence on the right makes it group to the right.

In the grammar above, the expression rule recurses on the left and descends on the right, so addition is left-associative. Swapping the sides would make it right-associative instead.

\[ E \to E + T \;\;\text{left} \qquad E \to T + E \;\;\text{right} \]

48. Verify the layered grammar is unambiguous

Worked example

Check that the repaired grammar assigns exactly one tree to the earlier problem string.

\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to (E) \;\mid\; n \]

Start at the top level

Why: The string contains a plus sign at the top level, so the expression rule must use its recursive alternative. Nothing else can produce a plus sign here.

\[ E \Longrightarrow E + T \]

Force the left operand

Why: Everything before the plus is a single number, so the left expression must descend all the way to a factor.

\[ E \Longrightarrow T \Longrightarrow F \Longrightarrow n \]

Force the right operand

Why: Everything after the plus contains a multiplication sign, so the term must use its recursive alternative.

\[ T \Longrightarrow T \times F \Longrightarrow F \times F \Longrightarrow n \times n \]

Observe there was no choice at any step

Why: At each point the symbols present determined which alternative could apply. A different choice would have produced a different string.

Verify the unique tree has the intended shape

Why: The addition sits at the root and the multiplication in its right subtree, which is the grouping that evaluates the product first. The grammar now encodes the conventional precedence structurally.

\[ n + (n \times n) \quad \text{— the only tree} \ \checkmark \]

49. Verify the layered grammar is unambiguous — line by line

Picture it

Animation

Shows: Each line of the worked example "Verify the layered grammar is unambiguous", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The addition sits at the root and the multiplication in its right subtree, which is the grouping that evaluates the product first. The grammar now encodes the conventional precedence structurally.

50. Guess the shape of the answer: Check left-associativity on a chain

Estimation

Predict first

Confirm the second repair took effect, on a chain of equal operators.

Commit before you compute: what does Check left-associativity on a chain come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the choice matters for a non-associative operator

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Replacing addition by subtraction, the left-leaning grammar gives the conventional reading and the right-leaning one does not.

51. Check left-associativity on a chain

Worked example

Confirm the second repair took effect, on a chain of equal operators.

\[ w = n + n + n \]

Apply the expression rule

Why: Its recursive occurrence is on the left, so the left child is an expression and the right child is a term.

\[ E \Longrightarrow E + T \]

Ask what the right child can span

Why: A term cannot contain a plus sign at its top level, since no term rule produces one. So the right child must be the final number alone.

Conclude the split is forced

Why: The last plus sign must be the one at the root, so the left child spans the first two numbers and their plus. The grouping leans left, and no other split is possible.

\[ (n + n) + n \]

Contrast with the right-recursive variant

Why: Had the rule been written with the recursion on the right, the same argument would force the first plus to the root, giving the opposite grouping.

Verify the choice matters for a non-associative operator

Why: Replacing addition by subtraction, the left-leaning grammar gives the conventional reading and the right-leaning one does not. So the side of the recursion is a real design decision, not a stylistic one.

\[ (5-3)-1 = 1 \;\text{(left)} \qquad 5-(3-1) = 3 \;\text{(right)} \ \checkmark \]

52. Check left-associativity on a chain — line by line

Picture it

Animation

Shows: Each line of the worked example "Check left-associativity on a chain", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Replacing addition by subtraction, the left-leaning grammar gives the conventional reading and the right-leaning one does not. So the side of the recursion is a real design decision, not a stylistic one.

53. The layering recipe

Pattern

Turning an ambiguous operator grammar into an unambiguous one is mechanical once the conventions are chosen.

  1. List the operators from loosest-binding to tightest.
  2. Create one variable per precedence level, plus one for the atoms.
  3. Give each level a recursive alternative for its own operators, and a non-recursive alternative descending to the next level.
  4. Put the recursive occurrence on the left for left-associative operators, on the right for right-associative ones.
  5. Put brackets and literals in the tightest level, with brackets returning to the loosest.

The bracket rule returning to the loosest level is what lets parentheses override precedence — inside them, the whole hierarchy starts again.

54. Something is wrong here: fixing precedence but forgetting associativity

Anomaly

Predict first

A student writes this, and it looks reasonable:

Repair the arithmetic grammar so it is unambiguous.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Split the single variable into an expression level for addition and a term level for multiplication, so the operators no longer compete.

Repair the arithmetic grammar so it is unambiguous.

Why: Split the single variable into an expression level for addition and a term level for multiplication, so the operators no longer compete.

55. Trap: fixing precedence but forgetting associativity

Trap

The trap

Repair the arithmetic grammar so it is unambiguous.

Add a level per operator

Why: Split the single variable into an expression level for addition and a term level for multiplication, so the operators no longer compete.

\[ E \to E + E \;\mid\; T, \qquad T \to T \times T \;\mid\; n \]

Declare the grammar fixed

Why: Mixed expressions now have a unique grouping, since the levels force the multiplication deeper. Precedence is enforced.

\[ n + n \times n \;\longrightarrow\; n + (n \times n) \]

Find the ambiguity that remains

Why: A chain of equal operators still groups either way, because each level's rule is recursive on both sides.

\[ n + n + n \;\longrightarrow\; (n+n)+n \;\text{ or }\; n+(n+n) \]

The fix

Repair the arithmetic grammar so it is unambiguous.

Add a level per operator, and recurse on one side only

Why: Each level keeps its recursive occurrence on the left and descends to the next level on the right. That fixes precedence and associativity together.

\[ E \to E + T \;\mid\; T, \qquad T \to T \times F \;\mid\; F, \qquad F \to n \]

Check the mixed case

Why: The multiplication is forced into the term level and therefore deeper, so precedence holds.

\[ n + n \times n \;\longrightarrow\; n + (n \times n) \]

Check the chain case

Why: A term cannot contain a plus sign, so the rightmost operand of an addition is always a single term. The grouping is forced to lean left, and the grammar is now unambiguous.

\[ n + n + n \;\longrightarrow\; (n+n)+n \;\text{ only} \ \checkmark \]

56. Decode the notation: Trap: fixing precedence but forgetting associativity

Notation

Annotate

From Trap: fixing precedence but forgetting associativity — read this one piece at a time. What is each part doing?

On: \( E \to E + E \;\mid\; T, \qquad T \to T \times T \;\mid\; n \)

  • Split the single variable into an expression level for addition and a term level for multiplication, so the operators no longer compete.
  • Mixed expressions now have a unique grouping, since the levels force the multiplication deeper. Precedence is enforced.
  • A chain of equal operators still groups either way, because each level's rule is recursive on both sides.

57. Check yourself: the repair

Check

Look at which side each recursive occurrence sits on.

Check your understanding

A grammar has the rule E → T ^ E | T, where the caret is exponentiation. What does this enforce?

  • A. Exponentiation groups to the right (correct)
  • B. Exponentiation groups to the left
  • C. Exponentiation has the lowest precedence
  • D. The grammar is ambiguous

Answer: A

Why: The recursive occurrence sits on the right of the operator, so the right operand may itself contain another exponentiation while the left may not. That forces a chain to group rightwards, which matches the usual mathematical convention.

Why B tempts people
Left grouping would require the recursive occurrence on the left, with the descent to the next level on the right. This rule is the mirror image of that.
Why C tempts people
Precedence is set by where the level sits in the hierarchy, not by which side recurses. This rule says nothing about precedence.
Why D tempts people
Exactly one side recurses, so a chain has only one possible grouping. Both-sides recursion is what causes ambiguity, and this rule avoids it.

58. Designing Real Grammars

Section

Section 4

59. What has to happen first: Add a third precedence level

Ranking

Put in order

Put the moves of Add a third precedence level into the order they have to happen.

  1. Decide where the new level goes
  2. Write the new level with right recursion
  3. Rewire the level above it
  4. Check a mixed expression
  5. Verify both conventions hold at once

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. It binds tighter than multiplication, so it sits between the term level and the atoms.

60. Add a third precedence level

Worked example

Extend the layered grammar with exponentiation, which binds tighter than multiplication and groups to the right.

Decide where the new level goes

Why: It binds tighter than multiplication, so it sits between the term level and the atoms.

\[ E \;>\; T \;>\; P \;>\; F \]

Write the new level with right recursion

Why: Exponentiation groups to the right, so the recursive occurrence goes on the right and the descent on the left.

\[ P \to F \;\hat{\;}\; P \;\mid\; F \]

Rewire the level above it

Why: Multiplication now descends to the new level rather than straight to the atoms.

\[ T \to T \times P \;\mid\; P \]

Check a mixed expression

Why: In a string with a product and a power, the power must sit deeper because the term level can only reach it through the new variable.

Verify both conventions hold at once

Why: A chain of powers groups rightwards, because the recursion is on the right; and a power binds tighter than a product, because its level is lower. Testing a string with both confirms the intended structure is the only one derivable.

\[ n \times n \hat{\;} n \hat{\;} n \;\longrightarrow\; n \times \big(n \hat{\;} (n \hat{\;} n)\big) \ \checkmark \]

61. Add a third precedence level — line by line

Picture it

Animation

Shows: Each line of the worked example "Add a third precedence level", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: In a string with a product and a power, the power must sit deeper because the term level can only reach it through the new variable.

62. The hierarchy is a total order on operators

Intuition

Layering only works because precedence is a linear ordering: every pair of operators has a definite winner.

If two operators were declared to have equal precedence but different associativity, no layering could express it — they would have to share a level, and a shared level cannot recurse on two different sides at once.

Real language specifications therefore assign every operator a distinct precedence, or group equal-precedence operators that share an associativity into one level. Anything else is not expressible by this technique.

\[ \text{equal precedence} \;\Rightarrow\; \text{same level} \;\Rightarrow\; \text{same associativity} \]

63. The dangling else

Concept

The most famous ambiguity in programming languages, and it appears in a three-rule grammar.

\[ S \to \mathrm{if}\;E\;S \;\mid\; \mathrm{if}\;E\;S\;\mathrm{else}\;S \;\mid\; s \]

A statement with two conditionals and one else clause can attach the else to either conditional. Both attachments use only these rules, so the grammar permits both.

The consequence is a genuine behavioural difference: the else branch runs under different conditions depending on which conditional it belongs to.

64. Complete the line: Exhibit the dangling-else ambiguity

Fill the middle

Fill in the blanks

From Exhibit the dangling-else ambiguity — finish the line. Write what belongs on the right of the equals sign before you look.

T_1 \neq T_2, \quad \mathrm\mathrm{yield}(T_2) \ \checkmark(T_1) = ___

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. The outer conditional uses the else-less alternative, and its body is a complete conditional with an else.

65. Exhibit the dangling-else ambiguity

Worked example

Find the two trees for the shortest problematic statement.

\[ \mathrm{if}\;E\;\mathrm{if}\;E\;s\;\mathrm{else}\;s \]

Attach the else to the inner conditional

Why: The outer conditional uses the else-less alternative, and its body is a complete conditional with an else.

\[ \mathrm{if}\;E\;\big(\mathrm{if}\;E\;s\;\mathrm{else}\;s\big) \]

Attach the else to the outer conditional

Why: The outer conditional uses the with-else alternative, and its then-branch is an else-less conditional.

\[ \big(\mathrm{if}\;E\;(\mathrm{if}\;E\;s)\;\mathrm{else}\;s\big) \]

Read off the behavioural difference

Why: In the first, the else runs when the outer test succeeds and the inner fails. In the second, it runs when the outer test fails — regardless of the inner test.

Note both are structurally legal

Why: Each tree uses only the three rules and yields the same token sequence, so the grammar genuinely permits both readings.

Verify the ambiguity is not merely theoretical

Why: The two readings disagree on the case where the outer test fails: one runs the else branch and the other runs nothing. A language leaving that unspecified would have programs whose behaviour depends on the parser, so real languages must resolve it.

\[ T_1 \neq T_2, \quad \mathrm{yield}(T_1) = \mathrm{yield}(T_2) \ \checkmark \]

66. Exhibit the dangling-else ambiguity — line by line

Picture it

Animation

Shows: Each line of the worked example "Exhibit the dangling-else ambiguity", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The two readings disagree on the case where the outer test fails: one runs the else branch and the other runs nothing. A language leaving that unspecified would have programs whose behaviour depends on the parser, so real languages must resolve it.

67. Resolving the dangling else

Concept

Three resolutions are used in practice, and they differ in what they change.

ResolutionChangesUsed by
attach to the nearest conditionalthe parser, by a stated ruleC, Java, most languages
split into matched and unmatched statementsthe grammartextbook presentations
require an explicit terminatorthe language itselfAda, and many modern languages

Only the second and third make the grammar unambiguous. The first leaves the grammar ambiguous and resolves the conflict outside it, which is why language specifications must state the rule explicitly.

68. What each one costs: Resolving the dangling else

Trade off

Comparison matrix

From Resolving the dangling else: every row here is a choice with a cost. Fill the Changes column, then say which row you would actually pick and what you give up for it.

ResolutionChangesUsed by
attach to the nearest conditionalthe parser, by a stated ruleC, Java, most languages
split into matched and unmatched statementsthe grammartextbook presentations
require an explicit terminatorthe language itselfAda, and many modern languages

69. State the rule before it runs: Repair the grammar by splitting the…

Hypothesis

Predict first

Repair the grammar by splitting the statement class is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Identify the invariant to enforce

Why: An else must attach to the nearest unmatched conditional. So a statement appearing before an else must itself be fully matched.

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

70. Repair the grammar by splitting the statement class

Worked example

The grammatical fix, which removes the ambiguity rather than deferring it.

Identify the invariant to enforce

Why: An else must attach to the nearest unmatched conditional. So a statement appearing before an else must itself be fully matched.

Split the statement variable in two

Why: One variable for statements whose conditionals all have else clauses, and one for statements that may end with an unmatched conditional.

Write the matched rules

Why: A matched statement is a plain statement, or a conditional both of whose branches are matched.

\[ M \to \mathrm{if}\;E\;M\;\mathrm{else}\;M \;\mid\; s \]

Write the unmatched rules

Why: An unmatched statement is a conditional with no else, or one whose else branch is unmatched — and whose then branch is matched.

\[ U \to \mathrm{if}\;E\;S \;\mid\; \mathrm{if}\;E\;M\;\mathrm{else}\;U, \qquad S \to M \;\mid\; U \]

Verify the problem string now has one tree

Why: The then-branch of a conditional with an else must be matched, and a bare conditional is not matched. So the else cannot attach to the outer conditional, and only the nearest-attachment tree survives — which is the intended reading.

\[ \mathrm{if}\;E\;\big(\mathrm{if}\;E\;s\;\mathrm{else}\;s\big) \quad\text{— the only tree} \ \checkmark \]

71. Repair the grammar by splitting the statement class — line by line

Picture it

Animation

Shows: Each line of the worked example "Repair the grammar by splitting the statement class", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: The then-branch of a conditional with an else must be matched, and a bare conditional is not matched. So the else cannot attach to the outer conditional, and only the nearest-attachment tree survives — which is the intended reading.

72. The repair pattern generalizes

Intuition

Both repairs in this lesson used the same move, and it is worth naming.

When a grammar permits an unwanted grouping, split the variable into two, one for the forms allowed in the constrained position and one for the rest. Then use the constrained variable exactly where the unwanted grouping would have occurred.

The layered arithmetic grammar did this by precedence level; the dangling-else repair did it by matched and unmatched. In both cases the number of variables grew and the number of trees fell to one.

\[ \text{unwanted grouping} \;\Rightarrow\; \text{split the variable, constrain the position} \]

73. Plan first: Design a grammar for nested structures with two bracket…

Step zero

Discussion prompt

Design a grammar for nested structures with two bracket kinds — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Write the shapes a member can have

Answer:

  1. Write the shapes a member can have
  2. Notice this grammar is ambiguous
  3. Remove the ambiguity by forcing a direction
  4. Check the language is unchanged
  5. Verify uniqueness on a string with two groups

74. Design a grammar for nested structures with two bracket kinds

Worked example

A design where ambiguity is easy to introduce accidentally.

\[ L = \text{properly nested strings over } \{\,(,\,),\,[,\,]\,\} \]

Write the shapes a member can have

Why: Empty; a member wrapped in round brackets; a member wrapped in square brackets; or two members side by side.

\[ S \to (S) \;\mid\; [S] \;\mid\; SS \;\mid\; \varepsilon \]

Notice this grammar is ambiguous

Why: The side-by-side alternative together with the empty alternative lets a single member be split into pieces in many ways, including empty pieces.

Remove the ambiguity by forcing a direction

Why: Use a second variable for a single bracketed group, and let the top level be a sequence of them built one at a time from the left.

\[ S \to A\,S \;\mid\; \varepsilon, \qquad A \to (S) \;\mid\; [S] \]

Check the language is unchanged

Why: Every properly nested string is a sequence of bracketed groups, and each group contains a properly nested string. Both grammars generate exactly those.

Verify uniqueness on a string with two groups

Why: The string with a round pair followed by a square pair has one tree: the first group is peeled, then the second, then the sequence ends. The empty alternative can only be used once, at the very end, so no alternative split exists.

\[ ()[] \in L(G), \quad \text{one tree} \ \checkmark \]

75. Design a grammar for nested structures with two… — line by line

Picture it

Animation

Shows: Each line of the worked example "Design a grammar for nested structures with two bracket kinds", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Every properly nested string is a sequence of bracketed groups, and each group contains a properly nested string. Both grammars generate exactly those.

76. Without one step: Designing without introducing ambiguity

Constraint

Discussion prompt

Run Designing without introducing ambiguity with this step confiscated:

Give each precedence level its own variable, descending in one direction only.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Never write a rule recursive on both sides unless you intend ambiguity.
  2. Never combine a sequencing rule with an empty alternative in the same variable.
  3. Give each precedence level its own variable, descending in one direction only.
  4. When a construct may or may not be complete, split it into two variables.
  5. After writing the grammar, test the shortest string with two operators or two groups.

77. Designing without introducing ambiguity

Pattern

Ambiguity is usually introduced by accident. These habits avoid most of it.

  1. Never write a rule recursive on both sides unless you intend ambiguity.
  2. Never combine a sequencing rule with an empty alternative in the same variable.
  3. Give each precedence level its own variable, descending in one direction only.
  4. When a construct may or may not be complete, split it into two variables.
  5. After writing the grammar, test the shortest string with two operators or two groups.

The second bullet is the one that catches people. A variable that can both split and vanish gives every string unboundedly many trees, since empty pieces can be inserted anywhere.

78. Where does it stop working: Designing without introducing ambiguity

Edge cases

Discussion prompt

Designing without introducing ambiguity works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

Ambiguity is usually introduced by accident. These habits avoid most of it.

79. How sure are you: Check yourself: designing

Commit first

Predict first

Why is the grammar S → SS | a | ε especially badly ambiguous?

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: Empty pieces can be inserted anywhere, so each string has unboundedly many trees

Why: The doubling rule can split off a piece that then derives the empty string, and that can be done any number of times at any position. So a single symbol already has infinitely many parse trees, not merely two.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

80. Check yourself: designing

Check

Look for the combination that creates unbounded trees.

Check your understanding

Why is the grammar S → SS | a | ε especially badly ambiguous?

  • A. Empty pieces can be inserted anywhere, so each string has unboundedly many trees (correct)
  • B. It generates no strings at all
  • C. It is not context-free
  • D. The variable S is never eliminated

Answer: A

Why: The doubling rule can split off a piece that then derives the empty string, and that can be done any number of times at any position. So a single symbol already has infinitely many parse trees, not merely two.

Why B tempts people
It generates every run of a's, including the empty string. Generation is not the problem.
Why C tempts people
Every rule has a single variable on its left-hand side, so it is a perfectly ordinary context-free grammar.
Why D tempts people
The empty alternative and the single-symbol alternative both eliminate the variable, so derivations terminate fine.

81. Inherent Ambiguity

Section

Section 5

82. Design a grammar for a list with separators

Worked example

A pattern that appears in every real syntax, and one where the empty case is easy to get wrong.

\[ L = \text{zero or more items separated by commas} \]

Write the wrong version first

Why: The obvious attempt combines a sequencing rule with an empty alternative — the third entry in the ambiguity table.

\[ L \to L\,,\,L \;\mid\; i \;\mid\; \varepsilon \]

See the problem

Why: A single item can be derived directly, or as two pieces with one empty, so the grammar is badly ambiguous.

Separate the empty case from the nonempty one

Why: Use one variable for the possibly-empty list and another for the nonempty list, so the empty alternative can fire only once, at the top.

\[ L \to N \;\mid\; \varepsilon, \qquad N \to i\,,\,N \;\mid\; i \]

Check the separator count

Why: A nonempty list of k items has exactly k minus one commas, because each use of the recursive alternative contributes one comma and one item, and the base contributes one item and no comma.

Verify on the three shortest lists

Why: The empty list uses the empty alternative alone; a single item uses the base of the nonempty variable; two items use one recursive step then the base. Each has exactly one derivation, and no trailing comma is derivable.

\[ \varepsilon,\ i,\ i\,,\,i \in L(G) \qquad i\,, \notin L(G) \ \checkmark \]

83. Design a grammar for a list with separators — line by line

Picture it

Animation

Shows: Each line of the worked example "Design a grammar for a list with separators", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: A nonempty list of k items has exactly k minus one commas, because each use of the recursive alternative contributes one comma and one item, and the base contributes one item and no comma.

84. The empty case belongs at the top

Concept

That repair used a move worth stating on its own, because it recurs in every list-like grammar.

Put the empty alternative on a variable that is used once, at the outermost position, and let a separate variable handle the nonempty recursion. Then emptiness is decided exactly once rather than at every level.

\[ L \to N \;\mid\; \varepsilon \qquad\text{rather than}\qquad N \to \cdots \;\mid\; \varepsilon \]

The same shape handles optional clauses, optional trailing separators, and optional whole sections. Whenever something may be absent, decide its absence at one place only.

85. Some languages cannot be repaired

Concept

Every ambiguity so far was the grammar's fault and could be rewritten away. That is not always possible.

inherently ambiguous language — A context-free language for which every context-free grammar is ambiguous.

The claim quantifies over all grammars, so exhibiting one bad grammar proves nothing. Establishing inherent ambiguity requires ruling out every grammar at once, much as non-regularity did in Lesson 10.

\[ \forall G \; \big(L(G) = L \;\Rightarrow\; G \text{ ambiguous}\big) \]

86. The standard example

Concept

The usual witness is a union of two languages that overlap in an awkward way.

\[ L = \{\, a^{i}b^{j}c^{k} : i = j \;\text{ or }\; j = k \,\} \]

Each half is easy to generate: one grammar matches the a's against the b's and lets the c's run free, and the other matches the b's against the c's. Their union is generated by combining them with a fresh start variable.

The difficulty is the overlap. Strings where both conditions hold belong to both halves, and any grammar must generate them somehow — which turns out to force two trees.

\[ a^{n}b^{n}c^{n} \in \text{both halves} \]

87. What has to be given first: See why the overlap forces two trees

Missing information

Discussion prompt

The full proof is beyond this course, but the obstruction is visible and worth seeing.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

One variable per half, joined by a fresh start variable with two alternatives.

88. See why the overlap forces two trees

Worked example

The full proof is beyond this course, but the obstruction is visible and worth seeing.

Write the obvious grammar

Why: One variable per half, joined by a fresh start variable with two alternatives.

\[ S \to S_1 \;\mid\; S_2 \]

Generate a string in the overlap

Why: A string with equal counts of all three symbols satisfies both conditions, so it can be derived through either alternative.

Observe the two trees

Why: One tree begins by choosing the first alternative and matches the a's against the b's; the other chooses the second and matches the b's against the c's. The trees differ at the root.

\[ T_1 \text{ via } S_1, \qquad T_2 \text{ via } S_2 \]

Try to repair by removing the overlap

Why: The natural fix is to make the second alternative generate only strings not already generated by the first. But that set is not context-free, so no grammar can express the restriction.

Verify why the obvious repairs all fail

Why: Every repair must either drop strings of the overlap, changing the language, or keep both routes to them, keeping the ambiguity. Since no context-free grammar can generate the difference of the two halves, no third option exists — which is the heart of the full proof.

\[ \text{overlap not removable} \;\Rightarrow\; \text{every grammar ambiguous} \ \checkmark \]

89. See why the overlap forces two trees — line by line

Picture it

Animation

Shows: Each line of the worked example "See why the overlap forces two trees", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Every repair must either drop strings of the overlap, changing the language, or keep both routes to them, keeping the ambiguity. Since no context-free grammar can generate the difference of the two halves, no third option exists — which is the heart of the full proof.

90. Why this is rare in practice

Intuition

Inherent ambiguity is a genuine phenomenon but almost never an engineering problem, and it is worth knowing why.

The witness languages are unions of two structurally different descriptions that overlap. Programming languages and data formats are not designed that way — their constructs are made distinguishable on purpose, usually with keywords or delimiters.

So when a real grammar is ambiguous, the overwhelmingly likely cause is a repairable defect: missing precedence, both-sided recursion, or an unconstrained optional part. Reach for the repairs of Section 3 before suspecting the language.

\[ \text{ambiguous grammar in practice} \;\Rightarrow\; \text{almost always repairable} \]

91. Deciding ambiguity is impossible

Concept

A natural next question: given a grammar, can a program tell whether it is ambiguous? The answer is no.

The problem of deciding whether an arbitrary context-free grammar is ambiguous is undecidable. Lesson 29 proves this by reduction from the Post Correspondence Problem.

Deciding whether a language is inherently ambiguous is undecidable too. So both questions in this section are, in general, beyond any algorithm — which is why parser generators report conflicts rather than proving ambiguity.

\[ \text{'is } G \text{ ambiguous?'} \;\text{ is undecidable} \]

92. Teach it back: Deciding ambiguity is impossible

Explain it

Discussion prompt

Explain Deciding ambiguity is impossible to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A natural next question: given a grammar, can a program tell whether it is ambiguous? The answer is no.

93. Guess the shape of the answer: Interpret a parser generator's conflict…

Estimation

Predict first

What a tool can and cannot tell you, given the undecidability just stated.

Commit before you compute: what does Interpret a parser generator's conflict report come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the asymmetry is forced by undecidability

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A tool that answered the ambiguity question in both directions would decide an undecidable problem.

94. Interpret a parser generator's conflict report

Worked example

What a tool can and cannot tell you, given the undecidability just stated.

Understand what the tool checks

Why: A parser generator tries to build a deterministic parser for a restricted class of grammars. It reports a conflict when its construction cannot proceed.

Note what a conflict does not prove

Why: A conflict means the grammar is outside the tool's class. It does not prove the grammar is ambiguous — many unambiguous grammars are outside the class.

Note what the absence of a conflict does prove

Why: If the construction succeeds, the grammar is unambiguous, because a deterministic parser assigns at most one tree per string.

\[ \text{no conflict} \;\Rightarrow\; \text{unambiguous} \]

Read the asymmetry correctly

Why: The tool gives a one-way guarantee, exactly like the pumping lemma did. Success is conclusive; failure is a prompt to investigate.

Verify the asymmetry is forced by undecidability

Why: A tool that answered the ambiguity question in both directions would decide an undecidable problem. So the one-way guarantee is not a limitation of any particular tool but a mathematical necessity.

\[ \text{two-way answer} \;\Rightarrow\; \text{decides the undecidable} \ \checkmark \]

95. Interpret a parser generator's conflict report — line by line

Picture it

Animation

Shows: Each line of the worked example "Interpret a parser generator's conflict report", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: A tool that answered the ambiguity question in both directions would decide an undecidable problem. So the one-way guarantee is not a limitation of any particular tool but a mathematical necessity.

96. Answer it before you see the options: Check yourself: inherent ambiguity

Prediction

Predict first

You write an ambiguous grammar for a language. What have you shown?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Nothing about the language — only that this grammar is ambiguous

Why: Inherent ambiguity quantifies over every grammar for the language, so a single ambiguous grammar is no evidence at all. The three-a example earlier in this lesson had an ambiguous grammar for a language with an easy unambiguous one.

97. Check yourself: inherent ambiguity

Check

Recall what the definition quantifies over.

Check your understanding

You write an ambiguous grammar for a language. What have you shown?

  • A. Nothing about the language — only that this grammar is ambiguous (correct)
  • B. That the language is inherently ambiguous
  • C. That the language is not context-free
  • D. That every grammar for the language is ambiguous

Answer: A

Why: Inherent ambiguity quantifies over every grammar for the language, so a single ambiguous grammar is no evidence at all. The three-a example earlier in this lesson had an ambiguous grammar for a language with an easy unambiguous one.

Why B tempts people
That would require ruling out every possible grammar, which exhibiting one cannot do.
Why C tempts people
You exhibited a context-free grammar generating it, which proves the language is context-free.
Why D tempts people
This restates inherent ambiguity, and it does not follow from one example any more than one slow algorithm proves a problem hard.

98. Diagnosis and Practice

Section

Section 6

99. Rebuild the recipe: Diagnosing an ambiguous grammar

Ranking

Put in order

These are the steps of Diagnosing an ambiguous grammar, scrambled. Put them back in order before the next slide shows you.

  1. A rule recursive on both sides of an operator — fix by recursing on one side.
  2. Two levels of precedence sharing one variable — fix by splitting into levels.
  3. A sequencing rule together with an empty alternative — fix by peeling one item at a time.
  4. An optional part that can attach in two places — fix by splitting into matched and unmatched.
  5. None of the above — consider whether the language itself is the problem.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

100. Diagnosing an ambiguous grammar

Pattern

When a grammar turns out ambiguous, the cause is usually one of four things. Check them in this order.

  1. A rule recursive on both sides of an operator — fix by recursing on one side.
  2. Two levels of precedence sharing one variable — fix by splitting into levels.
  3. A sequencing rule together with an empty alternative — fix by peeling one item at a time.
  4. An optional part that can attach in two places — fix by splitting into matched and unmatched.
  5. None of the above — consider whether the language itself is the problem.

The fifth line is genuinely last. Inherent ambiguity is rare enough that it should be suspected only after the first four have been ruled out.

101. Where this shows up: Grammar Design & Ambiguity

Real world

Discussion prompt

Outside this lesson: where does Grammar Design & Ambiguity actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of Diagnosing an ambiguous grammar is doing the work in it.

Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.

Answer:

Lesson 14 takes up what happens when one string has two structures. It defines ambiguity on parse trees rather than on derivations and explains why the distinction matters, then works the arithmetic grammar and its two separate defects, repairing precedence by layering variables and associativity by one-sided recursion. It covers the dangling-else problem with all three standard resolutions and the matched-unmatched grammatical repair, and the habits that keep you from introducing ambiguity by accident. It then treats inherent ambiguity with the standard overlapping-union witness, the undecidability of the ambiguity question, and how to read a parser generator's conflict report, closing with a diagnostic checklist that maps each symptom to its repair.

102. Plan first: Diagnose and repair an unfamiliar grammar

Step zero

Discussion prompt

Diagnose and repair an unfamiliar grammar — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Test the shortest suspicious string

Answer:

  1. Test the shortest suspicious string
  2. Identify the cause from the checklist
  3. Apply the corresponding repair
  4. Re-test the problem string
  5. Verify the language is unchanged and uniqueness now holds

103. Diagnose and repair an unfamiliar grammar

Worked example

Apply the checklist to a grammar you have not seen.

\[ S \to S\,;\,S \;\mid\; \mathrm{stmt} \;\mid\; \varepsilon \]

Test the shortest suspicious string

Why: A single statement can be derived directly, or by splitting into two pieces one of which is empty. Two trees already.

Identify the cause from the checklist

Why: This is the third item: a sequencing rule combined with an empty alternative. The empty piece can be inserted anywhere, unboundedly often.

Apply the corresponding repair

Why: Peel one item at a time, and let a separate variable handle the empty case so it can be used only once.

\[ S \to \mathrm{stmt}\,;\,S \;\mid\; \mathrm{stmt} \;\mid\; \varepsilon \]

Re-test the problem string

Why: A single statement now matches the second alternative and nothing else, since the first requires a semicolon and the third requires no symbols at all.

Verify the language is unchanged and uniqueness now holds

Why: Both grammars generate any number of statements separated by semicolons, including none. In the repaired grammar each string determines its derivation at every step — the number of semicolons decides which alternative applies — so each has exactly one tree.

\[ L(G_{\text{old}}) = L(G_{\text{new}}), \quad \text{one tree per string} \ \checkmark \]

104. Diagnose and repair an unfamiliar grammar — line by line

Picture it

Animation

Shows: Each line of the worked example "Diagnose and repair an unfamiliar grammar", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Both grammars generate any number of statements separated by semicolons, including none. In the repaired grammar each string determines its derivation at every step — the number of semicolons decides which alternative applies — so each has exactly one tree.

105. What has to happen first: Check a repaired grammar systematically

Ranking

Put in order

Put the moves of Check a repaired grammar systematically into the order they have to happen.

  1. Test the shortest string with two of the same construct
  2. Test the shortest string with two different constructs
  3. Test the empty case
  4. Confirm the language did not change
  5. Verify the four checks are sufficient in practice

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Two operators, two list items, or two nested groups — whichever the grammar builds.

106. Check a repaired grammar systematically

Worked example

After any repair, run this check rather than trusting the fix.

Test the shortest string with two of the same construct

Why: Two operators, two list items, or two nested groups — whichever the grammar builds. If it has one tree, the associativity defect is gone.

Test the shortest string with two different constructs

Why: One of each operator, or a group inside a sequence. If it has one tree, the precedence defect is gone.

Test the empty case

Why: Derive the empty string, if it belongs, and confirm exactly one derivation reaches it. Empty alternatives are the usual source of unbounded ambiguity.

Confirm the language did not change

Why: List the shortest few strings of the old grammar and check each is still derivable, and that no new short string appeared.

Verify the four checks are sufficient in practice

Why: Each check targets one row of the ambiguity table from Section 1, and the fourth guards against a repair that removed ambiguity by removing strings. Together they catch every defect this lesson has covered.

\[ \text{4 checks} \;\longleftrightarrow\; \text{4 causes} \ \checkmark \]

107. Check a repaired grammar systematically — line by line

Picture it

Animation

Shows: Each line of the worked example "Check a repaired grammar systematically", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Each check targets one row of the ambiguity table from Section 1, and the fourth guards against a repair that removed ambiguity by removing strings. Together they catch every defect this lesson has covered.

108. Repairs must preserve the language

Intuition

Every repair in this lesson was checked twice: once for uniqueness and once for the language being unchanged. The second check is the one people skip.

It is easy to remove ambiguity by accidentally removing strings. Forcing one-sided recursion can drop the empty case; splitting a variable can leave one branch unreachable; adding a level can make a construct underivable.

So a repair is only correct when both halves hold: one tree per string, and the same set of strings. Reporting the first without the second is the commonest error in grammar exercises.

\[ \text{repair correct} \iff \text{unambiguous} \;\text{and}\; L \text{ unchanged} \]

109. By analogy: Repairs must preserve the language

Analogy

Discussion prompt

Explain Repairs must preserve the language by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Every repair in this lesson was checked twice: once for uniqueness and once for the language being unchanged. The second check is the one people skip.

110. What ambiguity costs and what it buys

Concept

Ambiguous grammars are not always wrong. It is worth knowing when to tolerate one.

SituationAmbiguity acceptable?
you only need membershipyes — the tree is discarded anyway
you need a unique meaningno — resolve it
you use a general parser that returns all treessometimes — if downstream can choose
you use a deterministic parser generatorno — it will refuse to build

Natural-language grammars are deliberately ambiguous, because natural language is; the parser returns several trees and later stages pick. Programming-language grammars are deliberately not.

111. Fill in: Ambiguity acceptable? for What ambiguity costs and what it buys

Comparison

Comparison matrix

From What ambiguity costs and what it buys: refill the Ambiguity acceptable? column from what you know. The rest of the table is as it appeared.

SituationAmbiguity acceptable?
you only need membershipyes — the tree is discarded anyway
you need a unique meaningno — resolve it
you use a general parser that returns all treessometimes — if downstream can choose
you use a deterministic parser generatorno — it will refuse to build

112. Ambiguity and the chapters ahead

Intuition

The question raised here reappears in a stronger form twice more before the course ends.

So this lesson's practical question turns out to have a deep answer, and the tools to give it are still four chapters away.

113. Break it if you can: Ambiguity and the chapters ahead

Counterexample

Discussion prompt

The question raised here reappears in a stronger form twice more before the course ends.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

114. Rule out three: Check yourself: diagnosis

Elimination

Eliminate the wrong options

A grammar gives the string of three operands joined by two identical operators exactly two parse trees. Which repair applies?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Make the operator's rule recursive on one side only
  • B. Add a precedence level for that operator
  • C. Remove the empty alternative
  • D. Split the variable into matched and unmatched forms

Survives elimination: A

Why: A chain of one operator grouping two ways is the associativity defect, caused by the recursive variable appearing on both sides. Restricting the recursion to one side forces a consistent lean and leaves exactly one tree.

115. Check yourself: diagnosis

Check

Match the symptom to the standard cause.

Check your understanding

A grammar gives the string of three operands joined by two identical operators exactly two parse trees. Which repair applies?

  • A. Make the operator's rule recursive on one side only (correct)
  • B. Add a precedence level for that operator
  • C. Remove the empty alternative
  • D. Split the variable into matched and unmatched forms

Answer: A

Why: A chain of one operator grouping two ways is the associativity defect, caused by the recursive variable appearing on both sides. Restricting the recursion to one side forces a consistent lean and leaves exactly one tree.

Why B tempts people
Precedence levels separate different operators. With only one operator involved there is nothing to separate.
Why C tempts people
An empty alternative causes unboundedly many trees, not exactly two. The count here points elsewhere.
Why D tempts people
Matched and unmatched splitting fixes an optional part attaching in two places, which is the dangling-else shape rather than an operator chain.

116. Connect it up: Grammar Design & Ambiguity

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — What Ambiguity Is · The Arithmetic Example · Removing Ambiguity · Designing Real Grammars · Inherent Ambiguity · Diagnosis and Practice. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

117. What you can do now

Recap

You can tell a defective grammar from a defective language, and repair the first kind.

SituationMove
operators grouping either wayrecurse on one side only
different operators competingone variable per precedence level
unboundedly many treeslook for sequencing plus an empty alternative
an optional part attaching in two placessplit into matched and unmatched
all four repairs failconsider inherent ambiguity — but check again first

Lesson 15 puts grammars into a normal form, which makes the algorithms and proofs of the following lessons uniform.

Sources

  1. Sipser, Introduction to the Theory of Computation, 3rd ed., Ch. 2.1 (Ambiguity) — Cengage, 2013.
  2. Hopcroft, Motwani & Ullman, Introduction to Automata Theory, Languages, and Computation, 3rd ed., Ch. 5.4 (Ambiguity in grammars and languages) — Pearson, 2007.
  3. Aho, Lam, Sethi & Ullman, Compilers: Principles, Techniques, and Tools, 2nd ed., Ch. 4.3 (Writing a grammar; the dangling-else problem) — Pearson, 2006.
  4. Parikh, 'On context-free languages', Journal of the ACM 13(4) (inherently ambiguous languages) — ACM, 1966.
  5. Every grammar, parse tree and repair in this deck was checked by hand against the strings claimed, including the chain and mixed-operator cases. — Verified 2026-08-08.

Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108