Lesson 11 is a worked catalogue of non-regularity proofs. It gives the canonical matched-counts proof and the trap of choosing your own split, then counting languages and the line between bounded and unbounded quantities, and shape languages such as palindromes and repeated blocks, handled with the marker trick. It works arithmetic languages that need gap and factorization arguments, for the perfect squares and for the primes, and shows closure arguments as the shorter alternative. It ends with a language that pumps yet is not regular, together with the distinguishable-prefix fallback, and a diagnostic table for repairing failed attempts.
Subject: Theory of Computation · 120 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Theory of Computation · Lesson 11
A catalogue of non-regularity proofs: the string to choose, the repetition count to pick, and the cases where pumping is the wrong tool entirely.
Objectives
Lesson 10 built the tool. This lesson uses it, repeatedly, until the choices become automatic. By the end you can:
Warm-up
Discussion prompt
Before we open Applications of the Pumping Lemma: without looking back, what was the main idea of The Pumping Lemma for Regular Languages, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Lesson 10 supplies the first tool for proving a language has no finite automaton. Covers why every earlier technique proves only the positive direction, how finiteness plus the pigeonhole principle forces a repeated state and hence a pumpable cycle, the exact statement with its three conditions and the quantifier order, the adversary game that makes the alternation manageable, the proof of the lemma itself with each condition traced to its source, and how to set a proof up: choosing a string built from the pumping length with a uniform prefix, and choosing the repetition count.
Section
Section 1
Concept
Every proof in this lesson is the same five sentences from Lesson 10. Only the string and the repetition count change.
Sentence four is the one that varies most, and it is where a badly chosen string reveals itself: if you cannot say what the middle must be, the string was wrong.
Counterexample
Discussion prompt
Every proof in this lesson is the same five sentences from Lesson 10. Only the string and the repetition count change.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Sentence four is the one that varies most, and it is where a badly chosen string reveals itself: if you cannot say what the middle must be, the string was wrong.
Ranking
Put in order
Put the moves of The matched-counts language into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Suppose the language were regular.
Worked example
The standard first example. Prove that the strings of some number of zeros followed by the same number of ones is not regular.
\[ L = \{\, 0^{n}1^{n} : n \ge 0 \,\} \]
Assume regularity and take the pumping length
Why: Suppose the language were regular. The lemma then provides a pumping length p, whose value we do not know.
Choose the string
Why: Take p zeros followed by p ones. It is in the language, since the counts match, and its length is twice p, which is at least p.
\[ w = 0^{p}1^{p} \in L, \qquad |w| = 2p \ge p \]
Apply the lemma and pin down the middle
Why: The split has a nonempty middle with the first two parts together at most p long. The first p symbols of the string are all zeros, so the middle lies entirely inside that block.
\[ y = 0^{k}, \qquad 1 \le k \le p \]
Pump with two copies
Why: The middle is duplicated, adding k more zeros. The block of ones is untouched, since it sits entirely in the third part.
\[ xy^{2}z = 0^{p+k}1^{p} \]
Verify the result is outside the language, for every legal split
Why: Since k is at least one, the zero count strictly exceeds the one count, so the string does not have the required form. This holds for every value of k the adversary could have chosen, so the contradiction is complete and the language is not regular.
\[ p+k > p \;\Rightarrow\; 0^{p+k}1^{p} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "The matched-counts language", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Since k is at least one, the zero count strictly exceeds the one count, so the string does not have the required form. This holds for every value of k the adversary could have chosen, so the contradiction is complete and the language is not regular.
Intuition
Every requirement from the string-choice pattern is met, and it is worth checking them off explicitly this once.
| Requirement | How the string meets it |
|---|---|
| in the language | the counts match by construction |
| length at least p | its length is twice p |
| uniform first p symbols | the leading block is all zeros |
| pumping breaks the condition | only the zero count changes |
The third row is what makes the fourth possible. Because the middle is forced into the zeros, pumping can only ever change one of the two counts — and a condition demanding they stay equal cannot survive that.
Comparison
Comparison matrix
From Why this string is the right one: refill the How the string meets it column from what you know. The rest of the table is as it appeared.
| Requirement | How the string meets it |
|---|---|
| in the language | the counts match by construction |
| length at least p | its length is twice p |
| uniform first p symbols | the leading block is all zeros |
| pumping breaks the condition | only the zero count changes |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Prove the matched-counts language is not regular.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Take the string with p zeros and p ones, and let the middle part be the single zero at the boundary — the one just before the ones begin.
Prove the matched-counts language is not regular.
Why: Take the string with p zeros and p ones, and let the middle part be the single zero at the boundary — the one just before the ones begin.
Trap
Prove the matched-counts language is not regular.
Pick a convenient split
Why: Take the string with p zeros and p ones, and let the middle part be the single zero at the boundary — the one just before the ones begin.
\[ x = 0^{p-1}, \quad y = 0, \quad z = 1^{p} \]
Pump and declare victory
Why: Duplicating that zero unbalances the counts, so the result is outside the language and the proof is finished.
\[ xy^{2}z = 0^{p+1}1^{p} \notin L \]
Prove the matched-counts language is not regular.
Notice the split is not yours to choose
Why: The lemma promises that some valid split exists — it does not promise the one you picked. A proof must derive its contradiction for every split satisfying the two side conditions.
\[ \exists (x,y,z) \text{, not } \forall (x,y,z) \text{ of your choosing} \]
Argue from the conditions instead
Why: The early-split condition says the first two parts total at most p symbols. Since the first p symbols are all zeros, every legal middle is a nonempty run of zeros — which is all the argument ever needed.
\[ |xy| \le p \;\Rightarrow\; y = 0^{k}, \; 1 \le k \le p \]
Now the conclusion covers every case
Why: Pumping any such middle adds zeros only, so the counts differ for every legal split. The proof is now valid, and it is barely longer.
\[ 0^{p+k}1^{p} \notin L \text{ for all } 1 \le k \le p \ \checkmark \]
Notation
Annotate
From Trap: choosing the split yourself — read this one piece at a time. What is each part doing?
On: \( x = 0^{p-1}, \quad y = 0, \quad z = 1^{p} \)
Ranking
Put in order
These are the steps of Writing the proof so it is checkable, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
A correct proof and a plausible one look similar. These habits keep them apart.
The third bullet is the one graders look for. A proof that never mentions the early-split condition has almost certainly chosen its own split.
Elimination
Eliminate the wrong options
In the proof for the matched-counts language, which condition guarantees the middle part contains only zeros?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The condition bounds the combined length of the first two parts by p, so the middle lies inside the first p symbols of the string. Since those were all chosen to be zeros, the middle can contain nothing else.
Check
Think about which condition forces the middle part into the zeros.
Check your understanding
In the proof for the matched-counts language, which condition guarantees the middle part contains only zeros?
Answer: A
Why: The condition bounds the combined length of the first two parts by p, so the middle lies inside the first p symbols of the string. Since those were all chosen to be zeros, the middle can contain nothing else.
Step zero
Discussion prompt
A variant: more ones than zeros — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Choose the string carefully
Answer:
Worked example
The same shape of argument on a language whose condition is an inequality rather than an equality.
\[ L = \{\, 0^{m}1^{n} : n > m \ge 0 \,\} \]
Choose the string carefully
Why: Take p zeros followed by p plus one ones. The inequality holds by exactly one, which is the tightest margin available.
\[ w = 0^{p}1^{p+1} \in L \]
Pin down the middle
Why: The first p symbols are zeros, so the middle is a nonempty run of zeros.
\[ y = 0^{k}, \qquad 1 \le k \le p \]
Choose the repetition count to close the margin
Why: The margin was one, so adding even a single zero destroys it. Two copies add k zeros, and k is at least one.
\[ xy^{2}z = 0^{p+k}1^{p+1} \]
Check the inequality now fails
Why: The one count is p plus one and the zero count is p plus k, which is at least p plus one. So the ones no longer strictly outnumber the zeros.
\[ p+1 > p+k \text{ is false for } k \ge 1 \]
Verify the choice of margin was necessary
Why: Had the string used p zeros and two-p ones, adding k zeros would leave the ones still ahead whenever k is small, and no contradiction would follow. Choosing the tightest possible margin is what makes a single pump decisive.
\[ \text{tight margin} \;\Rightarrow\; \text{one pump suffices} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "A variant: more ones than zeros", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The one count is p plus one and the zero count is p plus k, which is at least p plus one. So the ones no longer strictly outnumber the zeros.
Intuition
The variant above illustrates a general principle worth carrying into every proof.
When the condition is an inequality, choose the string so the inequality is as close to failing as the language allows. Then the smallest possible change breaks it, and you do not have to reason about how large the middle part might be.
When the condition is an equality, any change breaks it, so the margin question does not arise — which is why matched-count languages are the easiest of all.
\[ \text{equality} : \text{any pump works} \qquad \text{inequality} : \text{make it tight first} \]
Analogy
Discussion prompt
Explain Tight margins make short proofs by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
When the condition is an equality, any change breaks it, so the margin question does not arise — which is why matched-count languages are the easiest of all.
Section
Section 2
Concept
A counting language constrains how many of something a string contains. Whether pumping breaks it depends on what the count is compared against.
| Condition | Regular? | Why |
|---|---|---|
| count is even | yes | only the parity matters — two values |
| count is at most five | yes | capped — six values |
| count is a multiple of seven | yes | remainder — seven values |
| two counts are equal | no | the difference is unbounded |
| one count exceeds another | no | the difference is unbounded |
The line is always the same: a count taken against a fixed bound or modulus is fine, and a count compared against another count is not.
Discrimination
Sort into buckets
Sort these by Regular?, from memory, without looking back at Counting conditions and what breaks them. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Estimation
Predict first
The counts need not be in separate blocks for the language to be irregular.
Commit before you compute: what does Equal zeros and ones, in any order come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify by comparing with the closure route
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Lesson 7 proved this same language irregular by intersecting it with the regular language of zeros-then-ones.
Worked example
The counts need not be in separate blocks for the language to be irregular.
\[ L = \{\, w \in \{0,1\}^{*} : \#_0(w) = \#_1(w) \,\} \]
Choose a string with a uniform prefix
Why: Take p zeros followed by p ones. It has equal counts, so it is in the language, and its prefix is uniform.
\[ w = 0^{p}1^{p} \]
Pin down the middle and pump
Why: The middle is a nonempty run of zeros, and duplicating it adds zeros without adding ones.
\[ xy^{2}z = 0^{p+k}1^{p} \]
Check the counts
Why: The zero count is p plus k and the one count is p. Since k is at least one, they differ.
Note that this also proves a stronger statement
Why: The same string works whether the language demands the zeros come first or allows any order, because the chosen string happens to have them in order. A proof for the general language is therefore no harder.
Verify by comparing with the closure route
Why: Lesson 7 proved this same language irregular by intersecting it with the regular language of zeros-then-ones. Both routes reach the same conclusion; this one is self-contained, and that one is shorter once the matched-counts result is already available.
\[ \#_0 \neq \#_1 \;\Rightarrow\; 0^{p+k}1^{p} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Equal zeros and ones, in any order", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The zero count is p plus k and the one count is p. Since k is at least one, they differ.
Missing information
Discussion prompt
Not every language with two counts is irregular. This one is, but the reason is worth seeing fail first.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Take p zeros followed by twice p ones. The condition holds and the prefix is uniform.
Worked example
Not every language with two counts is irregular. This one is, but the reason is worth seeing fail first.
\[ L = \{\, 0^{n}1^{2n} : n \ge 0 \,\} \]
Choose the string
Why: Take p zeros followed by twice p ones. The condition holds and the prefix is uniform.
\[ w = 0^{p}1^{2p} \]
Pin down the middle
Why: As before, the middle is a nonempty run of zeros from the leading block.
\[ y = 0^{k}, \qquad 1 \le k \le p \]
Pump and check the ratio
Why: Duplicating gives p plus k zeros and still twice p ones. The condition would require the ones to be twice the zeros, so it would need twice p plus twice k ones.
\[ 2(p+k) = 2p + 2k \neq 2p \]
Confirm the mismatch for every legal k
Why: Since k is at least one, the required count exceeds the actual count by at least two. The condition fails regardless of the split.
Verify that the fixed ratio did not save the language
Why: It might seem that a fixed multiplier makes the relationship more machine-friendly, but the quantity that must be remembered is still the zero count itself, which is unbounded. The proof goes through unchanged.
\[ 0^{p+k}1^{2p} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "A count against a fixed multiple", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: It might seem that a fixed multiplier makes the relationship more machine-friendly, but the quantity that must be remembered is still the zero count itself, which is unbounded. The proof goes through unchanged.
Intuition
A machine reading the zeros has to hand something to the part of itself that reads the ones. That handover is the whole difficulty, and multiplying by a constant does not shrink it.
Whether the machine must produce the same number, twice the number, or the number plus three, it must first know the number — and the number is unbounded.
The only conditions that escape are those where a bounded summary suffices: a parity, a remainder, or a capped count. Anything requiring the exact value fails, whatever is done with it afterwards.
\[ \text{exact unbounded value needed} \;\Rightarrow\; \text{not regular} \]
Explain it
Discussion prompt
Explain Why a ratio is no easier than an equality to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A machine reading the zeros has to hand something to the part of itself that reads the ones. That handover is the whole difficulty, and multiplying by a constant does not shrink it.
Worked example
A counting language where the thing being counted is a two-symbol block rather than a symbol.
\[ L = \{\, w : \#_{ab}(w) = \#_{ba}(w) + 1 \,\} \]
Ask whether the quantity is bounded
Why: The difference between the two block counts must be exactly one. The difference itself ranges over all integers as the string grows, so it is unbounded — which suggests irregularity.
Test the suggestion before proving it
Why: Try a few strings by hand. Every occurrence of the first block that is not immediately balanced seems recoverable later, which is a warning sign that the condition may be looser than it looks.
Discover the language is in fact regular
Why: Between any two occurrences of one block there must be an occurrence of the other, because the string has to return. So the difference never leaves the range from minus one to one, and three states suffice.
\[ \#_{ab}(w) - \#_{ba}(w) \in \{-1, 0, 1\} \]
Note what went wrong with the first instinct
Why: The quantity looked unbounded but was constrained by the structure of the string itself. Checking whether a quantity can actually reach large values, rather than whether it could in principle, is a step worth taking.
Verify by exhibiting the machine's behaviour on three strings
Why: The string ab has a difference of one and is accepted; ba has minus one and is rejected; abab has one and is accepted. All three agree with the three-state design, confirming the language is regular after all.
\[ ab,\ abab \in L \qquad ba,\ \varepsilon \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Counting occurrences of a block", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The string ab has a difference of one and is accepted; ba has minus one and is rejected; abab has one and is accepted. All three agree with the three-state design, confirming the language is regular after all.
Intuition
That example is worth dwelling on, because it is the most common way a classification goes wrong in the safe direction.
A quantity is only a problem if the strings of the language can actually drive it to arbitrarily large values. If the structure of the strings keeps it inside a fixed window, the machine can track it after all.
So the classification question is really: over the strings of this language, does the quantity take unboundedly many values? Answering it for arbitrary strings rather than for members of the language is the error.
\[ \text{unbounded over } \Sigma^{*} \;\not\Rightarrow\; \text{unbounded over } L \]
Prediction
Predict first
Which of these languages over the alphabet of zeros and ones is NOT regular?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Strings where the number of zeros equals the number of ones
Why: Deciding this requires remembering the running difference between the two counts, and that difference is unbounded — a prefix with a gap of three and one with a gap of four have different futures, so no finite state set can separate them all.
Check
Three of these are regular and one is not.
Check your understanding
Which of these languages over the alphabet of zeros and ones is NOT regular?
Answer: A
Why: Deciding this requires remembering the running difference between the two counts, and that difference is unbounded — a prefix with a gap of three and one with a gap of four have different futures, so no finite state set can separate them all.
Step zero
Discussion prompt
A counting language that is regular after all — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: State the language
Answer:
Worked example
Practise the classification by proving a suspicious-looking language regular instead.
State the language
Why: The strings over the alphabet of zeros and ones in which the number of zeros and the number of ones differ by at most two.
Ask what must be remembered
Why: The running difference — but only while it stays within the window from minus two to two. Once it leaves that window it can never come back into the language's favour without passing through, so the machine can cap it.
Count the values
Why: Five values inside the window, plus one state meaning 'has drifted too far'. Six states in total, and the drifted state is a dead state.
\[ \text{difference} \in \{-2,-1,0,1,2\} \;\cup\; \{\text{out of range}\} \]
Note why the cap is legitimate
Why: The condition is about the final difference, and a string whose difference has left the window could still return. So the cap is only legitimate here because leaving the window by more than two cannot be undone within the window's own bookkeeping — the drifted state must record the direction as well.
Verify the state count by checking the boundary
Why: Tracking direction as well as magnitude gives states for differences from minus three to three, with the two extremes absorbing — seven states. Testing the strings 000, 0011 and the empty string against this design gives reject, accept and accept, matching the specification.
\[ \varepsilon,\ 0011,\ 001 \in L \qquad 000 \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "A counting language that is regular after all", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Tracking direction as well as magnitude gives states for differences from minus three to three, with the two extremes absorbing — seven states. Testing the strings 000, 0011 and the empty string against this design gives reject, accept and accept, matching the specification.
Intuition
Half the value of these lessons is knowing which tool to reach for, and that decision takes seconds if made deliberately.
Attempting a pumping proof on a regular language wastes an hour and always fails, because the adversary genuinely has a winning strategy. Attempting a machine on an irregular language wastes the same hour.
Section
Section 3
Concept
A shape language constrains how parts of the string relate to each other — symmetry, repetition, or matching — rather than how many symbols it has.
| Language | Condition |
|---|---|
| palindromes | reads the same in both directions |
| a string followed by itself | the two halves are identical |
| a string followed by its reversal | the second half mirrors the first |
| balanced brackets | every opening has a matching closing |
All four are non-regular, and all four fail for the same underlying reason: the machine would have to remember an unbounded prefix in order to check it against what comes later.
Trade off
Comparison matrix
From Conditions about structure rather than count: every row here is a choice with a cost. Fill the Condition column, then say which row you would actually pick and what you give up for it.
| Language | Condition |
|---|---|
| palindromes | reads the same in both directions |
| a string followed by itself | the two halves are identical |
| a string followed by its reversal | the second half mirrors the first |
| balanced brackets | every opening has a matching closing |
Fill the middle
Fill in the blanks
From The palindromes — finish the line. Write what belongs on the right of the equals sign before you look.
xy^a^{p+k}\,b\,a^{p}z = ___
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Take p copies of a, then a single b, then p copies of a.
Worked example
Prove that the strings reading the same in both directions are not regular, over a two-symbol alphabet.
\[ L = \{\, w \in \{a,b\}^{*} : w = w^{R} \,\} \]
Choose a string with a uniform prefix and a marker
Why: Take p copies of a, then a single b, then p copies of a. It reads the same both ways, so it is in the language.
\[ w = a^{p}\,b\,a^{p} \]
Pin down the middle
Why: The first p symbols are all a's, so the middle is a nonempty run of a's from the leading block.
\[ y = a^{k}, \qquad 1 \le k \le p \]
Pump with two copies
Why: The leading block grows and the trailing block does not, since it lies entirely in the third part.
\[ xy^{2}z = a^{p+k}\,b\,a^{p} \]
Check the symmetry
Why: Reversing the pumped string gives p a's, then b, then p plus k a's. That differs from the pumped string, since the two runs have different lengths.
Verify the marker was necessary
Why: Without the b, the string would be a uniform run of a's, and every such run is a palindrome — pumping would produce another palindrome and the proof would fail. The single b is what makes the two blocks distinguishable.
\[ a^{p+k}ba^{p} \neq \big(a^{p+k}ba^{p}\big)^{R} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "The palindromes", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Reversing the pumped string gives p a's, then b, then p plus k a's. That differs from the pumped string, since the two runs have different lengths.
Intuition
The b in that proof did no counting and satisfied no condition. It was there purely to make the two blocks tell-apart-able.
This trick recurs constantly. When a language is about symmetry or repetition, insert a symbol that appears exactly once, so that 'before the marker' and 'after the marker' are unambiguous notions.
Without it, pumping often produces another member of the language and the proof collapses — not because the language is regular, but because the string was too uniform to expose the problem.
\[ \text{uniform string} \;\Rightarrow\; \text{pumping may stay inside } L \]
Sorting
Sort into buckets
These are the pieces of Applications of the Pumping Lemma, out of order. Put each one back under the part of the lesson it belongs to.
Hypothesis
Predict first
A string followed by itself is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Choose a string that resists an easy repair
Why: Take a block of p a's followed by a b, written twice. The whole string is in the language by construction.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Prove that the strings consisting of some block written twice are not regular.
\[ L = \{\, ww : w \in \{a,b\}^{*} \,\} \]
Choose a string that resists an easy repair
Why: Take a block of p a's followed by a b, written twice. The whole string is in the language by construction.
\[ s = a^{p}b\,a^{p}b \]
Pin down the middle
Why: The first p symbols are all a's, so the middle is a nonempty run of a's from the first block.
\[ y = a^{k}, \qquad 1 \le k \le p \]
Pump and inspect the halves
Why: Two copies give p plus k a's, then b, then p a's, then b. The total length is odd in the sense that the two halves can no longer match.
\[ xy^{2}z = a^{p+k}b\,a^{p}b \]
Show no split into two equal halves works
Why: For the string to be in the language it would have to split into two identical halves. The two b's are at positions p plus k plus one and p plus k plus p plus two, so the halves would have to place their b at the same offset — which requires the two a-runs to be equal, and they are not.
Verify the choice of block was necessary
Why: Had the block been a run of a's alone, the pumped string would be a run of a's of some length, and any even-length run of a's is a block written twice. The trailing b in each half is what prevents that escape.
\[ a^{p+k}b\,a^{p}b \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "A string followed by itself", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Had the block been a run of a's alone, the pumped string would be a run of a's of some length, and any even-length run of a's is a block written twice. The trailing b in each half is what prevents that escape.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Prove the palindromes are not regular.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Take a run of p copies of a. It is a palindrome, and its length is at least p.
Prove the palindromes are not regular.
Why: Take a run of p copies of a. It is a palindrome, and its length is at least p.
Trap
Prove the palindromes are not regular.
Choose the simplest long string
Why: Take a run of p copies of a. It is a palindrome, and its length is at least p.
\[ w = a^{p} \]
Pump it
Why: The middle is a run of a's, and duplicating gives a longer run of a's.
\[ xy^{2}z = a^{p+k} \]
Look for the contradiction and fail to find one
Why: A run of a's of any length reads the same in both directions, so the pumped string is still a palindrome. No contradiction appears, and the proof cannot be completed.
\[ a^{p+k} \in L \quad \text{— no contradiction} \]
Prove the palindromes are not regular.
Choose a string whose two ends must stay balanced
Why: Insert a marker between two uniform blocks, so that pumping the first block breaks the balance with the second.
\[ w = a^{p}\,b\,a^{p} \]
Pump it
Why: The middle lies in the leading block, so duplicating lengthens that block and leaves the trailing one alone.
\[ xy^{2}z = a^{p+k}\,b\,a^{p} \]
Find the contradiction
Why: The reversal places the longer run after the marker instead of before it, so the pumped string is not a palindrome. The proof closes.
\[ a^{p+k}ba^{p} \notin L \ \checkmark \]
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Ranking
Put in order
Put the moves of Balanced brackets into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Take p opening brackets followed by p closing brackets.
Worked example
The language behind every parser, and the reason the next chapters exist.
\[ L = \{\, w : \text{every opening bracket has a matching closing one} \,\} \]
Choose a string that sits exactly on the condition
Why: Take p opening brackets followed by p closing brackets. Every opening is matched, so it is in the language.
\[ w = \text{(}^{p}\,\text{)}^{p} \]
Pin down the middle
Why: The first p symbols are all opening brackets, so the middle is a nonempty run of them.
Pump with two copies
Why: The opening count rises and the closing count does not, so some opening bracket is left without a partner.
\[ \text{(}^{p+k}\,\text{)}^{p} \]
State why this language matters
Why: Nested structure — brackets, tags, block delimiters — is the everyday case that finite memory cannot handle. Recognizing it is exactly what the stack machines of Lesson 16 are built for.
Verify the mismatch for every legal split
Why: Since the middle was nonempty, at least one extra opening bracket appears with no closing partner, for every value of k the adversary could choose. So the pumped string is unbalanced and the language is not regular.
\[ k \ge 1 \;\Rightarrow\; \text{(}^{p+k}\text{)}^{p} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Balanced brackets", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Since the middle was nonempty, at least one extra opening bracket appears with no closing partner, for every value of k the adversary could choose. So the pumped string is unbalanced and the language is not regular.
Concept
The bracket example points at the real dividing line between this chapter and the next.
| Structure | Example | Regular? |
|---|---|---|
| flat, bounded lookback | ends in 01 | yes |
| flat, modular counting | even length | yes |
| nested, unbounded depth | balanced brackets | no |
| two independent nestings | equal a's, b's and c's | no, and not context-free either |
Rows three and four both defeat finite automata, but only row three is handled by the stack machines ahead. That distinction is what Lesson 18's pumping lemma is for.
Discrimination
Sort into buckets
Sort these by Regular?, from memory, without looking back at Nested versus flat structure. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Pattern
Shape languages need one more consideration than counting languages, and it is the one the trap illustrates.
Step five is the one skipped most often. For repetition languages especially, the pumped string can sometimes be re-split into halves in a way that satisfies the condition, and the proof must rule that out.
Check
Think about what pumping a uniform run produces.
Check your understanding
Why does the string consisting of p copies of a fail as a choice for proving the palindromes irregular?
Answer: A
Why: A run of a single symbol reads the same in both directions no matter how long it is, so lengthening or shortening it produces another member of the language. The proof needs a string whose pumped versions leave the language.
Section
Section 4
Concept
For counting and shape languages, deleting or duplicating the middle part breaks the condition immediately. Arithmetic languages resist that.
If the condition is that the length satisfies some arithmetic property, then pumping changes the length by a multiple of the middle part's size — and the new length may well satisfy the property again by accident.
\[ |xy^{i}z| = |w| + (i-1)|y| \]
So the repetition count must be chosen to land the length strictly between two values the property allows. That takes a small argument rather than a reflex.
Estimation
Predict first
Prove that the strings whose length is a perfect square are not regular.
Commit before you compute: what does Lengths that are perfect squares come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify no square lies in that gap, and conclude
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The new length sits strictly between two consecutive perfect squares, and there is no square strictly between them.
Worked example
Prove that the strings whose length is a perfect square are not regular.
\[ L = \{\, a^{n^{2}} : n \ge 0 \,\} \]
Choose the string
Why: Take a run of p squared copies of a. Its length is a perfect square, so it is in the language, and p squared is at least p.
\[ w = a^{p^{2}} \]
Pin down the middle
Why: Everything is the same symbol, so the middle is simply a nonempty run of at most p copies.
\[ y = a^{k}, \qquad 1 \le k \le p \]
Pump with two copies and compute the new length
Why: Duplicating adds k symbols, so the length becomes p squared plus k.
\[ |xy^{2}z| = p^{2} + k \]
Trap the new length strictly between consecutive squares
Why: The new length is strictly greater than p squared. It is also at most p squared plus p, which is strictly less than p squared plus twice p plus one — the next square.
\[ p^{2} < p^{2}+k \le p^{2}+p < p^{2}+2p+1 = (p+1)^{2} \]
Verify no square lies in that gap, and conclude
Why: The new length sits strictly between two consecutive perfect squares, and there is no square strictly between them. So the pumped string is outside the language for every legal k, and the contradiction holds.
\[ p^{2} < |xy^{2}z| < (p+1)^{2} \;\Rightarrow\; |xy^{2}z| \text{ is not a square} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Lengths that are perfect squares", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The new length sits strictly between two consecutive perfect squares, and there is no square strictly between them. So the pumped string is outside the language for every legal k, and the contradiction holds.
Intuition
The whole proof turns on one inequality, and it is worth seeing why it is available.
Consecutive squares are p squared and p squared plus twice p plus one, so the gap between them has width twice p plus one. The early-split condition caps the middle part at p symbols, so pumping once can add at most p.
Adding at most p to a square therefore cannot reach the next square, since that would need twice p plus one. The constraint that made the middle part small is exactly the constraint that makes the argument close.
\[ |y| \le p < 2p+1 = \text{gap between consecutive squares} \]
Step zero
Discussion prompt
Lengths that are prime — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Choose a string
Answer:
Worked example
A language where a large repetition count is the only route.
\[ L = \{\, a^{n} : n \text{ is prime} \,\} \]
Choose a string
Why: Let p be the pumping length and take a prime n at least as large as p plus one. Use a run of n copies of a.
\[ w = a^{n}, \qquad n \text{ prime}, \; n \ge p+1 \]
Pin down the middle
Why: The string is uniform, so the middle is a nonempty run of at most p copies. Write its length as k.
Choose the repetition count to force a factorization
Why: Take the count to be n plus one. Then the new length is n plus n times k, which factors.
\[ |xy^{n+1}z| = n + n\,k = n(1+k) \]
Check both factors exceed one
Why: The first factor is n, which is at least two since it is prime. The second is one plus k, which is at least two since k is at least one.
Verify the product is composite, and conclude
Why: A product of two factors each at least two is composite, so the pumped string's length is not prime and the string is outside the language. The contradiction holds for every legal split.
\[ n \ge 2, \; 1+k \ge 2 \;\Rightarrow\; n(1+k) \text{ composite} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Lengths that are prime", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The first factor is n, which is at least two since it is prime. The second is one plus k, which is at least two since k is at least one.
Worked example
A third arithmetic case, with the same gap technique as the squares but a much wider gap.
\[ L = \{\, a^{2^{n}} : n \ge 0 \,\} \]
Choose the string
Why: Take a run whose length is a power of two at least as large as p. Such a power exists because the powers grow without bound.
\[ w = a^{m}, \qquad m = 2^{n} \ge p \]
Pin down the middle
Why: The string is uniform, so the middle is a nonempty run of at most p copies, and p is at most m.
\[ 1 \le k \le p \le m \]
Pump with two copies and bound the new length
Why: The new length is m plus k. It exceeds m, and since k is at most m it is at most twice m.
\[ m < m + k \le 2m \]
Exclude the endpoint
Why: The new length equals twice m only when k equals m, which would require the middle part to be the entire string — possible only if the first part is empty and the middle is everything, and then k equals m is allowed. Handle it by choosing the string longer than p, so that k is at most p which is strictly less than m.
\[ k \le p < m \;\Rightarrow\; m < m+k < 2m \]
Verify no power of two lies strictly between consecutive ones
Why: The next power after m is twice m, and the new length lies strictly between them. So it is not a power of two, the pumped string is outside the language, and the contradiction holds for every legal split.
\[ m < |xy^{2}z| < 2m \;\Rightarrow\; |xy^{2}z| \text{ is not a power of two} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Lengths that are powers of two", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The next power after m is twice m, and the new length lies strictly between them. So it is not a power of two, the pumped string is outside the language, and the contradiction holds for every legal split.
Intuition
The prime proof did something the earlier ones did not: it picked the count to make the resulting length factor.
The general move is to write the pumped length as a formula in the count, then choose the count so the formula visibly fails the property — by factoring, by landing in a gap, or by matching a forbidden residue.
\[ |xy^{i}z| = |x| + i|y| + |z| \]
| Property of the length | Count to choose |
|---|---|
| a perfect square | two — land inside the gap |
| prime | the length itself plus one — force a factorization |
| a power of two | two — land strictly between powers |
| equal to another count | zero or two — any change suffices |
Commit first
Predict first
In the perfect-squares proof, why can pumping once never reach the next square?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: The middle part is at most p long, and the gap between consecutive squares exceeds p
Why: The early-split condition caps the middle part at p symbols, so one duplication adds at most p. The gap from p squared to the next square is twice p plus one, which is strictly larger, so the new length lands strictly inside the gap.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Think about how much pumping once can add.
Check your understanding
In the perfect-squares proof, why can pumping once never reach the next square?
Answer: A
Why: The early-split condition caps the middle part at p symbols, so one duplication adds at most p. The gap from p squared to the next square is twice p plus one, which is strictly larger, so the new length lands strictly inside the gap.
Section
Section 5
Missing information
Discussion prompt
The mirror-image relative of the repeated-block language, and it needs the same marker trick.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Take p a's, then a b, then a b, then p a's. That is the block of p a's followed by a b, written and then mirrored.
Worked example
The mirror-image relative of the repeated-block language, and it needs the same marker trick.
\[ L = \{\, w\,w^{R} : w \in \{a,b\}^{*} \,\} \]
Choose a string sitting exactly on the condition
Why: Take p a's, then a b, then a b, then p a's. That is the block of p a's followed by a b, written and then mirrored.
\[ s = a^{p}b\,b\,a^{p} \]
Pin down the middle
Why: The first p symbols are all a's, so the middle is a nonempty run of a's from the leading block.
Pump and check the mirror condition
Why: Two copies lengthen the leading run of a's while leaving the trailing run alone, so the string no longer mirrors about its centre.
\[ a^{p+k}b\,b\,a^{p} \]
Rule out an alternative reading
Why: For the pumped string to be in the language it would have to split into some block and that block reversed. The two b's sit adjacent at a position no longer central, so no such split exists.
Verify the doubled marker was necessary
Why: With a single b the string would have odd length and could never be a block followed by its reversal at all, so it would not be in the language to begin with. Doubling the marker keeps the length even and places it exactly at the centre, which is what makes the string a legitimate member.
\[ a^{p+k}bba^{p} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "A string followed by its reversal", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: With a single b the string would have odd length and could never be a block followed by its reversal at all, so it would not be in the language to begin with. Doubling the marker keeps the length even and places it exactly at the centre, which is what makes the string a legitimate member.
Concept
Three shape proofs have now used a marker, and each placed it differently. The placement is not arbitrary.
| Language | Marker | Why there |
|---|---|---|
| palindromes | one symbol at the centre | keeps the string symmetric and odd-length |
| a block written twice | one symbol ending each half | makes the half-boundary identifiable |
| a block then its reversal | two adjacent symbols at the centre | keeps the length even and centres the mirror |
In every case the marker must leave the string inside the language while making the two tied parts impossible to confuse. Getting the first half right and the second half wrong is the usual failure.
Comparison
Comparison matrix
From Marker placement is part of the design: refill the Why there column from what you know. The rest of the table is as it appeared.
| Language | Marker | Why there |
|---|---|---|
| palindromes | one symbol at the centre | keeps the string symmetric and odd-length |
| a block written twice | one symbol ending each half | makes the half-boundary identifiable |
| a block then its reversal | two adjacent symbols at the centre | keeps the length even and centres the mirror |
Concept
Once one non-regular language is established, Lesson 7's closure properties multiply it cheaply, and the resulting proofs are often three lines.
The move: assume the candidate is regular, intersect it with a simple regular language to expose a known non-regular one, and take the contradiction.
\[ L \cap R = N, \; R \text{ regular}, \; N \text{ not regular} \;\Rightarrow\; L \text{ not regular} \]
This needs no string choice, no split analysis and no repetition count. When a closure argument is available it is almost always the better proof.
Estimation
Predict first
Take a language where a direct pumping proof is fiddly and a closure argument is immediate.
Commit before you compute: what does Prove a language irregular by closure instead come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the intersection equality in both directions
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Any string in the intersection has equal counts and zeros first, so it has the matched form; and any matched string has equal counts and zeros first, so it is in the intersection.
Worked example
Take a language where a direct pumping proof is fiddly and a closure argument is immediate.
\[ L = \{\, w \in \{0,1\}^{*} : \#_0(w) = \#_1(w) \,\} \]
Assume the candidate is regular
Why: For contradiction, suppose some machine recognizes it.
Choose a regular helper that isolates a known shape
Why: Take the strings with all zeros before all ones. A three-state machine recognizes it, so it is regular.
\[ R = 0^{*}1^{*} \]
Compute the intersection
Why: A string with equal counts whose zeros all precede its ones is exactly a matched-counts string.
\[ L \cap R = \{\, 0^{n}1^{n} : n \ge 0 \,\} \]
Invoke closure and take the contradiction
Why: The regular languages are closed under intersection, so the result would have to be regular. It is not, by Section 1. So the assumption fails.
Verify the intersection equality in both directions
Why: Any string in the intersection has equal counts and zeros first, so it has the matched form; and any matched string has equal counts and zeros first, so it is in the intersection. Getting this equality exactly right is where closure proofs usually go wrong, and here both inclusions are immediate.
\[ L \cap R = \{0^{n}1^{n}\} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove a language irregular by closure instead", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Any string in the intersection has equal counts and zeros first, so it has the matched form; and any matched string has equal counts and zeros first, so it is in the intersection. Getting this equality exactly right is where closure proofs usually go wrong, and here both inclusions are immediate.
Constraint
Discussion prompt
Run Choosing between pumping and closure with this step confiscated:
Is it the first irregular language of its kind? Use pumping — closure has nothing to build on.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
A quick decision procedure, worth running before writing anything.
The third line is the reason Lesson 10 exists at all. Closure arguments are derivative; something has to be proved from scratch first, and the pumping lemma is what does it.
Edge cases
Discussion prompt
Choosing between pumping and closure works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
A quick decision procedure, worth running before writing anything.
Concept
The pumping lemma is a one-way implication, and here is the language that proves the converse genuinely fails.
Consider the strings over a three-symbol alphabet that either begin with the symbol c and are otherwise unconstrained, or contain no c at all and have equal counts of a and b.
\[ L = \{\, c\,x : x \in \Sigma^{*} \,\} \;\cup\; \{\, w \in \{a,b\}^{*} : \#_a(w) = \#_b(w) \,\} \]
Every sufficiently long string of this language satisfies the pumping condition, because the first branch is so permissive that a valid split always exists. Yet the language is not regular, since the second branch requires unbounded counting.
So exhibiting a valid split proves nothing. Only failure to pump is informative, exactly as Lesson 10's trap warned.
Step zero
Discussion prompt
Handle it with distinguishable prefixes instead — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Find an infinite family of candidate prefixes
Answer:
Worked example
When pumping is silent, the Myhill-Nerode argument from Lesson 6 still works, because it is a genuine characterization.
Find an infinite family of candidate prefixes
Why: Take the strings consisting of some number of a's, one family member per count.
\[ u_i = a^{i}, \qquad i = 0, 1, 2, \dots \]
Show any two of them are distinguishable
Why: Take two with different counts. Append the matching number of b's for the first one.
\[ z = b^{i}, \qquad u_i z \in L, \; u_j z \notin L \text{ for } j \neq i \]
Check the separating string does its job
Why: The first extension has equal counts and no c, so it is in the language. The second has unequal counts and no c, so it is not.
Apply the counting argument
Why: Pairwise distinguishable prefixes can never share a state, so a machine would need at least as many states as there are family members — and there are infinitely many.
\[ \text{infinitely many distinguishable prefixes} \;\Rightarrow\; \text{no finite } Q \]
Verify the family really is pairwise distinguishable
Why: For any two distinct counts, the separating string built from the smaller one sends exactly one of the two extensions into the language. So every pair is separated, not merely consecutive pairs, and the conclusion follows.
\[ \forall i \neq j \; \exists z : \text{exactly one of } u_iz, u_jz \in L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Handle it with distinguishable prefixes instead", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The first extension has equal counts and no c, so it is in the language. The second has unequal counts and no c, so it is not.
Intuition
Collected, so the choice is automatic.
| Tool | Use when | Cost |
|---|---|---|
| closure argument | a known irregular language is one intersection away | three lines |
| pumping lemma | the language is a fresh case with a tight condition | five sentences |
| distinguishable prefixes | pumping is silent, or an exact state count is wanted | an infinite family plus separators |
Try them in that order. The last one always works but takes the most writing, so it is the fallback rather than the default.
Prediction
Predict first
You verify that every long string of some language can be pumped. What follows?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Nothing — the lemma only says regular implies pumpable
Why: The lemma is a one-way implication, and its only legitimate use is the contrapositive: failure to pump proves irregularity. Successful pumping is consistent with both regular and irregular languages, and this lesson exhibits an irregular language that pumps.
Check
Recall which implication the pumping lemma states.
Check your understanding
You verify that every long string of some language can be pumped. What follows?
Answer: A
Why: The lemma is a one-way implication, and its only legitimate use is the contrapositive: failure to pump proves irregularity. Successful pumping is consistent with both regular and irregular languages, and this lesson exhibits an irregular language that pumps.
Section
Section 6
Concept
A proof that will not close is usually failing for one of four identifiable reasons.
| Symptom | Cause | Fix |
|---|---|---|
| pumped string is still in the language | string too uniform | add a marker between the tied parts |
| cannot say what the middle contains | prefix not uniform | put p copies of one symbol first |
| the condition still holds after pumping | margin too loose | make the condition as tight as allowed |
| every string seems to work for the adversary | the language may be regular | try to build a machine instead |
The last row is worth taking seriously rather than pushing harder. A genuine machine is the fastest way to find out that no pumping proof exists.
Trade off
Comparison matrix
From Diagnosing a failed attempt: every row here is a choice with a cost. Fill the Cause column, then say which row you would actually pick and what you give up for it.
| Symptom | Cause | Fix |
|---|---|---|
| pumped string is still in the language | string too uniform | add a marker between the tied parts |
| cannot say what the middle contains | prefix not uniform | put p copies of one symbol first |
| the condition still holds after pumping | margin too loose | make the condition as tight as allowed |
| every string seems to work for the adversary | the language may be regular | try to build a machine instead |
Ranking
Put in order
Put the moves of Diagnose and repair a failing attempt into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Take p copies of c. It has zero a's and zero b's, so the counts match and it is in the language.
Worked example
Work through a repair, on the language of strings with equally many a's and b's over a three-symbol alphabet including c.
Try the obvious string
Why: Take p copies of c. It has zero a's and zero b's, so the counts match and it is in the language.
\[ w = c^{p} \]
Watch the attempt fail
Why: The middle is a nonempty run of c's, and pumping produces another run of c's — which still has zero a's and zero b's, so it is still in the language.
\[ c^{p+k} \in L \quad \text{— no contradiction} \]
Diagnose
Why: The string is too uniform in the wrong way: it satisfies the condition trivially, so changing its length cannot break it. The chosen symbols do not participate in the condition at all.
Repair by choosing a string that engages the condition
Why: Take p copies of a followed by p copies of b. The counts match, the prefix is uniform, and the two blocks are tied together by the condition.
\[ w = a^{p}b^{p} \]
Verify the repaired proof closes
Why: The middle is a nonempty run of a's, and pumping adds a's without adding b's, so the counts differ. The repaired string engages the condition and the contradiction follows for every legal split.
\[ a^{p+k}b^{p} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Diagnose and repair a failing attempt", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The middle is a nonempty run of a's, and pumping adds a's without adding b's, so the counts differ. The repaired string engages the condition and the contradiction follows for every legal split.
Intuition
The repair above illustrates the single most useful diagnostic question: does my string actually exercise the property that makes this language hard?
A string satisfying the condition trivially — zero of everything, or a uniform run — cannot be broken by pumping, because there is nothing to break. The string must sit at a point where the condition is delicately balanced.
For counting languages that means equal nonzero counts; for shape languages it means two blocks tied across a marker; for arithmetic languages it means a length sitting exactly on the property.
\[ \text{trivially satisfied} \;\Rightarrow\; \text{pumping cannot break it} \]
Elimination
Eliminate the wrong options
Your pumped string keeps landing back inside the language. What is the most likely cause?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A string that meets the condition trivially — a uniform run, or one with zero of the relevant symbols — has no delicate balance to disturb. The repair is to choose a string sitting exactly on the condition, with the tied parts separated so pumping breaks the tie.
Check
Match the symptom to its cause.
Check your understanding
Your pumped string keeps landing back inside the language. What is the most likely cause?
Answer: A
Why: A string that meets the condition trivially — a uniform run, or one with zero of the relevant symbols — has no delicate balance to disturb. The repair is to choose a string sitting exactly on the condition, with the tied parts separated so pumping breaks the tie.
Intuition
A proof that will not close is information, but it is weaker information than it feels.
It does not show the language is regular. The lemma is one-way, so a failure to find a breaking string is consistent with an irregular language that pumps, or simply with a poor choice of string.
What it should trigger is the diagnostic table: try a marker, tighten the margin, engage the condition. If three genuinely different strings all fail, the useful next move is to spend ten minutes trying to build a machine — succeeding settles the question, and failing usually reveals the unbounded quantity you need.
\[ \text{proof fails} \;\not\Rightarrow\; \text{language regular} \]
Explain it
Discussion prompt
Explain What a failed proof does and does not tell you to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A proof that will not close is information, but it is weaker information than it feels.
Concept
The standard examples, with the choices that make each proof work. Worth memorizing the middle column.
| Language | String to choose | Count |
|---|---|---|
| matched counts | p zeros then p ones | zero or two |
| equal counts, any order | p zeros then p ones | zero or two |
| more ones than zeros | p zeros then p plus one ones | two |
| palindromes | p a's, one b, p a's | two |
| a block written twice | p a's, b, p a's, b | two |
| length a perfect square | p squared copies | two, with a gap argument |
| length prime | a prime run at least p plus one | the length plus one |
Comparison
Comparison matrix
From The catalogue: refill the String to choose column from what you know. The rest of the table is as it appeared.
| Language | String to choose | Count |
|---|---|---|
| matched counts | p zeros then p ones | zero or two |
| equal counts, any order | p zeros then p ones | zero or two |
| more ones than zeros | p zeros then p plus one ones | two |
| palindromes | p a's, one b, p a's | two |
| a block written twice | p a's, b, p a's, b | two |
| length a perfect square | p squared copies | two, with a gap argument |
| length prime | a prime run at least p plus one | the length plus one |
Ranking
Put in order
These are the steps of The complete procedure, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Everything from Lessons 10 and 11, as one routine.
If step three cannot produce a string that engages the condition, the language may pump despite being irregular — fall back to distinguishable prefixes.
Real world
Discussion prompt
Outside this lesson: where does Applications of the Pumping Lemma actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The complete procedure is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
Lesson 11 is a worked catalogue of non-regularity proofs. It gives the canonical matched-counts proof and the trap of choosing your own split, then counting languages and the line between bounded and unbounded quantities, and shape languages such as palindromes and repeated blocks, handled with the marker trick. It works arithmetic languages that need gap and factorization arguments, for the perfect squares and for the primes, and shows closure arguments as the shorter alternative. It ends with a language that pumps yet is not regular, together with the distinguishable-prefix fallback, and a diagnostic table for repairing failed attempts.
Intuition
Seven languages, one argument. It is worth naming what they all have in common.
In every case the machine must carry a quantity from an early part of the string to a later part — a count, a block, a length. The early-split condition forces the cycle into that early part, and pumping corrupts what was being carried.
So the proofs are not seven tricks but one: put the thing that must be remembered inside the first p symbols, and let the lemma damage it. Choosing the string well is entirely a matter of arranging that.
\[ \text{what must be carried} \;\subseteq\; \text{first } p \text{ symbols} \]
Analogy
Discussion prompt
Explain Why the same argument keeps working by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Seven languages, one argument. It is worth naming what they all have in common.
Concept
With this lesson the regular languages are fully mapped: four ways to show membership, three ways to show non-membership.
The natural next question is what a machine would need in order to recognize the matched-counts language. The answer is an unbounded memory with a restricted discipline — a stack — and that is the model of Lessons 13 to 19.
Those machines recognize the matched-counts language easily, and they have their own pumping lemma with its own catalogue of languages beyond reach. The pattern of this lesson repeats one level up.
\[ \text{regular} \;\subsetneq\; \text{context-free} \;\subsetneq\; \cdots \]
Counterexample
Discussion prompt
With this lesson the regular languages are fully mapped: four ways to show membership, three ways to show non-membership.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The Canonical Proof · Counting Languages · Shape Languages · Arithmetic Languages · When Pumping Is the Wrong Tool · Diagnosis and Catalogue. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can prove languages irregular fluently, and you know which of three tools to reach for.
| Situation | Move |
|---|---|
| two counts must match | p of each, pump by zero or two |
| the string must be symmetric | uniform block, marker, uniform block |
| the length must satisfy arithmetic | choose the count to land in a gap or force a factor |
| a known irregular language is nearby | intersect with a simple pattern |
| pumping succeeds but the language feels irregular | distinguishable prefixes |
Lesson 12 looks at some surprising positive results about regular languages before the course moves on to grammars in Lesson 13.
Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.