Lesson 10 supplies the first tool for proving that a language has no finite automaton. It explains why every earlier technique proves only the positive direction, then shows how finiteness together with the pigeonhole principle forces a repeated state and hence a pumpable cycle. It gives the exact statement with its three conditions and its quantifier order, the adversary game that makes the alternation manageable, and the proof of the lemma itself with each condition traced back to its source. It then shows how to set a proof up: choosing a string built from the pumping length with a uniform prefix, and choosing the repetition count. It includes the one-way-implication trap and the fixed-string trap.
Subject: Theory of Computation · 110 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Theory of Computation · Lesson 10
Every regular language has a property forced on it by finiteness. Find a language lacking it, and you have proved no machine exists.
Objectives
Lessons 4 to 9 built four ways to show a language is regular. None of them can show one is not. This lesson supplies the missing tool. By the end you can:
Warm-up
Discussion prompt
Before we open The Pumping Lemma for Regular Languages: without looking back, what was the main idea of DFA to Regular Expression, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Lesson 9 completes Kleene's theorem by converting any finite automaton back into a regular expression. It explains why this direction cannot be done by structural recursion, then introduces generalized automata whose arrows carry expressions, the four conditions of normal form, and the state-elimination rip-out rule, including the starred self-loop term people drop. It works conversions, among them machines with dead states, shows how the elimination order controls the size of the result, and presents Kleene's original build-up recurrence and its kinship with Floyd-Warshall. It closes with the four now-equivalent definitions of a regular language.
Section
Section 1
Concept
To show a language is regular you exhibit an object: a machine, or an expression. That is a single witness, and finding one settles the question.
To show a language is not regular you would have to rule out every machine and every expression. There are infinitely many of each, so no amount of trying settles anything.
\[ \text{regular} : \exists M \qquad \text{not regular} : \lnot \exists M \equiv \forall M \lnot(\cdots) \]
A universal claim needs a different kind of argument: find a property that every regular language must have, then show the candidate lacks it.
Counterexample
Discussion prompt
To show a language is regular you exhibit an object: a machine, or an expression. That is a single witness, and finding one settles the question.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
To show a language is not regular you would have to rule out every machine and every expression. There are infinitely many of each, so no amount of trying settles anything.
Intuition
This is the standard way to prove something impossible, and it is worth recognizing in the abstract before meeting the details.
The same shape proves an integer is not a perfect square by looking at its last digit, or that a graph is not planar by counting edges. The work is always in step one: finding a property that is genuinely forced and genuinely checkable.
Concept
What is forced on a regular language? Only one thing is available: the machine has finitely many states, and the input can be arbitrarily long.
Read a string longer than the number of states and the machine must revisit some state. There is nowhere else to go — this is the pigeonhole principle from Lesson 2, applied to a run.
\[ |w| \ge |Q| \;\Longrightarrow\; \text{some state repeats during the run on } w \]
A repeated state means the run contains a cycle. And a cycle can be gone round again, or skipped — producing different strings the machine treats identically.
Analogy
Discussion prompt
Explain The property comes from finiteness by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
What is forced on a regular language? Only one thing is available: the machine has finitely many states, and the input can be arbitrarily long.
Picture it
Figure (svg): Automaton with states q0, q1, q2
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Take a three-state machine and feed it a string of length four. Track the states.
Worked example
Take a three-state machine and feed it a string of length four. Track the states.
Figure (svg): Automaton with states q0, q1, q2
Count the states visited
Why: A run on a string of length four visits five states in total: one before reading anything, and one after each of the four symbols.
\[ |w| = 4 \;\Rightarrow\; 5 \text{ states visited} \]
Apply the pigeonhole principle
Why: Five visits into three available states means at least two visits land on the same state. There is no way to avoid it.
\[ 5 > 3 \;\Rightarrow\; \text{some state visited twice} \]
Find the repeat concretely
Why: Run 0010. The states visited are the start, the middle, the start again, the middle, then the accepting state. The start state is visited at positions zero and two.
\[ q_0 \xrightarrow{0} q_1 \xrightarrow{0} q_0 \xrightarrow{1} q_1 \xrightarrow{0} q_0 \]
Read the cycle off the run
Why: Between the two visits to the repeated state, the machine consumed the substring 00 and returned to where it started. That is a cycle in the diagram.
Verify the cycle can be repeated or skipped
Why: Removing that substring gives 10 and inserting a second copy gives 000010; both runs pass through the same states at the same points and end where the original ended. So the machine cannot tell the three strings apart, which is exactly the leverage the lemma will use.
\[ 10, \; 0010, \; 000010 \text{ all end in the same state} \ \checkmark \]
Notation
Annotate
From Watch a long string force a repeat — read this one piece at a time. What is each part doing?
On: \( |w| = 4 \;\Rightarrow\; 5 \text{ states visited} \)
Intuition
The consequence is worth stating in plain terms, because it is the whole content of the lemma.
If a machine goes round a cycle while reading part of a string, then it treats one trip, two trips and no trips as indistinguishable. It has lost track of how many times it went round.
So any language that requires an exact count over an unbounded range is in trouble: the machine will be forced into a cycle before the count is finished, and from then on it cannot tell the counts apart.
\[ \text{unbounded exact count} \;+\; \text{finite states} \;=\; \text{contradiction} \]
Explain it
Discussion prompt
Explain A cycle means the machine has stopped counting to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The consequence is worth stating in plain terms, because it is the whole content of the lemma.
Concept
pumping — Removing or duplicating the portion of a string consumed while the machine traverses a cycle, producing new strings the machine treats the same way.
Cut the string into three pieces: what came before the cycle, what the cycle consumed, and what came after.
\[ w = xyz, \qquad y \text{ is consumed on the cycle} \]
Then every string with the middle piece repeated any number of times — including zero — drives the machine to exactly the same final state, so all of them get the same verdict.
\[ \hat{\delta}(q_0, xy^{i}z) = \hat{\delta}(q_0, xyz) \quad \text{for every } i \ge 0 \]
Ranking
Put in order
These are the steps of How to tell whether a language needs this tool, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Before writing any proof, spend thirty seconds classifying the language. The answer usually decides the whole approach.
Step three saves the most time. A surprising number of languages that look like they need counting need only a parity or a remainder, and a five-minute machine settles them.
Elimination
Eliminate the wrong options
Why must a run on a sufficiently long string revisit a state?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A run on a string of length n visits n plus one states. Once n plus one exceeds the number of states available, the pigeonhole principle forces two of those visits onto the same state, and the portion between them is a cycle.
Check
Think about which finite thing forces the repeat.
Check your understanding
Why must a run on a sufficiently long string revisit a state?
Answer: A
Why: A run on a string of length n visits n plus one states. Once n plus one exceeds the number of states available, the pigeonhole principle forces two of those visits onto the same state, and the portion between them is a cycle.
Concept
It is worth separating the two questions cleanly, because they need opposite kinds of evidence.
| Question | Evidence needed | Difficulty |
|---|---|---|
| is this language regular? | one machine or expression | usually easy |
| is this language not regular? | an argument covering all machines | needs a theorem |
| is this the smallest machine? | a lower bound on states | distinguishable prefixes |
The middle row is what this lesson is for. Notice the third row is closely related — it also rules out machines, but only ones below a certain size. Push that argument to infinity and it becomes a non-regularity proof, which is the Myhill-Nerode route.
Comparison
Comparison matrix
From Two questions that look alike and are not: refill the Evidence needed column from what you know. The rest of the table is as it appeared.
| Question | Evidence needed | Difficulty |
|---|---|---|
| is this language regular? | one machine or expression | usually easy |
| is this language not regular? | an argument covering all machines | needs a theorem |
| is this the smallest machine? | a lower bound on states | distinguishable prefixes |
Ranking
Put in order
Put the moves of Show a language is regular, for contrast into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The strings over the alphabet of zeros and ones whose number of zeros is even.
Worked example
Do the easy direction once, so the contrast with what follows is sharp.
Take a language that looks like it needs counting
Why: The strings over the alphabet of zeros and ones whose number of zeros is even. It mentions a count, which suggests trouble.
Ask whether the count needs to be unbounded
Why: Only the parity matters, and parity has two values. So the machine needs to distinguish two situations, not infinitely many.
\[ \#_0(w) \bmod 2 \in \{0,1\} \]
Exhibit the machine
Why: Two states, one per parity, with the zero-arrows crossing between them and the one-arrows looping. The even state is accepting.
Note that exhibiting it is the entire proof
Why: One object settles the question. No argument about other machines is needed, because regularity only asserts that some machine exists.
Verify the machine on three strings
Why: The empty string has zero zeros, an even count, and is accepted since the start state is the even one. The string 0 is rejected, and 00 is accepted. All three match the specification, so the language is regular.
\[ \varepsilon,\ 00,\ 0110 \in L \qquad 0,\ 010 \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Show a language is regular, for contrast", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The empty string has zero zeros, an even count, and is accepted since the start state is the even one. The string 0 is rejected, and 00 is accepted. All three match the specification, so the language is regular.
Intuition
It is worth being precise about which languages are in danger, because the intuition transfers directly to the choices in Lesson 11.
A machine can track any quantity that takes boundedly many values: a parity, a remainder, a position within a fixed-length window, a count capped at some maximum.
It cannot track a quantity whose range grows with the input: a running difference between two counts, the length of a prefix to be matched later, or the value of an unbounded counter.
| Quantity | Trackable? |
|---|---|
| count modulo 7 | yes — seven values |
| count, capped at 3 | yes — four values |
| count of zeros minus count of ones | no — unbounded |
| the whole prefix read so far | no — unbounded |
Discrimination
Sort into buckets
Sort these by Trackable?, from memory, without looking back at Exactly where finiteness bites. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Pattern
Before the formal statement, here is the whole argument in outline. Everything that follows is detail.
Step three is where the skill is. The string must be long enough, must be in the language, and must be chosen so that every possible cycle position leads to a contradiction.
Section
Section 2
Concept
Here is the theorem in full. Read the quantifiers slowly — their order is the entire content.
pumping lemma — If a language is regular, then there is a length p such that every string in the language of length at least p can be split into three parts, with the middle part nonempty and starting within the first p symbols, so that repeating the middle part any number of times keeps the string in the language.
\[ L \text{ regular} \;\Longrightarrow\; \exists p \; \forall w \in L, |w| \ge p \; \exists x,y,z \]
The three conditions on the split come next, and each one does real work in the proofs.
Concept
Given the split into three parts, the lemma guarantees all three of these.
\[ w = xyz \]
| Condition | Written | Why it holds |
|---|---|---|
| pumping works | every repetition stays in the language | the cycle can be run any number of times |
| the middle is nonempty | the middle part has positive length | a cycle consumes at least one symbol |
| the split is early | the first two parts together are at most p long | the repeat happens within the first p states visited |
\[ xy^{i}z \in L \;\forall i \ge 0, \qquad |y| > 0, \qquad |xy| \le p \]
Trade off
Comparison matrix
From The three conditions: every row here is a choice with a cost. Fill the Why it holds column, then say which row you would actually pick and what you give up for it.
| Condition | Written | Why it holds |
|---|---|---|
| pumping works | every repetition stays in the language | the cycle can be run any number of times |
| the middle is nonempty | the middle part has positive length | a cycle consumes at least one symbol |
| the split is early | the first two parts together are at most p long | the repeat happens within the first p states visited |
Intuition
The two side conditions look like technicalities. They are the reason most proofs work at all.
Nonempty middle stops the adversary handing you an empty cycle, which could be pumped forever without changing the string. Without it the lemma would be vacuous.
Early split confines the cycle to the first p symbols. That is what lets you design a string whose first p symbols are all the same, so you know exactly what the middle part must consist of.
\[ |xy| \le p \;\Rightarrow\; y \text{ lies inside the first } p \text{ symbols} \]
Almost every proof in Lesson 11 chooses its string precisely to exploit the early-split condition.
Concept
Pumping is usually pictured as repeating, but the index may be zero, and that case is often the easiest contradiction.
\[ i = 0 \;\Rightarrow\; xy^{0}z = xz \]
Taking zero copies deletes the middle part entirely, which corresponds to skipping the cycle. The resulting string is shorter than the original and must still be in the language.
For counting languages, deleting is often more obviously fatal than repeating, so it is worth trying first.
Step zero
Discussion prompt
Read the statement carefully on a regular language — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Take a regular language and a machine for it
Answer:
Worked example
The lemma says something true about every regular language. Check it on one, to see the pieces before using them against a language.
Take a regular language and a machine for it
Why: Use the strings over the alphabet of zeros and ones ending in 1, recognized by the two-state machine of Lesson 4.
Find a pumping length that works
Why: The number of states is two, so any string of length two or more forces a repeat. Take the pumping length to be two.
\[ p = 2 \]
Take a long string of the language and split it
Why: Take 001. The first two symbols are both 0, and the machine loops on the start state while reading them. So the cycle is available inside that prefix.
\[ x = \varepsilon, \quad y = 0, \quad z = 01 \]
Check the three conditions
Why: The middle part is nonempty, the first two parts have combined length one which is at most two, and the split reassembles the original string.
Verify pumping keeps every result in the language
Why: Taking zero copies gives 01, one copy gives 001, and three copies give 00001 — every one of them ends in 1 and is therefore in the language. The lemma's promise holds, as it must for a regular language.
\[ 01,\ 001,\ 0001,\ 00001 \in L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Read the statement carefully on a regular language", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The middle part is nonempty, the first two parts have combined length one which is at most two, and the split reassembles the original string.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A language satisfies the pumping condition. What may be concluded?
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The lemma links regularity and pumpability, so a language that pumps must be regular and one that does not must not be.
A language satisfies the pumping condition. What may be concluded?
Why: The lemma links regularity and pumpability, so a language that pumps must be regular and one that does not must not be.
Trap
A language satisfies the pumping condition. What may be concluded?
Treat the lemma as a characterization
Why: The lemma links regularity and pumpability, so a language that pumps must be regular and one that does not must not be.
\[ L \text{ pumps} \overset{?}{\iff} L \text{ regular} \]
Conclude regularity from pumping
Why: Since the candidate pumps, declare it regular and stop looking for a machine.
A language satisfies the pumping condition. What may be concluded?
Read the arrow's direction
Why: The lemma states one implication only: regular implies pumpable. It says nothing about languages that happen to pump.
\[ L \text{ regular} \;\Longrightarrow\; L \text{ pumps} \]
Use only the contrapositive
Why: The single legitimate use is the contrapositive: a language that fails to pump is not regular. Nothing follows from a language that does pump.
\[ L \text{ does not pump} \;\Longrightarrow\; L \text{ not regular} \ \checkmark \]
Remember there are non-regular languages that pump
Why: Such languages exist, so the converse is genuinely false rather than merely unproved. The Myhill-Nerode condition from Lesson 6 is a true characterization; the pumping lemma is not.
Translation
\( L \text{ regular} \;\Longrightarrow\; L \text{ pumps} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Concept
Three misreadings account for nearly every incorrect proof. All three come from over-reading the statement.
The two things you do choose are the string, and the number of repetitions. Everything else is chosen against you, which is exactly what the next section formalizes.
Prediction
Predict first
In the pumping lemma, who or what determines the split of the string into three parts?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: It is given by the lemma; the proof must handle every valid split
Why: The split is existentially quantified inside the statement, so the lemma merely promises that some valid split exists. A proof by contradiction must therefore derive its contradiction for every split satisfying the two side conditions, not for one hand-picked split.
Check
Look at which quantifier governs which variable.
Check your understanding
In the pumping lemma, who or what determines the split of the string into three parts?
Answer: A
Why: The split is existentially quantified inside the statement, so the lemma merely promises that some valid split exists. A proof by contradiction must therefore derive its contradiction for every split satisfying the two side conditions, not for one hand-picked split.
Section
Section 3
Intuition
Lesson 6 already had a way to rule out machines, by counting pairwise distinguishable prefixes. It is worth knowing how the two relate.
| Pumping lemma | Distinguishable prefixes | |
|---|---|---|
| proves | not regular | not regular, and gives the exact state count |
| form | one implication | a genuine characterization |
| effort | usually five lines | needs an infinite family and a separating string for each pair |
| fails when | the language pumps but is irregular | never |
So the prefix argument is strictly stronger and strictly more work. The usual practice is to reach for pumping first and fall back to prefixes only when pumping cannot close the case.
Trade off
Comparison matrix
From How the lemma compares with distinguishable prefixes: every row here is a choice with a cost. Fill the Pumping lemma column, then say which row you would actually pick and what you give up for it.
| Pumping lemma | Distinguishable prefixes | |
|---|---|---|
| proves | not regular | not regular, and gives the exact state count |
| form | one implication | a genuine characterization |
| effort | usually five lines | needs an infinite family and a separating string for each pair |
| fails when | the language pumps but is irregular | never |
Concept
The alternating quantifiers are hard to keep straight in symbols and easy to keep straight as a game between two players.
You are trying to prove the language is not regular. An adversary is defending the claim that it is. The quantifiers say exactly who moves when.
| Move | Who plays it | Quantifier |
|---|---|---|
| the pumping length | adversary | there exists p |
| the string | you | for all w |
| the split into three parts | adversary | there exists a split |
| the number of repetitions | you | for all i |
You win if the pumped string falls outside the language. If you can always win, no matter how the adversary plays, the language is not regular.
Discrimination
Sort into buckets
Sort these by Who plays it, from memory, without looking back at Turn the quantifiers into a game. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Intuition
The rule is mechanical once you see it: whoever the quantifier belongs to gets to move, and existentials belong to whoever is defending.
The adversary is defending regularity, so they supply the witnesses the lemma promises — the pumping length and the split. You are attacking, so you supply the counterexamples — the string and the repetition count.
This also tells you the order of play. The adversary picks the pumping length first, so your string may depend on it. You pick the string before the split, so your string cannot depend on where the split falls.
\[ p \;\to\; w \;\to\; (x,y,z) \;\to\; i \]
Sorting
Sort into buckets
These are the pieces of The Pumping Lemma for Regular Languages, out of order. Put each one back under the part of the lesson it belongs to.
Concept
Getting the information flow right is what stops the two classic errors.
So a proof reads: let the adversary give any p; here is my string, built from p; whatever split they choose subject to the two conditions, here is the repetition count that kills it.
Step zero
Discussion prompt
Play the game on a language that is regular — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: The adversary names a pumping length
Answer:
Worked example
Play against a regular language first, to feel what losing looks like — this is the case where the adversary can always survive.
The adversary names a pumping length
Why: Take the strings over the alphabet of a and b containing at least one a. The adversary names a pumping length of one, since its machine has two states.
You choose a long string in the language
Why: Any string of length at least one containing an a will do. Take a single a.
The adversary splits it
Why: With the string being one symbol, the only split with a nonempty middle and an early cut has empty outer parts and the a in the middle.
\[ x = \varepsilon, \; y = a, \; z = \varepsilon \]
You choose a repetition count and fail
Why: Zero copies gives the empty string, which has no a and is outside the language — but wait: the empty string does fail. Try instead the string ba, where the adversary can split so the middle is the b.
Verify the adversary survives with a better string choice
Why: The lemma promises only that some split works. For ba the adversary offers the middle part b, and pumping it gives bba, bbba and a — all containing an a, so all in the language. The adversary survives, correctly, because the language really is regular.
\[ a,\ ba,\ bba,\ bbba \in L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Play the game on a language that is regular", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The lemma promises only that some split works. For ba the adversary offers the middle part b, and pumping it gives bba, bbba and a — all containing an a, so all in the language. The adversary survives, correctly, because the language really is regular.
Intuition
The previous example is not luck. Against a genuinely regular language, the adversary has a guaranteed strategy.
They simply take the machine, run it on your string, find the state that repeats within the first p steps, and offer the cycle as the middle part. The lemma's proof guarantees this always exists and always pumps.
So the game is a faithful reformulation: you can win exactly when no machine exists. Losing is not a failure of ingenuity — it is evidence the language is regular.
Hypothesis
Predict first
Play the game on a language that is not regular is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: The adversary names a pumping length
Why: They name some p. You do not know its value, so your string must work for every possible value.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Now the same game against the standard example, to see a win. Take the strings of some number of zeros followed by the same number of ones.
The adversary names a pumping length
Why: They name some p. You do not know its value, so your string must work for every possible value.
You choose a string built from p
Why: Take p zeros followed by p ones. It is in the language, and its length is at least p, so it is a legal move.
\[ w = 0^{p}1^{p} \]
The adversary splits it, constrained by the early-split condition
Why: The first two parts together are at most p symbols long, and the first p symbols of your string are all zeros. So the middle part consists entirely of zeros, and it is nonempty.
\[ |xy| \le p \;\Rightarrow\; y = 0^{k} \text{ with } k > 0 \]
You choose a repetition count
Why: Take two copies. The result has more zeros than it had, and exactly as many ones as before.
\[ xy^{2}z = 0^{p+k}1^{p} \]
Verify the pumped string is outside the language, for every split
Why: Since the middle part was nonempty, the count of zeros strictly increased while the count of ones did not, so the two counts differ. That holds for every legal split, because the early-split condition forced the middle into the zeros. You win regardless of how the adversary played.
\[ p + k \neq p \;\Rightarrow\; 0^{p+k}1^{p} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Play the game on a language that is not regular", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Since the middle part was nonempty, the count of zeros strictly increased while the count of ones did not, so the two counts differ. That holds for every legal split, because the early-split condition forced the middle into the zeros. You win regardless of how the adversary played.
Ranking
Put in order
Put the moves of Play the game against the palindromes into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The palindromes over the alphabet of a and b are the strings that read the same forwards and backwards.
Worked example
A second win, on a language whose condition is about shape rather than counting.
Take the language and let the adversary move
Why: The palindromes over the alphabet of a and b are the strings that read the same forwards and backwards. The adversary names some pumping length.
Choose a string built from p
Why: Take p copies of a, then a single b, then p copies of a again. It reads the same both ways, so it is in the language, and its length exceeds p.
\[ w = a^{p}\,b\,a^{p} \]
Use the early-split condition
Why: The first p symbols are all a's, so the middle part is a nonempty run of a's taken from the leading block.
\[ y = a^{k}, \qquad 1 \le k \le p \]
Pump and inspect the shape
Why: Taking two copies gives more a's before the b than after it. The string no longer reads the same in both directions.
\[ xy^{2}z = a^{p+k}\,b\,a^{p} \]
Verify the pumped string fails for every legal split
Why: Reversing it puts the longer run of a's after the b instead of before, so it differs from the original unless the two runs are equal — and they are not, since the middle part was nonempty. The contradiction holds whatever split the adversary chose.
\[ k > 0 \;\Rightarrow\; a^{p+k}ba^{p} \text{ is not a palindrome} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Play the game against the palindromes", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Reversing it puts the longer run of a's after the b instead of before, so it differs from the original unless the two runs are equal — and they are not, since the middle part was nonempty. The contradiction holds whatever split the adversary chose.
Concept
Knowing how a competent adversary plays tells you which strings are bad choices, before you waste a proof on one.
Each strategy corresponds to one requirement in the string-choice pattern of Section 5. A string satisfying all four requirements leaves the adversary with no strategy at all.
Intuition
The game is not a heuristic. A winning strategy for you is a proof, and the reason is worth stating.
If the language were regular, the lemma guarantees the adversary has a reply at every move: a valid pumping length exists, and for every long string a valid split exists. So a regular language gives the adversary a strategy that never loses.
Therefore, if you can win against every possible play, no such guarantee can hold, so the language is not regular. The game is exactly the contrapositive of the lemma, dressed up.
\[ \text{you always win} \;\Rightarrow\; \text{the lemma's promise fails} \;\Rightarrow\; \text{not regular} \]
Constraint
Discussion prompt
Run The game, written as a proof with this step confiscated:
By the lemma it splits into three parts with the middle nonempty and the first two parts at most p long.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Every non-regularity proof is this game transcribed. Use these five sentences as a template.
Sentence four is the load-bearing one. If your string does not let you pin down what the middle part must be, the proof cannot be completed and you need a different string.
Edge cases
Discussion prompt
The game, written as a proof works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Every non-regularity proof is this game transcribed. Use these five sentences as a template.
Check
Think about the order of play.
Check your understanding
In the adversary game, may your choice of string depend on the pumping length?
Answer: A
Why: The pumping length is existentially quantified outermost, so it is played first. Your string may therefore be built from it, which is exactly why the standard move is a string with p copies of some symbol.
Section
Section 4
Concept
The lemma is a theorem, not an axiom. Proving it is short, and it explains where each of the three conditions comes from.
Assume the language is regular, so some machine recognizes it. Let the pumping length be the number of states of that machine.
\[ p = |Q| \]
That choice is the whole trick. Every other part of the proof follows from the pigeonhole principle applied to a run of length at least p.
Step zero
Discussion prompt
Prove the pumping lemma — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Take a long accepted string and look at its run
Answer:
Worked example
The full proof, in five steps.
Take a long accepted string and look at its run
Why: Let the string have length at least p. Its run visits one more state than the string has symbols, so it visits at least p plus one states.
\[ |w| = n \ge p \;\Rightarrow\; \text{the run visits } n+1 \text{ states} \]
Apply the pigeonhole principle to the first p plus one visits
Why: Those visits land in a set of only p states, so two of them coincide. Say the run is in the same state after reading j symbols and after reading k symbols, with j strictly less than k.
\[ 0 \le j < k \le p, \qquad \hat{\delta}(q_0, w_{1..j}) = \hat{\delta}(q_0, w_{1..k}) \]
Cut the string at those two points
Why: The first part is what was read before the first visit, the middle is what was read between the two visits, and the third is the rest.
\[ x = w_{1..j}, \quad y = w_{j+1..k}, \quad z = w_{k+1..n} \]
Check the two side conditions
Why: The middle is nonempty because the two indices differ, and the first two parts together have length k, which is at most p because both indices were taken from the first p plus one visits.
\[ |y| = k - j > 0, \qquad |xy| = k \le p \]
Verify that pumping preserves the final state, and conclude
Why: Reading the middle part takes the machine from the repeated state back to itself, so it may be read any number of times — or none — without changing where the run is. The remaining part then drives it to the same accepting state as before, so every pumped string is accepted and therefore in the language.
\[ \hat{\delta}(q_0, xy^{i}z) = \hat{\delta}(q_0, xyz) \in F \;\forall i \ge 0 \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove the pumping lemma", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The middle is nonempty because the two indices differ, and the first two parts together have length k, which is at most p because both indices were taken from the first p plus one visits.
Concept
The proof takes the pumping length to be the state count, but nothing breaks if it is larger, and that flexibility is occasionally useful.
A longer string still forces a repeat within its first p plus one visits, since those visits still outnumber the states available. So every value at least as large as the state count is a valid pumping length.
\[ p \ge |Q| \;\Rightarrow\; p \text{ is a valid pumping length} \]
This is why a proof may never assume the pumping length is small, or even that it exceeds any particular constant. The adversary picks it, and they may pick it enormous.
Intuition
Reading the proof backwards makes it obvious that pumpability cannot characterize regularity.
The proof derives pumpability from a machine. Nothing in it says a pumpable language must have one — the argument simply never runs in that direction.
Concretely, a language can happen to contain enough repetitive strings to satisfy the condition while still requiring unbounded memory to decide. Lesson 11 exhibits one. The moral is to treat the lemma strictly as a one-way tool.
\[ \text{machine} \;\Rightarrow\; \text{pumpable}, \quad \text{pumpable} \;\not\Rightarrow\; \text{machine} \]
Concept
Tracing the three conditions back to their source makes them memorable, and makes it obvious why they cannot be strengthened.
| Condition | Source |
|---|---|
| pumping preserves membership | the middle part is a cycle on the repeated state |
| the middle is nonempty | the two pigeonhole indices are distinct |
| the first two parts are at most p long | the pigeonhole was applied to the first p plus one visits only |
Notice what is not claimed: nothing says the middle is short, nothing says the first part is nonempty, and nothing says the split is unique. Those would all be false.
Explain it
Discussion prompt
Explain Where each condition came from to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Tracing the three conditions back to their source makes them memorable, and makes it obvious why they cannot be strengthened.
Intuition
Students often treat the pumping length as an unknowable constant. In the proof it has a concrete value.
It is the number of states of some machine for the language. Any larger value also works, since a longer string still forces a repeat within its first p plus one visits.
\[ p = |Q| \text{ works, and so does any } p' \ge |Q| \]
In a non-regularity proof you never learn its value, because you are arguing against every machine at once. But knowing where it comes from is what makes the adversary's move feel principled rather than arbitrary.
Analogy
Discussion prompt
Explain The pumping length is not mysterious by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Students often treat the pumping length as an unknowable constant. In the proof it has a concrete value.
Picture it
Figure (svg): Automaton with states r0, r1, r2
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Make the proof concrete by carrying it out on an actual machine and string.
Worked example
Make the proof concrete by carrying it out on an actual machine and string.
Figure (svg): Automaton with states r0, r1, r2
Choose a string long enough
Why: The machine has three states, so take a string of length at least three that it accepts. Take abaa.
Write out the run
Why: Reading a moves to the middle state, b returns to the start, a moves to the middle again, and a reaches the accepting state.
\[ r_0 \xrightarrow{a} r_1 \xrightarrow{b} r_0 \xrightarrow{a} r_1 \xrightarrow{a} r_2 \]
Find the first repeat
Why: The start state appears at visit zero and visit two. Those indices are both within the first four visits, so the early-split condition is satisfied.
\[ j = 0, \quad k = 2, \quad k \le p = 3 \]
Read off the split
Why: The first part is empty, the middle is the first two symbols, and the third part is the remainder.
\[ x = \varepsilon, \quad y = ab, \quad z = aa \]
Verify that pumping this split really works
Why: Zero copies gives aa, one gives abaa, two give ababaa. Running each: aa reaches the accepting state, and each extra ab returns the machine to the start before the final aa. All three are accepted, exactly as the proof promises.
\[ aa,\ abaa,\ ababaa \in L(M) \ \checkmark \]
Notation
Annotate
From Locate the cycle in a concrete run — read this one piece at a time. What is each part doing?
On: \( aa,\ abaa,\ ababaa \in L(M) \ \checkmark \)
Commit first
Predict first
In the proof of the pumping lemma, what is the pumping length taken to be?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: The number of states of a machine recognizing the language
Why: Setting the pumping length to the state count is what makes the pigeonhole argument fire: a string that long produces more visits than there are states, forcing a repeat within the first p plus one of them.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Recall what value the proof gives the pumping length.
Check your understanding
In the proof of the pumping lemma, what is the pumping length taken to be?
Answer: A
Why: Setting the pumping length to the state count is what makes the pigeonhole argument fire: a string that long produces more visits than there are states, forcing a repeat within the first p plus one of them.
Section
Section 5
Pattern
Almost all the difficulty in these proofs is the choice of string. Four requirements, and they must all hold.
Requirement three is the one that turns a hard proof into an easy one. A string beginning with p copies of a single symbol leaves the adversary no room.
Intuition
The early-split condition says the first two parts together are at most p symbols long. That is a statement about the prefix of your string.
If the first p symbols are all the same, then whatever the adversary chooses, the middle part is a nonempty run of that one symbol. You know its composition exactly, even though you do not know its length or position.
\[ w = 0^{p}\cdots \;+\; |xy| \le p \;\Rightarrow\; y = 0^{k}, \; k > 0 \]
If instead the prefix were mixed, the middle could be any of several shapes and each would need its own case. That is why a badly chosen string turns a five-line proof into a three-case slog.
Counterexample
Discussion prompt
The early-split condition says the first two parts together are at most p symbols long. That is a statement about the prefix of your string.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
If instead the prefix were mixed, the middle could be any of several shapes and each would need its own case. That is why a badly chosen string turns a five-line proof into a three-case slog.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Prove that the strings of some number of zeros followed by the same number of ones is not regular.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Take a string with a hundred zeros and a hundred ones.
Prove that the strings of some number of zeros followed by the same number of ones is not regular.
Why: Take a string with a hundred zeros and a hundred ones. It is in the language and it is certainly long.
Trap
Prove that the strings of some number of zeros followed by the same number of ones is not regular.
Pick a concrete long string
Why: Take a string with a hundred zeros and a hundred ones. It is in the language and it is certainly long.
\[ w = 0^{100}1^{100} \]
Split it and pump
Why: If the middle part lies in the zeros, pumping unbalances the counts. Declare the contradiction and finish.
Notice the gap when the adversary names a large pumping length
Why: The adversary may name a pumping length of a thousand. Then the chosen string is shorter than p, the lemma imposes nothing about it, and the argument never starts.
\[ |w| = 200 < 1000 = p \]
Prove that the strings of some number of zeros followed by the same number of ones is not regular.
Build the string from the pumping length
Why: Take p zeros followed by p ones, whatever p turns out to be. It is in the language and its length is at least p by construction.
\[ w = 0^{p}1^{p}, \qquad |w| = 2p \ge p \]
Now the early-split condition bites
Why: The first p symbols are all zeros, so the middle part is a nonempty run of zeros — no matter what value the adversary chose for p.
\[ |xy| \le p \;\Rightarrow\; y = 0^{k}, \; k > 0 \]
Pump and finish
Why: Two copies give more zeros than ones, so the result is outside the language. The argument holds for every p, which is what the proof requires.
\[ 0^{p+k}1^{p} \notin L \ \checkmark \]
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Concept
Once the middle part is pinned down, the repetition count is usually forced, and there are only three sensible attempts.
| Count | Effect | Good against |
|---|---|---|
| zero | deletes the middle part, shortening the string | exact-count conditions |
| two | duplicates it, lengthening the string | exact-count conditions |
| a large value | makes one part dominate | ordering or ratio conditions |
For a counting language either zero or two works, and it is worth stating which you chose rather than saying 'pump it'. For languages about primes or squares, a carefully chosen large value is needed, as Lesson 11 shows.
Comparison
Comparison matrix
From Choosing the repetition count: refill the Effect column from what you know. The rest of the table is as it appeared.
| Count | Effect | Good against |
|---|---|---|
| zero | deletes the middle part, shortening the string | exact-count conditions |
| two | duplicates it, lengthening the string | exact-count conditions |
| a large value | makes one part dominate | ordering or ratio conditions |
Step zero
Discussion prompt
Set up a proof without finishing it — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Assume regularity and take the pumping length
Answer:
Worked example
Practise the setup alone, on the language of strings whose length is a perfect square. The finish belongs to Lesson 11.
Assume regularity and take the pumping length
Why: Suppose the language were regular and let p be the pumping length the lemma provides.
Choose a string in the language, built from p
Why: The length must be a perfect square and at least p. Take a string of p squared symbols, all the same.
\[ w = a^{p^{2}}, \qquad |w| = p^{2} \ge p \]
Apply the lemma and pin down the middle
Why: The split has a nonempty middle inside the first p symbols, so the middle is a run of between one and p copies of the symbol.
\[ y = a^{k}, \qquad 1 \le k \le p \]
Identify what pumping controls
Why: Pumping changes only the length, adding k symbols per extra copy. So the question becomes whether the new length can still be a perfect square.
\[ |xy^{2}z| = p^{2} + k \]
Verify the setup is sound before going further
Why: The string is in the language, its length exceeds p, and the middle part is pinned to a bounded nonempty run — all four requirements from the pattern hold. The remaining work is purely arithmetic: showing the new length falls strictly between two consecutive squares, which Lesson 11 completes.
\[ p^{2} < p^{2}+k \le p^{2}+p < (p+1)^{2} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Set up a proof without finishing it", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The string is in the language, its length exceeds p, and the middle part is pinned to a bounded nonempty run — all four requirements from the pattern hold. The remaining work is purely arithmetic: showing the new length falls strictly between two consecutive squares, which Lesson 11 completes.
Ranking
Put in order
Put the moves of Set up a proof for a language with two blocks into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The strings over the alphabet of a and b consisting of some number of a's followed by strictly more b's.
Worked example
A second setup, on a language where the obvious string needs a moment's thought.
State the language
Why: The strings over the alphabet of a and b consisting of some number of a's followed by strictly more b's.
\[ L = \{\, a^{m}b^{n} : n > m \ge 0 \,\} \]
Choose a string built from p
Why: Take p a's followed by p plus one b's. The count condition holds, and the length exceeds p.
\[ w = a^{p}b^{p+1} \]
Pin down the middle part
Why: The first p symbols are all a's, so the middle part is a nonempty run of a's from the leading block.
\[ y = a^{k}, \qquad 1 \le k \le p \]
Choose the repetition count deliberately
Why: Repeating adds a's, which pushes the counts closer together. Taking two copies gives p plus k a's against p plus one b's, and since k is at least one the a's now weakly outnumber.
\[ xy^{2}z = a^{p+k}b^{p+1} \]
Verify the contradiction holds for every legal k
Why: With k at least one, the a-count is at least p plus one, which is not strictly less than the b-count of p plus one. So the condition fails for every allowed k, and the pumped string is outside the language regardless of the adversary's split.
\[ p + k \ge p+1 \;\Rightarrow\; a^{p+k}b^{p+1} \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Set up a proof for a language with two blocks", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: With k at least one, the a-count is at least p plus one, which is not strictly less than the b-count of p plus one. So the condition fails for every allowed k, and the pumped string is outside the language regardless of the adversary's split.
Concept
Non-regular languages met in practice fall into a few recognizable families, and the family suggests the string to choose.
| Family | Example | String to try |
|---|---|---|
| two counts must match | equal zeros then ones | p of the first, p of the second |
| a block must be repeated | a string followed by itself | a uniform block, then a marker |
| shape must be symmetric | palindromes | uniform block, marker, uniform block |
| length must satisfy an arithmetic condition | length a perfect square | a uniform run of the required length |
In every row the string opens with a uniform run of length at least p, which is what makes the early-split condition do the work. Lesson 11 works through one example of each.
Comparison
Comparison matrix
From The languages these proofs usually target: refill the String to try column from what you know. The rest of the table is as it appeared.
| Family | Example | String to try |
|---|---|---|
| two counts must match | equal zeros then ones | p of the first, p of the second |
| a block must be repeated | a string followed by itself | a uniform block, then a marker |
| shape must be symmetric | palindromes | uniform block, marker, uniform block |
| length must satisfy an arithmetic condition | length a perfect square | a uniform run of the required length |
Prediction
Predict first
You are proving the equal-zeros-then-ones language is not regular, and the middle part has been pinned to a nonempty run of zeros. Which repetition count works?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Either zero or two — both change the zero count while leaving the ones alone
Why: Deleting the middle part removes zeros and duplicating it adds them, while in both cases the block of ones is untouched. Since the middle was nonempty, either move makes the two counts differ, which breaks the defining condition.
Check
Think about which count breaks an exact-match condition fastest.
Check your understanding
You are proving the equal-zeros-then-ones language is not regular, and the middle part has been pinned to a nonempty run of zeros. Which repetition count works?
Answer: A
Why: Deleting the middle part removes zeros and duplicating it adds them, while in both cases the block of ones is untouched. Since the middle was nonempty, either move makes the two counts differ, which breaks the defining condition.
Intuition
It is worth knowing in advance which languages resist this tool, so you do not spend an hour on an impossible proof.
The lemma proves non-regularity by breaking a condition through pumping. If a language's defining condition survives pumping — because it is loose, or because it is closed under the repetitions involved — no choice of string will work.
Some non-regular languages genuinely satisfy the pumping condition. For those the lemma is simply silent, and the Myhill-Nerode argument from Lesson 6 must be used instead. Lesson 11 shows one such example.
\[ \text{pumping fails} \;\Rightarrow\; \text{not regular}, \quad \text{but not conversely} \]
Elimination
Eliminate the wrong options
Why is a string beginning with p copies of a single symbol such a good choice?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The condition bounds the combined length of the first two parts by p, so the middle part lies entirely within the first p symbols. Making those symbols uniform means you know exactly what the middle consists of, without knowing where it starts or how long it is.
Check
Recall which requirement the early-split condition exploits.
Check your understanding
Why is a string beginning with p copies of a single symbol such a good choice?
Answer: A
Why: The condition bounds the combined length of the first two parts by p, so the middle part lies entirely within the first p symbols. Making those symbols uniform means you know exactly what the middle consists of, without knowing where it starts or how long it is.
Ranking
Put in order
These are the steps of The complete proof template, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Collected, so it can be applied directly to the examples in Lesson 11.
Write the five steps out in full every time, even when the language is easy. Skipping step four is what produces proofs that look convincing and prove nothing.
Real world
Discussion prompt
Outside this lesson: where does The Pumping Lemma for Regular Languages actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The complete proof template is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
Lesson 10 supplies the first tool for proving a language has no finite automaton. Covers why every earlier technique proves only the positive direction, how finiteness plus the pigeonhole principle forces a repeated state and hence a pumpable cycle, the exact statement with its three conditions and the quantifier order, the adversary game that makes the alternation manageable, the proof of the lemma itself with each condition traced to its source, and how to set a proof up: choosing a string built from the pumping length with a uniform prefix, and choosing the repetition count.
Concept
This lesson has the statement, the game, the proof and the setup. The remaining skill is choosing well when the language resists the obvious move.
All of them use the template above unchanged. What varies is the string, and Lesson 11 is essentially a catalogue of good choices.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why a New Tool Is Needed · The Statement · The Adversary Game · Proving the Lemma · Setting Up a Proof. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You have the first tool in this course that proves a language has no machine, and you know exactly what it does and does not claim.
| Situation | Move |
|---|---|
| asked to show a language is regular | build a machine or an expression — not this lemma |
| asked to show one is not | the five-step template |
| choosing a string | build it from p, uniform in the first p symbols |
| choosing a repetition count | try zero, then two, then a large value |
| the language pumps but feels irregular | use distinguishable prefixes instead |
Lesson 11 applies the template to a catalogue of languages, including the ones where the obvious string does not work.
Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.