Lesson 7 establishes that the regular languages are closed under the three regular operations and more. It explains what closure means and why it is a theorem rather than a definition, then does union by both the product construction and the free-move construction, concatenation as a guessed split with its demotion trap, and Kleene star with the empty-string subtlety and why a naive loop-back is wrong. It adds intersection, complement, difference, reversal, and homomorphism, and ends by using closure properties as a proof technique, including the standard trick of proving a language non-regular by intersecting it with a simple regular pattern.
Subject: Theory of Computation · 115 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Theory of Computation · Lesson 7
Combine two regular languages and you get another one. Proving that — by building machines — is what makes regular expressions possible.
Objectives
Lessons 4 to 6 built machines and proved the two models equivalent. This lesson uses that freedom to combine languages. By the end you can:
Warm-up
Discussion prompt
Before we open Regular Operations & Closure: without looking back, what was the main idea of Equivalence of DFAs & NFAs, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Lesson 6 proves that nondeterminism buys convenience but not power. It states the equivalence theorem and disposes of its trivial direction, then builds the subset construction component by component, including the empty subset as a dead state and the accepting rule people get wrong, the epsilon-closure in the start subset, and the transition rule. It works conversions from a worklist both with and without free moves, proves correctness by induction on the input length, and gives the pigeonhole lower bound showing that the exponential blowup is unavoidable. It ends with the consequences: complementing an NFA in the right order, and deciding emptiness and infiniteness by graph search.
Section
Section 1
Concept
A class of languages is closed under an operation when applying that operation to members of the class always produces another member.
\[ L_1, L_2 \in \mathcal{R} \;\Longrightarrow\; L_1 \circ L_2 \in \mathcal{R} \]
closure property — A theorem stating that a class of languages is preserved by a given operation — the result never escapes the class.
The word is borrowed from algebra. The integers are closed under addition and multiplication but not under division, since dividing 1 by 2 leaves the integers behind. Closure is never automatic; it must be checked operation by operation.
Counterexample
Discussion prompt
A class of languages is closed under an operation when applying that operation to members of the class always produces another member.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Intuition
It is tempting to read 'the regular languages are closed under union' as part of what regular means. It is not. Regular was defined in Lesson 4 by the existence of a machine, and nothing in that definition mentions union.
So the claim has real content: given two machines, one must produce a third. Every closure proof in this lesson is a construction plus an argument that the construction is correct.
That also explains why the failures matter. Later classes in this course are closed under some of these operations and not others, and knowing which is often the fastest way to tell two classes apart.
Concept
Three operations on languages are singled out and called the regular operations, because they are exactly what regular expressions will be built from in Lesson 8.
| Operation | Members of the result | Read as |
|---|---|---|
| union | strings in either language | either one |
| concatenation | a string of the first followed by one of the second | one then the other |
| Kleene star | any number of members joined together | zero or more copies |
Union is set-theoretic and familiar. The other two are new: they are operations on strings lifted to sets of strings, and both involve cutting a string into pieces.
Comparison
Comparison matrix
From The three regular operations: refill the Members of the result column from what you know. The rest of the table is as it appeared.
| Operation | Members of the result | Read as |
|---|---|---|
| union | strings in either language | either one |
| concatenation | a string of the first followed by one of the second | one then the other |
| Kleene star | any number of members joined together | zero or more copies |
Concept
Both new operations deserve their definitions written out, because the boundary cases are where the errors live.
\[ L_1L_2 = \{\, xy \;:\; x \in L_1,\ y \in L_2 \,\} \]
Note that the concatenation is not 'strings containing a member of the first followed by a member of the second' — the two pieces must together account for the entire string, with nothing left over.
For star, the case of zero copies is the empty product, which is the empty string. So the empty string belongs to the star of every language, including the empty one.
\[ L^{*} = \bigcup_{k \ge 0} L^{k}, \qquad L^{0} = \{\varepsilon\}, \qquad \varnothing^{*} = \{\varepsilon\} \]
Analogy
Discussion prompt
Explain Concatenation and star, precisely by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Both new operations deserve their definitions written out, because the boundary cases are where the errors live.
Concept
Star is defined through powers, so the powers are worth their own table.
\[ L^{0} = \{\varepsilon\}, \qquad L^{k+1} = L^{k}L \]
So the first power is the language itself, the second is every string formed by concatenating two members (not necessarily distinct), and so on.
| Power | For the language holding just a and ab |
|---|---|
| zeroth | the empty string |
| first | a, ab |
| second | aa, aab, aba, abab |
| star | all of the above, and every longer combination |
Look at the top row. Even for the empty language, the zeroth power holds the empty string — because the empty product consults no member at all. That is exactly why starring the empty language gives the empty string back, rather than nothing.
Trade off
Comparison matrix
From Powers of a language: every row here is a choice with a cost. Fill the For the language holding just a and ab column, then say which row you would actually pick and what you give up for it.
| Power | For the language holding just a and ab |
|---|---|
| zeroth | the empty string |
| first | a, ab |
| second | aa, aab, aba, abab |
| star | all of the above, and every longer combination |
Intuition
Two misreadings account for most early mistakes, and both are worth stating explicitly.
Concatenation is not 'both conditions hold'. It cuts the string into two consecutive parts, each satisfying its own condition. The parts are disjoint stretches of the input, not two views of the whole.
Star does not repeat one fixed member. Each repetition may choose a different member of the language, independently of the others. That independence is what makes star powerful and what makes it easy to over-claim.
\[ L = \{a, b\} \;\Rightarrow\; ab \in L^{*} \]
Explain it
Discussion prompt
Explain Concatenation is not intersection, and star is not repetition of one string to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Two misreadings account for most early mistakes, and both are worth stating explicitly.
Ranking
Put in order
These are the steps of How to prove a closure property, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Every proof in this lesson has the same five-part shape. Recognizing it makes the rest of the lesson repetitive in the good sense.
Steps three and four are the two inclusions. Skipping the second is the most common gap: it is the direction that catches a machine accepting too much.
Elimination
Eliminate the wrong options
How many strings are in the concatenation of those two languages?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Pairing each member of the first with each member of the second gives ab, a, abb and ab. The string ab appears twice, and a set holds no duplicates, so the result is the three-string set holding a, ab and abb.
Check
Let the first language hold the strings a and ab, and let the second hold b and the empty string.
Check your understanding
How many strings are in the concatenation of those two languages?
Answer: A
Why: Pairing each member of the first with each member of the second gives ab, a, abb and ab. The string ab appears twice, and a set holds no duplicates, so the result is the three-string set holding a, ab and abb.
Ranking
Put in order
Put the moves of See closure fail somewhere familiar into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Take the integers as the class and division as the operation.
Worked example
Before proving closures, it helps to see what a closure failure looks like, so the proofs feel like they are doing work.
Pick a class and an operation
Why: Take the integers as the class and division as the operation. Both are entirely familiar, which is the point.
Exhibit two members whose result escapes
Why: Take 1 and 2. Both are integers, and their quotient is not.
\[ 1, 2 \in \mathbb{Z} \quad\text{but}\quad 1/2 \notin \mathbb{Z} \]
Note the asymmetry of evidence
Why: One escaping pair refutes the property outright. No number of pairs that stay inside would ever establish it, because closure quantifies over all pairs.
Carry the lesson over
Why: So each of the constructions ahead must work for every pair of regular languages, not for the examples they are demonstrated on. That is why the proofs are general and the diagrams are only illustrations.
Verify the counterexample has the right shape
Why: Both inputs belong to the class and the result does not — exactly the shape a closure failure must have. The same shape appears again in Section 5, where the context-free languages fail to be closed under intersection.
\[ a, b \in \mathcal{C} \; \text{ and } \; a \circ b \notin \mathcal{C} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "See closure fail somewhere familiar", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Both inputs belong to the class and the result does not — exactly the shape a closure failure must have. The same shape appears again in Section 5, where the context-free languages fail to be closed under intersection.
Section
Section 2
Concept
The first construction runs both machines at once on the same input, exactly as in Lesson 4. It stays entirely deterministic.
\[ Q = Q_1 \times Q_2, \qquad \delta\big((p,q),a\big) = \big(\delta_1(p,a), \delta_2(q,a)\big) \]
The only choice left is which pairs to accept. For union, accept a pair when either component is accepting.
\[ F = (F_1 \times Q_2) \;\cup\; (Q_1 \times F_2) \]
Read that as: accept if the first component is accepting or the second is. Since each component tracks its own machine faithfully, the pair accepts exactly the union.
Step zero
Discussion prompt
Prove the product construction correct — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: State the invariant
Answer:
Worked example
Show that the product machine accepts exactly the union of the two languages.
State the invariant
Why: After reading any string, each component of the pair is the state that component's own machine would have reached on that same string. This is immediate from the transition rule, by induction on length.
\[ \hat{\delta}\big((q_1,q_2), w\big) = \big(\hat{\delta}_1(q_1,w), \hat{\delta}_2(q_2,w)\big) \]
Take a string of the union and show it is accepted
Why: It lies in at least one of the two languages, so at least one component finishes accepting. The accepting rule for pairs then fires.
Take a string the machine accepts and show it lies in the union
Why: Accepting means some component finished accepting, so that component's machine accepts the string, so the string belongs to that language and hence to the union.
\[ (\hat{\delta}_1(q_1,w) \in F_1) \; \text{ or } \; (\hat{\delta}_2(q_2,w) \in F_2) \]
Note the free bonus
Why: Changing only the accepting set turns this into intersection or difference. The states, arrows and start state are untouched, which is why Section 5 gets those properties almost for nothing.
Verify the accepting condition on each half separately
Why: A string of the first language drives the first component into its accepting set, so the pair accepts no matter what the second does; the mirror argument covers the second language. A string in neither leaves both components outside, so the pair rejects. Both inclusions hold.
\[ (p,q) \in F \iff p \in F_1 \ \text{ or } \ q \in F_2 \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove the product construction correct", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: A string of the first language drives the first component into its accepting set, so the pair accepts no matter what the second does; the mirror argument covers the second language. A string in neither leaves both components outside, so the pair rejects. Both inclusions hold.
Concept
The second construction is nondeterministic and much shorter. It is legitimate because Lesson 6 proved the models equivalent.
Place the two machines side by side with disjoint state sets, add one fresh start state, and give it an ε-arrow into each original start state.
\[ Q = \{s\} \cup Q_1 \cup Q_2, \qquad \delta(s,\varepsilon) = \{q_1, q_2\} \]
Keep both accepting sets, unioned together. The new start state has no other transitions, so the only thing it can do is hand the input to both machines at once.
The disjointness requirement is not cosmetic: if the two machines shared a state name, arrows from one would be confused with arrows from the other.
\[ Q_1 \cap Q_2 = \varnothing \quad \text{(rename if necessary)} \]
Estimation
Predict first
Show that wiring a fresh start state into both machines recognizes exactly the union.
Commit before you compute: what does Prove the free-move construction correct come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify that the new start state cannot cheat
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. It has only its two ε-arrows and no incoming arrows and no symbol-arrows, so every accepting run leaves it once, at the very beginning, and then commits to one machine for good.
Worked example
Show that wiring a fresh start state into both machines recognizes exactly the union.
Describe what the opening closure does
Why: Before any symbol is read, the ε-closure of the new start state contains both original start states. So both machines begin running on the whole input.
\[ E\big(\{s\}\big) = \{s, q_1, q_2\} \]
Take a string of the union and build an accepting path
Why: It lies in one of the two languages, so that machine has an accepting path on it. Prefixing the ε-arrow into that machine's start state gives an accepting path in the combined machine.
Take an accepting path and locate which machine it used
Why: The path leaves the new start state by exactly one ε-arrow and then stays inside that machine, since the state sets are disjoint and no arrow crosses between them.
Note where disjointness was used
Why: Exactly there. Without it a path could drift from one machine into the other mid-string and accept something in neither language.
\[ Q_1 \cap Q_2 = \varnothing \]
Verify that the new start state cannot cheat
Why: It has only its two ε-arrows and no incoming arrows and no symbol-arrows, so every accepting run leaves it once, at the very beginning, and then commits to one machine for good. No run can mix the two, so nothing outside the union is ever accepted.
\[ \delta(s,a) = \varnothing \ \text{ for every } a \in \Sigma \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove the free-move construction correct", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: It has only its two ε-arrows and no incoming arrows and no symbol-arrows, so every accepting run leaves it once, at the very beginning, and then commits to one machine for good. No run can mix the two, so nothing outside the union is ever accepted.
Intuition
Both are correct, so the choice is about what you want out of it.
| Product machine | Fresh start with free moves | |
|---|---|---|
| result is | deterministic | nondeterministic |
| state count | the product of the two | the sum, plus one |
| also gives | intersection and difference | nothing else |
| best when | you need a DFA, or need intersection | you want the shortest proof |
The wiring construction generalizes: the same idea handles concatenation and star, and Lesson 8 turns all three into a single mechanical procedure. The product construction does not generalize at all — there is no product machine for concatenation.
Missing information
Discussion prompt
Let the first language be the strings over the alphabet of a and b with an even number of a's, and the second be the strings ending in b. Build both union machines and compare.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The first is a two-state parity machine on a, ignoring b. The second is a two-state machine remembering whether the last symbol was b.
Worked example
Let the first language be the strings over the alphabet of a and b with an even number of a's, and the second be the strings ending in b. Build both union machines and compare.
Build the two ingredient machines
Why: The first is a two-state parity machine on a, ignoring b. The second is a two-state machine remembering whether the last symbol was b.
Build the product machine
Why: Four pairs. A pair is accepting when the parity component is even, or the last-symbol component says b, or both.
\[ |Q| = 2 \cdot 2 = 4, \qquad |F| = 3 \]
Build the wiring machine
Why: One fresh start state plus the two originals, joined by two ε-arrows. Five states, and nondeterministic.
\[ |Q| = 1 + 2 + 2 = 5 \]
Note when each construction wins
Why: The product is smaller here and already deterministic. The wiring machine would win if the ingredients were large, since sums beat products — and it is the one that keeps working for the operations ahead.
Verify both machines on the same four strings
Why: Take one string in the first language only, one in the second only, one in both, and one in neither: aa, b, aab and a. Both machines accept the first three and reject the last, so they agree everywhere it matters.
\[ aa,\ b,\ aab \in L_1 \cup L_2 \qquad a \notin L_1 \cup L_2 \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Build a union machine concretely", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Take one string in the first language only, one in the second only, one in both, and one in neither: aa, b, aab and a. Both machines accept the first three and reject the last, so they agree everywhere it matters.
Prediction
Predict first
Using the product construction for their union, how many states does the result have?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Twelve
Why: The product machine's states are pairs, one component from each machine, so the count is the product of the two counts. Choosing which pairs are accepting selects the operation, but never changes how many states there are.
Check
One DFA has 3 states and another has 4 states.
Check your understanding
Using the product construction for their union, how many states does the result have?
Answer: A
Why: The product machine's states are pairs, one component from each machine, so the count is the product of the two counts. Choosing which pairs are accepting selects the operation, but never changes how many states there are.
Section
Section 3
Concept
To accept the concatenation, the machine must verify that the input can be cut into a prefix drawn from the first language and a suffix drawn from the second. But the cut is not marked anywhere in the input.
A deterministic machine has no way to try several cuts. It reads each symbol once and must commit. That is precisely the situation nondeterminism was introduced for in Lesson 5.
So the construction guesses. At every moment the machine may decide 'the first half ends here', and the acceptance rule credits it if any such guess works out.
\[ w \in L_1L_2 \iff \exists \text{ a split } w = xy \text{ with } x \in L_1,\ y \in L_2 \]
Concept
Take the two NFAs with disjoint state sets, and glue them together with ε-arrows.
Point four is the one people drop. If the first machine's accepting states stayed accepting, the machine would accept the prefix on its own — so the whole first language would slip into the result, which is wrong unless the second language contains the empty string.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Concatenate the machine for the single string a onto the machine for the single string b. The concatenation holds exactly one string: ab.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The first machine's accepting state still looks like a finish line, so leave it in the accepting set and simply add the crossing arrow on top.
Concatenate the machine for the single string a onto the machine for the single string b. The concatenation holds exactly one string: ab.
Why: The first machine's accepting state still looks like a finish line, so leave it in the accepting set and simply add the crossing arrow on top.
Trap
Concatenate the machine for the single string a onto the machine for the single string b. The concatenation holds exactly one string: ab.
Wire the ε-arrow and keep both accepting sets
Why: The first machine's accepting state still looks like a finish line, so leave it in the accepting set and simply add the crossing arrow on top.
\[ F = F_1 \cup F_2 \]
Trace the input a
Why: After reading a, the run is sitting on the first machine's accepting state — which is still accepting. So the machine says yes to a, and the language it recognizes is too big.
\[ a \in L(N) \quad \text{but} \quad a \notin \{ab\} \]
Concatenate the machine for the single string a onto the machine for the single string b. The concatenation holds exactly one string: ab.
Wire the ε-arrow and demote the first accepting set
Why: Finishing the first half is not finishing the job. Only the second machine's accepting states may end an accepting run.
\[ F = F_2 \]
Trace the input a
Why: After reading a the run reaches the first machine's accepting state and may cross by ε — but it then sits at the second machine's start state, which is not accepting. So a is rejected and ab is accepted, exactly as required.
\[ a \notin L(N), \qquad ab \in L(N) \ \checkmark \]
Notation
Annotate
From Trap: leaving the first machine's accepting states accepting — read this one piece at a time. What is each part doing?
On: \( a \in L(N) \quad \text{but} \quad a \notin \{ab\} \)
Intuition
Every time the run reaches an accepting state of the first machine it faces a decision: cross now, or keep reading inside that machine.
The ε-arrow is exactly that decision drawn on the page. Taking it says 'the first half ended here'; declining it says 'not yet'. Because the machine explores both, every possible cut is tried at once.
This is why the construction needs no analysis of where the cut might be. The nondeterminism does the searching, and the acceptance rule keeps only the successful search.
\[ \text{one } \varepsilon\text{-arrow} \;\longleftrightarrow\; \text{one candidate split} \]
Sorting
Sort into buckets
These are the pieces of Regular Operations & Closure, out of order. Put each one back under the part of the lesson it belongs to.
Fill the middle
Fill in the blanks
From Prove the concatenation construction correct — finish the line. Write what belongs on the right of the equals sign before you look.
L_1L_2 \subseteq L(N) \ \textL_1L_2 \ \checkmark \ L(N) \subseteq L_1L_2 \;\Rightarrow\; L(N) = ___
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. By definition it splits as a prefix in the first language and a suffix in the second.
Worked example
Show that the wired machine accepts exactly the concatenation.
Take a string of the concatenation and build an accepting run
Why: By definition it splits as a prefix in the first language and a suffix in the second. Run the first machine on the prefix, arriving at one of its accepting states.
Take the free arrow at exactly the right moment
Why: From that accepting state, follow the new ε-arrow to the second machine's start. No input is consumed, so the machine is positioned exactly at the start of the suffix.
\[ \text{after reading } x: \ \text{a state of } F_1 \quad\xrightarrow{\ \varepsilon\ }\quad \text{the start state of } N_2 \]
Finish inside the second machine
Why: Run it on the suffix, arriving in its accepting set, which is the accepting set of the whole. So every string of the concatenation is accepted.
Now take an accepting run and read a split off it
Why: The run starts inside the first machine and ends in the second machine's accepting set, so at some point it crossed. The only crossings are the new ε-arrows, and they leave from accepting states of the first machine.
Verify that both inclusions were genuinely proved
Why: The first half took an arbitrary string of the concatenation and produced a run; the second took an arbitrary run and produced a split, with the prefix consumed before the crossing and the suffix after. Two inclusions in opposite directions give equality, and neither half assumed the other.
\[ L_1L_2 \subseteq L(N) \ \text{ and } \ L(N) \subseteq L_1L_2 \;\Rightarrow\; L(N) = L_1L_2 \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove the concatenation construction correct", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The first half took an arbitrary string of the concatenation and produced a run; the second took an arbitrary run and produced a split, with the prefix consumed before the crossing and the suffix after. Two inclusions in opposite directions give equality, and neither half assumed the other.
Step zero
Discussion prompt
Concatenate two small machines and test the boundary — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Build the two ingredients
Answer:
Worked example
Let the first language hold just the string a, and the second hold b and the empty string. Build the concatenation machine and check it.
\[ L_1L_2 = \{ab, a\} \]
Build the two ingredients
Why: The first has a start state and an a-arrow to its accepting state. The second has a start state that is itself accepting, because the empty string belongs, plus a b-arrow to a second accepting state.
Wire them
Why: Add an ε-arrow from the first machine's accepting state to the second machine's start. The start state of the whole is the first machine's start, and the accepting set is the second machine's.
Trace the input a
Why: Read a to reach the first accepting state, then close: the ε-arrow adds the second machine's start state, which is accepting. So a is accepted — correctly, since a is a followed by the empty string.
\[ \hat{\delta}(p_0, a) = \{p_1, r_0\}, \qquad \{p_1, r_0\} \cap F = \{r_0\} \neq \varnothing \]
See what would break if the first accepting state stayed accepting
Why: Nothing here, because the empty string in the second language makes a acceptable anyway. But with the second language holding only b, the string a would have to be rejected, and leaving that state accepting would wrongly accept it. Always demote.
Verify all four short strings at once
Why: The concatenation holds exactly a and ab, and both are accepted. The two nearest non-members, the empty string and the single b, are both rejected — the first because the a-arrow must be taken, the second because there is no b-arrow out of the start.
\[ a,\ ab \in L(N) \qquad \varepsilon,\ b \notin L(N) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Concatenate two small machines and test the boundary", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The concatenation holds exactly a and ab, and both are accepted. The two nearest non-members, the empty string and the single b, are both rejected — the first because the a-arrow must be taken, the second because there is no b-arrow out of the start.
Concept
It is worth seeing precisely what goes wrong if you try the same wiring deterministically.
Suppose you added a deterministic jump from each accepting state of the first machine to the second machine's start. Then reaching an accepting state of the first would force the machine to leave it immediately.
But the correct split might come later. If the first language holds a and aa while the second holds only a, then the three-symbol string must be cut after two symbols — yet a forced jump after the first symbol commits to a cut after one, and the leftover is not in the second language.
\[ L_1 = \{a, aa\},\ L_2 = \{a\}: \quad aaa = aa \cdot a \in L_1L_2, \quad a \cdot aa \notin L_1L_2 \]
The machine would have to try both cuts, and a deterministic machine cannot try two things. Determinizing afterwards is fine — Lesson 6 guarantees it — but the construction itself has to be nondeterministic.
Check
You concatenate a 3-state NFA with two accepting states onto a 4-state NFA.
Check your understanding
How many new ε-arrows does the concatenation construction add?
Answer: A
Why: The construction adds an ε-arrow from every accepting state of the first machine to the second machine's start state. With two accepting states that is two arrows, each representing a place the split could legitimately occur.
Section
Section 4
Concept
Star must accept any number of members of the language, joined end to end — including no members at all.
\[ L^{*} = \{\, x_1x_2\cdots x_k \;:\; k \ge 0,\ x_i \in L \,\} \]
Two obligations follow, and they pull in different directions. The machine must be able to return to the beginning after finishing a member, and it must accept the empty string even when the language does not contain it.
\[ \varepsilon \in L^{*} \quad \text{always} \]
A construction that satisfies the first and forgets the second is wrong; so is one that satisfies the second by making the wrong state accepting.
Hypothesis
Predict first
Why the naive construction is wrong is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Apply it to a concrete machine
Why: Take a machine accepting exactly the strings starting with a and ending with b — say, a start state reading a into a middle state that loops on both symbols and reads b into the accepting state.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
The obvious idea is: keep the machine as it is, add ε-arrows from every accepting state back to the start state, and make the start state accepting. Watch it fail.
Apply it to a concrete machine
Why: Take a machine accepting exactly the strings starting with a and ending with b — say, a start state reading a into a middle state that loops on both symbols and reads b into the accepting state.
Add the loop-back arrows and promote the start
Why: One ε-arrow from the accepting state back to the start, and the start state is now accepting so that the empty string is covered.
Find a string the modified machine accepts that it should not
Why: The star of that language contains only the empty string and strings that start with a. But the modified machine can enter the loop, come back to the start state mid-string, and the promoted start state is now reachable from inside — so runs can finish there after reading input that is not a full member.
Diagnose the cause
Why: Promoting the original start state is the error. It has incoming arrows, so a run can arrive there partway through and be declared accepting even though the current member is unfinished.
Verify that the fix removes the problem
Why: Add a fresh start state instead, make that one accepting, and give it an ε-arrow into the original start. Because nothing points into the fresh state, no run can arrive there mid-string, so the empty string is accepted and nothing else is wrongly added.
\[ \text{fresh } s: \quad \delta(s,\varepsilon) = \{q_0\}, \quad \text{no arrows into } s \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Why the naive construction is wrong", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Add a fresh start state instead, make that one accepting, and give it an ε-arrow into the original start. Because nothing points into the fresh state, no run can arrive there mid-string, so the empty string is accepted and nothing else is wrongly added.
Concept
Four ingredients, and each one is there to handle a specific case.
The fresh state is inert by design: it has no incoming arrows and no symbol-arrows, so it can only be occupied before any input is read.
\[ E\big(\{s\}\big) = \{s, q_0\} \ni s \]
Ranking
Put in order
Put the moves of Build a star machine and check the boundaries into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The ingredient reads a then b into an accepting state.
Worked example
Star the machine for the single string ab, and test the cases the construction was designed for.
Build the ingredient and add the four pieces
Why: The ingredient reads a then b into an accepting state. Add a fresh accepting start state, an ε-arrow into the original start, and an ε-arrow from the accepting state back to the original start.
Test zero copies
Why: The empty string leaves the run on the fresh start state, which is accepting. So the empty string is accepted, as star requires.
\[ \varepsilon \in L^{*} \]
Test one copy
Why: Read a then b. The run reaches the original accepting state, which stayed accepting, so ab is accepted.
Test two copies
Why: After the first ab, the loop-back ε-arrow returns the run to the original start, and the second ab is read the same way. So abab is accepted.
Verify the two boundary cases and one non-member
Why: The empty string and ab are accepted, and so is abab by looping. The string aba is rejected, since after the second a the run needs a b to finish the member and the input has ended. Zero copies, one copy, two copies, and a partial copy all behave correctly.
\[ \varepsilon,\ ab,\ abab \in L^{*} \qquad aba \notin L^{*} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Build a star machine and check the boundaries", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The empty string and ab are accepted, and so is abab by looping. The string aba is rejected, since after the second a the run needs a b to finish the member and the input has ended. Zero copies, one copy, two copies, and a partial copy all behave correctly.
Concept
A close relative comes up constantly, and the two differ in exactly one place.
\[ L^{+} = \bigcup_{k \ge 1} L^{k} = LL^{*} \]
The construction is identical except that the fresh start state is not accepting. Everything else — the entry arrow, the loop-back, the original accepting states — is unchanged.
\[ F_{+} = F \quad\text{(not } F \cup \{s\}\text{)} \]
So the two agree exactly when the language already contains the empty string, and differ by that one string otherwise.
\[ L^{*} = L^{+} \iff \varepsilon \in L \]
Estimation
Predict first
Take the star machine just built and turn it into a plus machine.
Commit before you compute: what does Build a plus machine from a star machine come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify that the empty string is now excluded and nothing else changed
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The empty string is rejected while ab and abab are still accepted, so exactly one string moved.
Worked example
Take the star machine just built and turn it into a plus machine.
Change exactly one thing
Why: Remove the fresh start state from the accepting set. Leave every arrow and every other accepting state alone.
Re-test zero copies
Why: The empty string now leaves the run on a non-accepting state, so it is rejected — which is what plus requires.
Re-test one and two copies
Why: Both still reach the original accepting state, which was never demoted. So ab and abab remain accepted.
Note the general lesson
Why: Keeping the fresh start state gave a clean place to encode a single yes-or-no decision. Had the original start been promoted instead, this edit would have been impossible without disturbing the loop.
Verify that the empty string is now excluded and nothing else changed
Why: The empty string is rejected while ab and abab are still accepted, so exactly one string moved. That is precisely the difference between star and plus for a language not containing the empty string.
\[ \varepsilon \in L^{+} \iff \varepsilon \in L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Build a plus machine from a star machine", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The empty string is rejected while ab and abab are still accepted, so exactly one string moved. That is precisely the difference between star and plus for a language not containing the empty string.
Intuition
Union needed a fresh start for convenience. Star needs one for correctness, and it is worth being clear about the difference.
Star is the only one of the three operations that must accept the empty string unconditionally. The only way to arrange that without also accepting unfinished members is to have a state which is accepting and unreachable from inside the machine.
A state with no incoming arrows is exactly that. It can only be occupied at time zero, so making it accepting adds the empty string and nothing else.
\[ \text{no arrows into } s \;\Rightarrow\; s \text{ occupied only before reading} \]
Commit first
Predict first
In the Kleene star construction, why is a fresh start state added instead of simply making the original start state accepting?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: The original start has incoming arrows, so runs could finish there mid-member
Why: The loop-back arrows point into the original start state, so a run can arrive there partway through the input. Promoting it would accept strings that stop in the middle of a member, which star does not contain.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Recall which state the construction promotes.
Check your understanding
In the Kleene star construction, why is a fresh start state added instead of simply making the original start state accepting?
Answer: A
Why: The loop-back arrows point into the original start state, so a run can arrive there partway through the input. Promoting it would accept strings that stop in the middle of a member, which star does not contain.
Pattern
All three regular operations are now wired the same way, and Lesson 8 will turn this table into a mechanical procedure.
| Operation | Fresh states | New arrows |
|---|---|---|
| union | one start | into each machine's start |
| concatenation | none | from the first machine's accepting states to the second's start |
| star | one accepting start | into the start, and from accepting states back to the start |
Two habits carry across all three: rename states so the machines are disjoint, and check the empty string last. Nearly every bug in these constructions shows up on the empty string.
Edge cases
Discussion prompt
The three wiring constructions side by side works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
All three regular operations are now wired the same way, and Lesson 8 will turn this table into a mechanical procedure.
Section
Section 5
Concept
Two more properties follow from constructions already proved, with no new machinery.
Complement: determinize, then swap the accepting set, as in Lessons 4 and 6. The determinization step is essential and is the only work.
\[ L(M') = \overline{L(M)} = \Sigma^{*} \setminus L(M) \]
Intersection: use the product machine and accept a pair when both components are accepting. Or derive it from union and complement by De Morgan, which needs no new construction at all.
\[ L_1 \cap L_2 = \overline{\overline{L_1} \cup \overline{L_2}} \]
Difference then follows too, since it is an intersection with a complement.
Concept
Reversing a language reverses every one of its strings.
\[ L^{R} = \{\, w^{R} : w \in L \,\} \]
The construction is exactly what it looks like: flip every arrow, make the old start state the only accepting state, and make the old accepting states the start states.
The result is nondeterministic in general — flipping can produce several arrows with one label out of a state, and several start states. Both are fine: Lesson 5 allows the first and permits the second as a variant, and Lesson 6 determinizes if needed.
Step zero
Discussion prompt
Reverse a machine and check it — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Reverse the arrows
Answer:
Worked example
Let a machine accept exactly the string 01: a start state, a 0-arrow to a middle state, and a 1-arrow to an accepting state. Build a machine for the reversal, which holds just the string 10.
Reverse the arrows
Why: The 0-arrow from the start to the middle becomes an arrow from the middle back to the start. The 1-arrow from the middle to the accepting state becomes one from the accepting state to the middle.
\[ b \xrightarrow{\,0\,} a, \qquad c \xrightarrow{\,1\,} b \]
Swap the roles of start and accepting
Why: The old accepting state becomes the start; the old start becomes the only accepting state. There was exactly one old accepting state, so no fresh start state is needed here.
Trace the string that should be accepted
Why: Start at the old accepting state, read 1 to reach the middle, read 0 to reach the old start — which is accepting. So 10 is accepted.
Note why the result is generally nondeterministic
Why: With several old accepting states, flipping would produce several start states; and two arrows into one state with the same label become two arrows out of it. Neither breaks anything.
Verify on the member and on its mirror image
Why: The string 10 is accepted as traced. The string 01 is rejected: from the new start there is no 0-arrow at all, so the run dies immediately. Exactly one of the two is accepted, which is what reversing a one-string language must do.
\[ 10 \in L^{R}, \qquad 01 \notin L^{R} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Reverse a machine and check it", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The string 10 is accepted as traced. The string 01 is rejected: from the new start there is no 0-arrow at all, so the run dies immediately. Exactly one of the two is accepted, which is what reversing a one-string language must do.
Concept
A homomorphism replaces each symbol by a fixed string, then extends to whole strings by concatenation.
\[ h : \Sigma \to \Gamma^{*}, \qquad h(a_1a_2\cdots a_n) = h(a_1)h(a_2)\cdots h(a_n) \]
Applied to a language, it maps every member and collects the images. The regular languages are closed under this too.
The construction is a substitution on the machine: replace each arrow labelled with a symbol by a chain of arrows spelling that symbol's image, inserting fresh intermediate states. An image that is the empty string becomes an ε-arrow.
\[ q \xrightarrow{\,a\,} p \quad\Longrightarrow\quad q \xrightarrow{\,h(a)\,} p \quad \text{(a chain of } |h(a)| \text{ arrows)} \]
Explain it
Discussion prompt
Explain Homomorphism: renaming symbols to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A homomorphism replaces each symbol by a fixed string, then extends to whole strings by concatenation.
Concept
The inverse direction asks which strings map into the language, and it is closed as well.
\[ h^{-1}(L) = \{\, w \in \Sigma^{*} : h(w) \in L \,\} \]
The construction is unusually neat: keep the machine's states exactly as they are, and relabel each arrow. On symbol a, the new machine moves as the old one would on the whole string that a maps to.
\[ \delta'(p,a) = \hat{\delta}\big(p, h(a)\big) \]
No states are added at all. This is the cheapest of all the closure constructions, and it is the one that makes several later reduction arguments short.
Analogy
Discussion prompt
Explain Inverse homomorphism by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The inverse direction asks which strings map into the language, and it is closed as well.
Fill the middle
Fill in the blanks
From Apply an inverse homomorphism — finish the line. Write what belongs on the right of the equals sign before you look.
h^*\{a,b\}^{}**(L) = ___
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Each a contributes the two-symbol block; each b contributes nothing at all.
Worked example
Let the target language be every repetition of the block 01, and let the homomorphism send a to that block and b to the empty string. Find the inverse image.
Work out what each symbol contributes
Why: Each a contributes the two-symbol block; each b contributes nothing at all.
\[ h(a) = 01, \qquad h(b) = \varepsilon \]
Apply the homomorphism to an arbitrary string
Why: A string of a's and b's maps to one copy of the block per a, in order, with the b's vanishing. So the image is always some number of blocks in a row.
\[ h(w) = (01)^{\,\#_a(w)} \]
Ask which images land in the target
Why: Every run of blocks is in the target language, including the empty one. So every string of a's and b's has its image inside, and the inverse image is everything.
\[ h^{-1}(L) = \{a,b\}^{*} \]
Note why the construction is finite
Why: The relabelling looks up where the old machine would go on a fixed string, which is a finite computation done once per arrow. No states are added, so the new machine is the same size as the old.
Verify on a mixed string
Why: Take abba. Its image is the block, nothing, nothing, the block — that is 0101, which is two blocks and therefore in the target. So abba belongs to the inverse image, consistent with the claim that everything does.
\[ h(abba) = 0101 \in L \;\Rightarrow\; abba \in h^{-1}(L) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Apply an inverse homomorphism", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Take abba. Its image is the block, nothing, nothing, the block — that is 0101, which is two blocks and therefore in the target. So abba belongs to the inverse image, consistent with the claim that everything does.
Prediction
Predict first
Which combination of closure properties settles it most directly?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Complement the second, then intersect with the first
Why: The difference of two sets is exactly the intersection of the first with the complement of the second. Both operations are closed, so the result is regular — and the product machine delivers it in one step by choosing the right accepting pairs.
Check
You know two languages are regular, and you want to show that the set of strings lying in the first but not in the second is regular.
Check your understanding
Which combination of closure properties settles it most directly?
Answer: A
Why: The difference of two sets is exactly the intersection of the first with the complement of the second. Both operations are closed, so the result is regular — and the product machine delivers it in one step by choosing the right accepting pairs.
Missing information
Discussion prompt
Show that the strings over the alphabet of zeros and ones which contain the block 001 and have even length form a regular language.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Two conditions joined by 'and': one about a substring, one about length. Each is a language in its own right.
Worked example
Show that the strings over the alphabet of zeros and ones which contain the block 001 and have even length form a regular language.
Split the description into conditions
Why: Two conditions joined by 'and': one about a substring, one about length. Each is a language in its own right.
\[ L = \{w : 001 \text{ occurs in } w\} \;\cap\; \{w : |w| \text{ even}\} \]
Show each piece is regular on its own
Why: The substring language has the four-state machine designed in Lesson 4. The even-length language has a two-state parity machine. Both are exhibited, so both are regular.
Name the operation joining them
Why: The word 'and' is intersection, and the regular languages are closed under it by the product construction.
Conclude, and note the size
Why: The product machine has eight states, and the language is regular. Building it directly by asking what must be remembered would have needed the same eight, discovered the hard way.
\[ 4 \cdot 2 = 8 \text{ states} \]
Verify that every piece was regular before combining
Why: Each component was backed by an actual machine, and each operation used was one the class is closed under. Both halves matter: a closure argument resting on a piece that is not regular proves nothing at all, as the next section shows.
\[ \text{regular pieces} \;+\; \text{closure operations} \;\Rightarrow\; \text{regular} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Assemble a regular language from pieces", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Each component was backed by an actual machine, and each operation used was one the class is closed under. Both halves matter: a closure argument resting on a piece that is not regular proves nothing at all, as the next section shows.
Section
Section 6
Concept
So far closure has been used to build regular languages. The same theorems, read in the other direction, prove that certain languages are not regular.
The idea is a proof by contradiction. Suppose the language in question were regular. Combine it, using a closed operation, with a language known to be regular. The result must then be regular too — so if it is known not to be, the assumption was wrong.
\[ L \cap R = N, \quad R \text{ regular}, \; N \text{ not regular} \;\Longrightarrow\; L \text{ not regular} \]
This turns one known non-regular language into a whole family of them, at very little cost per example.
Estimation
Predict first
Assume it is already known that the language of some number of zeros followed by the same number of ones is not regular — Lesson 11 proves this. Use that to show the following is not regular.
Commit before you compute: what does Prove a language non-regular using closure come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the intersection really is the known language, and finish
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Both inclusions check: any string of the intersection has equal counts and zeros first, so it has the required form; and any string of that form has equal counts and zeros first.
Worked example
Assume it is already known that the language of some number of zeros followed by the same number of ones is not regular — Lesson 11 proves this. Use that to show the following is not regular.
State the candidate
Why: Take the strings over the alphabet of zeros and ones with equally many zeros as ones, in any order.
\[ L = \{\, w \in \{0,1\}^{*} : \#_0(w) = \#_1(w) \,\} \]
Assume it is regular, for contradiction
Why: If it were, then every closure property applies to it, and in particular it may be intersected with any regular language.
Choose a regular language that isolates the known one
Why: Take all the zeros first, then all the ones — a language with an obvious three-state machine, hence regular.
\[ R = 0^{*}1^{*} \quad \text{regular} \]
Compute the intersection
Why: A string with equal counts whose zeros all precede its ones is exactly some number of zeros followed by the same number of ones. That is the known non-regular language.
\[ L \cap R = \{\, 0^{n}1^{n} : n \ge 0 \,\} \]
Verify the intersection really is the known language, and finish
Why: Both inclusions check: any string of the intersection has equal counts and zeros first, so it has the required form; and any string of that form has equal counts and zeros first. The intersection of two regular languages must be regular, but this one is not — contradiction, so the candidate is not regular.
\[ L \cap R = \{0^{n}1^{n}\} \text{ not regular} \;\Rightarrow\; L \text{ not regular} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove a language non-regular using closure", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Both inclusions check: any string of the intersection has equal counts and zeros first, so it has the required form; and any string of that form has equal counts and zeros first. The intersection of two regular languages must be regular, but this one is not — contradiction, so the candidate is not regular.
Ranking
Put in order
These are the steps of The closure-argument recipe, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Every proof of this shape has the same five steps. Writing them out keeps the argument honest.
Step four is where the work is, and where proofs go wrong. The result must equal the known language exactly, not merely resemble it.
\[ \text{closure argument} = \text{known non-regular } N \;+\; \text{regular } R \;+\; \text{an operation} \]
Real world
Discussion prompt
Outside this lesson: where does Regular Operations & Closure actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The closure-argument recipe is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
Lesson 7 establishes that the regular languages are closed under the three regular operations and more. Covers what closure means and why it is a theorem rather than a definition, union by both the product construction and the free-move construction, concatenation as a guessed split with the demotion trap, Kleene star with the empty-string subtlety and why a naive loop-back is wrong, plus intersection, complement, difference, reversal and homomorphism.
Step zero
Discussion prompt
Choose the right helper for a closure argument — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: State the candidate
Answer:
Worked example
A poorly chosen helper produces an intersection that proves nothing. Compare two attempts on the same candidate.
State the candidate
Why: Take the strings with strictly more zeros than ones.
\[ L = \{\, w \in \{0,1\}^{*} : \#_0(w) > \#_1(w) \,\} \]
Try a helper that is too permissive
Why: Intersecting with everything leaves the candidate unchanged. The result is not a language already known to be non-regular, so nothing follows.
\[ L \cap \Sigma^{*} = L \]
Try a helper that is too restrictive
Why: Intersecting with the single string 0 gives a one-string language, which is regular. A regular result is consistent with the candidate being regular, so again nothing follows.
Choose a helper that isolates a known shape
Why: Take zeros followed by ones. Now the intersection is the strings with more zeros than ones in that order, which is a known non-regular language of the same family.
\[ R = 0^{*}1^{*}, \qquad L \cap R = \{\, 0^{m}1^{n} : m > n \,\} \]
Verify both requirements on the helper chosen
Why: The helper must be regular, or the closure step is unavailable — and this one has a three-state machine. And the intersection must be known non-regular, or the contradiction never arrives — and this one is, by the same argument as the equal-counts language. Both conditions hold, so the proof goes through.
\[ R \text{ regular}, \quad L \cap R \text{ known non-regular} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Choose the right helper for a closure argument", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The helper must be regular, or the closure step is unavailable — and this one has a three-state machine. And the intersection must be known non-regular, or the contradiction never arrives — and this one is, by the same argument as the equal-counts language. Both conditions hold, so the proof goes through.
Intuition
The technique is powerful but derivative: it always needs a non-regular language to start from.
The very first non-regular language has to be proved from scratch, without appealing to any other. That is what the pumping lemma of Lesson 10 is for, and it is why that lesson cannot be skipped.
Once one example exists, closure arguments multiply it cheaply. In practice most non-regularity proofs are closure arguments resting on a single pumping-lemma proof done once.
\[ \text{one pumping proof} \;+\; \text{closure} \;\Rightarrow\; \text{many non-regular languages} \]
Counterexample
Discussion prompt
The technique is powerful but derivative: it always needs a non-regular language to start from.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The very first non-regular language has to be proved from scratch, without appealing to any other. That is what the pumping lemma of Lesson 10 is for, and it is why that lesson cannot be skipped.
Concept
Every property in the table is closed, but the constructions differ enormously in cost, and conflating the two leads to bad engineering decisions.
| Operation | States in the result |
|---|---|
| union by wiring | the sum, plus one |
| concatenation | the sum |
| star | one more than the original |
| intersection by product | the product |
| complement | unchanged, after determinizing |
| complement of an NFA | exponential in the worst case |
The last row is the one that bites. Complementing is free for a deterministic machine and potentially exponential for a nondeterministic one, because the determinization has to happen first. Closure says the result exists; it does not say it is small.
Comparison
Comparison matrix
From Closure does not mean the operation is cheap: refill the States in the result column from what you know. The rest of the table is as it appeared.
| Operation | States in the result |
|---|---|
| union by wiring | the sum, plus one |
| concatenation | the sum |
| star | one more than the original |
| intersection by product | the product |
| complement | unchanged, after determinizing |
| complement of an NFA | exponential in the worst case |
Intuition
The reason to tabulate these at all is that later classes have different tables, and the differences are diagnostic.
The context-free languages of Lesson 19 are closed under union, concatenation and star — but not under intersection or complement. That single difference is often the fastest way to prove a language is not context-free.
So the habit worth forming now: when meeting a new class, ask which operations it is closed under before asking anything else. The answer determines which proof techniques are available.
| Class | Union | Intersection | Complement |
|---|---|---|---|
| regular | yes | yes | yes |
| context-free | yes | no | no |
Trade off
Comparison matrix
From Closure properties are how classes are told apart: every row here is a choice with a cost. Fill the Union column, then say which row you would actually pick and what you give up for it.
| Class | Union | Intersection | Complement |
|---|---|---|---|
| regular | yes | yes | yes |
| context-free | yes | no | no |
Ranking
Put in order
Put the moves of Use a closure failure to identify a class into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. One holds strings with as many a's as b's followed by any number of c's; the other holds any number of a's followed by equally many b's as c's.
Worked example
Practise the diagnostic in advance, on the class this course reaches in Lesson 19.
Take two languages that will turn out to be context-free
Why: One holds strings with as many a's as b's followed by any number of c's; the other holds any number of a's followed by equally many b's as c's. Each needs one matched pair, which a stack can handle.
Intersect them
Why: A string in both must match a's against b's and b's against c's simultaneously, so all three counts agree.
\[ L_1 \cap L_2 = \{\, a^{n}b^{n}c^{n} : n \ge 0 \,\} \]
Note that the result is not context-free
Why: Lesson 18 proves this. Two independent matched pairs are one more than a single stack can track.
Draw the conclusion about the class
Why: Both inputs are in the class and the result is not, so the class is not closed under intersection. That is exactly the shape of the counterexample from Section 1.
Verify the argument against the regular case, to see the contrast
Why: The same move cannot work for the regular languages, because the product construction proves closure under intersection outright. So the failure genuinely distinguishes the two classes rather than reflecting a weakness in the argument.
\[ \text{regular: closed} \qquad \text{context-free: not closed} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Use a closure failure to identify a class", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The same move cannot work for the regular languages, because the product construction proves closure under intersection outright. So the failure genuinely distinguishes the two classes rather than reflecting a weakness in the argument.
Concept
Everything proved in this lesson, in one place. This table gets extended in Lesson 19 for a larger class, and the differences between the two tables are the point.
| Operation | Regular languages closed? | Construction |
|---|---|---|
| union | yes | product, or fresh start with free moves |
| concatenation | yes | wire accepting states to the second start |
| Kleene star | yes | fresh accepting start, loop back |
| intersection | yes | product, accept when both accept |
| complement | yes | determinize, then swap |
| difference | yes | intersect with a complement |
| reversal | yes | flip the arrows, swap start and accepting |
| homomorphism | yes | expand each arrow into a chain |
Every row says yes, which is unusual and is what makes the regular languages such a comfortable class to work in.
Comparison
Comparison matrix
From The closure table, collected: refill the Regular languages closed? column from what you know. The rest of the table is as it appeared.
| Operation | Regular languages closed? | Construction |
|---|---|---|
| union | yes | product, or fresh start with free moves |
| concatenation | yes | wire accepting states to the second start |
| Kleene star | yes | fresh accepting start, loop back |
| intersection | yes | product, accept when both accept |
| complement | yes | determinize, then swap |
| difference | yes | intersect with a complement |
| reversal | yes | flip the arrows, swap start and accepting |
| homomorphism | yes | expand each arrow into a chain |
Elimination
Eliminate the wrong options
In a closure argument proving a language L is not regular, what must be true of the helper language R?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The argument needs the closure step to be available, which requires the helper to be regular, and it needs the result to be a language already known not to be regular, which is what delivers the contradiction. Both conditions are essential.
Check
Think about what each ingredient of the argument has to supply.
Check your understanding
In a closure argument proving a language L is not regular, what must be true of the helper language R?
Answer: A
Why: The argument needs the closure step to be available, which requires the helper to be regular, and it needs the result to be a language already known not to be regular, which is what delivers the contradiction. Both conditions are essential.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — What Closure Means · Union, Two Ways · Concatenation: Guessing the Split · Kleene Star: Looping Safely · More Closure Properties · Closure as a Proof Technique. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can combine regular languages freely, and you can turn the same theorems around to prove languages irregular.
| Situation | Move |
|---|---|
| two languages, either one | product machine, or wire a fresh start |
| one language then another | wire accepting states to the second start, and demote them |
| any number of repetitions | fresh accepting start, loop back, never promote the old start |
| both conditions at once | product machine, accept when both accept |
| show a language is not regular | intersect with a simple pattern to expose a known one |
Lesson 8 turns the three wiring constructions into a notation — regular expressions — and Lesson 9 converts machines back into that notation, completing the circle.
Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.