Lesson 8 introduces the notation that describes exactly the languages finite automata recognize. It covers the six-rule inductive syntax and the semantic clauses that assign a language to every expression, then the precedence order that puts star before concatenation before union, and design idioms built around the starred alphabet. It gives the algebraic laws, including the identities, distribution, and the star laws, and Thompson's construction, which converts any expression into an epsilon-NFA of linear size. It ends with Kleene's theorem and what the equivalence buys you.
Subject: Theory of Computation · 122 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Theory of Computation · Lesson 8
Six rules of syntax, one inductive meaning, and a bridge back to automata. Everything a finite machine can recognize, written as a formula.
Objectives
Lessons 4 through 7 built machines. This lesson builds a notation that describes exactly the same languages, and proves it. By the end you can:
Warm-up
Discussion prompt
Before we open Regular Expressions: without looking back, what was the main idea of Regular Operations & Closure, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Lesson 7 establishes that the regular languages are closed under the three regular operations and more. Covers what closure means and why it is a theorem rather than a definition, union by both the product construction and the free-move construction, concatenation as a guessed split with the demotion trap, Kleene star with the empty-string subtlety and why a naive loop-back is wrong, plus intersection, complement, difference, reversal and homomorphism.
Section
Section 1
Concept
A regular expression is a finite string of symbols built by fixed rules. On its own it computes nothing — it is a name for a language.
regular expression — A formal expression built from the alphabet symbols and the constants for the empty set and the empty string, using union, concatenation and star.
Every expression denotes a language. Writing the expression and knowing which language it denotes are two separate steps, and this lesson keeps them apart on purpose.
\[ R \quad \longmapsto \quad L(R) \subseteq \Sigma^{*} \]
Counterexample
Discussion prompt
A regular expression is a finite string of symbols built by fixed rules. On its own it computes nothing — it is a name for a language.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Every expression denotes a language. Writing the expression and knowing which language it denotes are two separate steps, and this lesson keeps them apart on purpose.
Intuition
A DFA answers a question: hand it a string and it says yes or no. An expression answers a different question: it tells you the shape of every string in the language, all at once.
That difference is why both survive. You would not hand a search tool a five-state diagram, and you would not run an expression symbol by symbol on hardware — but each converts to the other.
| Automaton | Expression | |
|---|---|---|
| shape | a labelled graph | a formula |
| natural question | does this string belong? | what do the strings look like? |
| good for | running, deciding | writing, communicating, searching |
| built in | Lessons 4 to 6 | this lesson |
Comparison
Comparison matrix
From Machines recognize; expressions describe: refill the Automaton column from what you know. The rest of the table is as it appeared.
| Automaton | Expression | |
|---|---|---|
| shape | a labelled graph | a formula |
| natural question | does this string belong? | what do the strings look like? |
| good for | running, deciding | writing, communicating, searching |
| built in | Lessons 4 to 6 | this lesson |
Concept
The definition is inductive: three base cases you may write down freely, then three ways to combine what you already have.
| Expression | Denotes | Size of that language |
|---|---|---|
| a symbol of the alphabet | the one-symbol string | one string |
| the empty-string constant | the empty string alone | one string |
| the empty-set constant | no strings at all | zero strings |
The last two are constants, not symbols of the alphabet. They are part of the notation, and they are the two that beginners confuse.
\[ L(a) = \{a\}, \qquad L(\varepsilon) = \{\varepsilon\}, \qquad L(\varnothing) = \varnothing \]
Trade off
Comparison matrix
From The atoms: three ways to start: every row here is a choice with a cost. Fill the Denotes column, then say which row you would actually pick and what you give up for it.
| Expression | Denotes | Size of that language |
|---|---|---|
| a symbol of the alphabet | the one-symbol string | one string |
| the empty-string constant | the empty string alone | one string |
| the empty-set constant | no strings at all | zero strings |
Concept
Given two expressions you already trust, three constructions make a new one. They are the same three regular operations proved closed in Lesson 7 — which is the whole reason this notation can work.
| Written | Called | Meaning |
|---|---|---|
| R union S | union | a string matching either one |
| R S | concatenation | a string matching R, then one matching S |
| R star | Kleene star | zero or more strings, each matching R, joined |
\[ L(R \cup S) = L(R) \cup L(S), \qquad L(RS) = L(R)L(S), \qquad L(R^{*}) = L(R)^{*} \]
Nothing else is allowed. There is no complement operator, no intersection operator, no 'not this symbol' — those languages are still regular, but they are not written directly in this notation.
Analogy
Discussion prompt
Explain The operators: three ways to combine by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Given two expressions you already trust, three constructions make a new one. They are the same three regular operations proved closed in Lesson 7 — which is the whole reason this notation can work.
Concept
Here is the complete grammar. Anything that can be produced by finitely many applications of these six rules is a regular expression over the alphabet, and nothing else is.
The phrase finitely many applications matters. Every expression is a finite object with a finite parse tree, even when the language it denotes is infinite.
\[ |R| < \infty \quad \text{always}, \qquad |L(R)| \text{ may be infinite} \]
Explain it
Discussion prompt
Explain The full inductive definition, in one place to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The phrase finitely many applications matters. Every expression is a finite object with a finite parse tree, even when the language it denotes is infinite.
Ranking
Put in order
Put the moves of Build the parse tree of an expression into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Scan the expression at bracket depth zero.
Worked example
Take the expression that starts with any number of a's and b's and then ends with abb. Find its structure.
\[ (a \cup b)^{*}abb \]
Find the outermost operator first
Why: Scan the expression at bracket depth zero. There is no union at depth zero — the union sits inside the parentheses — so the top-level operator is concatenation.
Split at the top-level concatenation
Why: The expression is four things joined in a row: the starred group, then a, then b, then b. Concatenation is associative, so the grouping among those four does not matter.
\[ \underbrace{(a \cup b)^{*}}_{\text{factor 1}} \cdot a \cdot b \cdot b \]
Descend into the starred factor
Why: Its operator is star, applied to whatever is inside the parentheses. So the star node has exactly one child.
Descend once more, into the parentheses
Why: Inside is a union of two atoms. Both children are alphabet symbols, so this branch of the tree is finished.
\[ \text{star} \to \text{union} \to \{a, b\} \]
Verify by rebuilding the expression from the tree
Why: Read the tree back out: union of a and b, starred, then concatenated with a, b, b. That reproduces the original string of symbols exactly, so no operator was misplaced and no parenthesis was dropped.
\[ \text{tree} \;\longrightarrow\; (a \cup b)^{*}abb \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Build the parse tree of an expression", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Read the tree back out: union of a and b, starred, then concatenated with a, b, b. That reproduces the original string of symbols exactly, so no operator was misplaced and no parenthesis was dropped.
Ranking
Put in order
These are the steps of How to read any regular expression aloud, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Reading an expression correctly is a mechanical procedure. Do it in this order every time.
Doing this out loud catches precedence mistakes before they reach the page, because a misread expression almost always sounds wrong.
Elimination
Eliminate the wrong options
Over the alphabet of a and b, which language does the expression ab* denote?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Star binds more tightly than concatenation, so the star applies to b alone, not to ab. The expression is therefore a single a concatenated with zero or more b's, giving a, ab, abb, abbb, and so on.
Check
Read the expression carefully before choosing. Precedence decides this one.
Check your understanding
Over the alphabet of a and b, which language does the expression ab* denote?
Answer: A
Why: Star binds more tightly than concatenation, so the star applies to b alone, not to ab. The expression is therefore a single a concatenated with zero or more b's, giving a, ab, abb, abbb, and so on.
Intuition
There is a useful way to picture the difference. An expression is a set of instructions for producing members of the language, and a machine is a procedure for testing a candidate.
Reading an expression left to right, each star says 'loop as many times as you like here' and each union says 'take either road'. Every set of choices you can make produces one string of the language, and every string arises from at least one set of choices.
That is why an expression makes the shape of a language obvious while a diagram makes membership easy to decide. They are the same information organized for two different jobs.
\[ \text{choices in } R \;\longleftrightarrow\; \text{strings of } L(R) \]
Anomaly
Predict first
A student writes this, and it looks reasonable:
Are the two constants of the notation interchangeable? Simplify the expression that unions the empty-set constant onto a.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The empty set is nothing and the empty string is nothing visible, so treat them the same and let either one contribute an option that matches no input.
Are the two constants of the notation interchangeable? Simplify the expression that unions the empty-set constant onto a.
Why: The empty set is nothing and the empty string is nothing visible, so treat them the same and let either one contribute an option that matches no input.
Trap
Are the two constants of the notation interchangeable? Simplify the expression that unions the empty-set constant onto a.
Treat both constants as 'nothing'
Why: The empty set is nothing and the empty string is nothing visible, so treat them the same and let either one contribute an option that matches no input.
\[ \varnothing \overset{?}{=} \varepsilon \]
Conclude the union adds a possibility
Why: If the constant behaves like the empty string, then the union offers a second choice: match a, or match nothing. So the language should hold two strings.
\[ L(a \cup \varnothing) \overset{?}{=} \{a, \varepsilon\} \]
Are the two constants of the notation interchangeable? Simplify the expression that unions the empty-set constant onto a.
Separate the two: one is an empty language, the other a one-string language
Why: The empty-set constant denotes a language with no members at all. The empty-string constant denotes a language with exactly one member, the string of length zero. Their languages have different sizes, so they cannot be the same expression.
\[ |L(\varnothing)| = 0 \qquad \text{but} \qquad |L(\varepsilon)| = 1 \]
Apply the union
Why: Unioning a language with the empty language adds no members, so the result is just the language of a. The empty-set constant is the identity for union, and the empty-string constant is the identity for concatenation — two different roles.
\[ L(a \cup \varnothing) = \{a\}, \qquad L(a\varepsilon) = \{a\} \]
Notation
Annotate
From Trap: the empty-set constant and the empty-string constant — read this one piece at a time. What is each part doing?
On: \( \varnothing \overset{?}{=} \varepsilon \)
Section
Section 2
Concept
The syntax was defined inductively, so the meaning must be too. There is one semantic clause for each of the six syntax rules, and each clause is stated only in terms of smaller expressions.
This is the same move as the extended transition function in Lesson 4: define the easy case outright, then define the hard case in terms of a case that is one step easier.
| Expression | Its language |
|---|---|
| a | the set holding the one-symbol string a |
| the empty-string constant | the set holding the empty string |
| the empty-set constant | the empty set |
| R union S | the union of the two languages |
| R S | the concatenation of the two languages |
| R starred | the Kleene star of the language |
Concept
Look at the last three rows. Nothing new is being invented: the meaning of the notation is spelled out entirely in the language operations already proved closed.
\[ L(R \cup S) = L(R) \cup L(S) \]
\[ L(RS) = L(R)L(S) \]
\[ L(R^{*}) = \bigcup_{k \ge 0} L(R)^{k} \]
That is exactly why every expression denotes a regular language: each construction step stays inside the class, by the closure theorems of Lesson 7, and there are only finitely many steps.
Step zero
Discussion prompt
Compute the language of an expression from the clauses — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Start at the leaves
Answer:
Worked example
Work out the language of the expression below, without guessing.
\[ (0 \cup 1)0^{*} \]
Start at the leaves
Why: The three atoms denote one-string languages. Nothing has been combined yet, so these are simply read off the base clauses.
\[ L(0) = \{0\}, \quad L(1) = \{1\}, \quad L(0) = \{0\} \]
Apply the union clause
Why: The parenthesized subexpression is a union, so its language is the union of the two one-string languages — a two-string language.
\[ L(0 \cup 1) = \{0\} \cup \{1\} = \{0, 1\} \]
Apply the star clause
Why: Starring the one-string language of 0 gives every finite run of zeros, including the run of length zero.
\[ L(0^{*}) = \{\varepsilon, 0, 00, 000, \dots\} \]
Apply the concatenation clause
Why: Every member of the result is one string from the left language followed by one from the right. So the strings are exactly: one symbol, either 0 or 1, then any number of zeros.
\[ L\big((0 \cup 1)0^{*}\big) = \{\, x0^{k} \;:\; x \in \{0,1\},\ k \ge 0 \,\} \]
Verify by testing a member and a non-member
Why: The string 1000 splits as the symbol 1 followed by three zeros, so it belongs. The string 010 would need the first symbol 0 followed by zeros only, but a 1 appears later, so it does not belong. Both verdicts match the description just derived.
\[ 1000 \in L, \qquad 010 \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Compute the language of an expression from the clauses", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The string 1000 splits as the symbol 1 followed by three zeros, so it belongs. The string 010 would need the first symbol 0 followed by zeros only, but a 1 appears later, so it does not belong. Both verdicts match the description just derived.
Estimation
Predict first
The boundary cases are where the clauses earn their keep. Compute both of these carefully.
Commit before you compute: what does Two expressions that look empty but are not come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify by counting members on each side
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The first result has exactly one member and the second has none, so they are certainly different expressions.
Worked example
The boundary cases are where the clauses earn their keep. Compute both of these carefully.
Start with the empty-set constant starred
Why: The star clause takes the union over all powers, and the zeroth power of any language is the language holding just the empty string. That clause never consults a member of the language, so it survives even when there are none.
\[ L(\varnothing^{*}) = \bigcup_{k \ge 0} \varnothing^{k} = \varnothing^{0} = \{\varepsilon\} \]
Now the empty-set constant concatenated with anything
Why: Concatenation needs one member from each side. There is no member on the left, so no string can be formed at all, no matter how rich the right-hand language is.
\[ L(\varnothing R) = \varnothing L(R) = \varnothing \]
Contrast with the empty-string constant concatenated with anything
Why: Here the left side does have a member — the empty string — and gluing it onto any string changes nothing. So this concatenation is the identity.
\[ L(\varepsilon R) = \{\varepsilon\}L(R) = L(R) \]
Verify by counting members on each side
Why: The first result has exactly one member and the second has none, so they are certainly different expressions. The third has the same members as R itself. All three counts follow from the clauses alone, with no appeal to intuition.
\[ |L(\varnothing^{*})| = 1, \quad |L(\varnothing R)| = 0, \quad L(\varepsilon R) = L(R) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Two expressions that look empty but are not", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The first result has exactly one member and the second has none, so they are certainly different expressions. The third has the same members as R itself. All three counts follow from the clauses alone, with no appeal to intuition.
Intuition
It is tempting to define the meaning by saying 'the language is whatever strings match'. That is circular: matching is what we are trying to define.
Induction on structure escapes the circle. Each clause explains a compound expression in terms of strictly smaller ones, and the smallest expressions are explained outright. Since every expression is finite, the unwinding always terminates.
This also hands you a proof technique for free: to prove something about every regular expression, prove it for the three atoms and show each of the three operators preserves it. That is exactly the shape of the proofs later in this lesson.
Concept
Writing every parenthesis would be unreadable, so the notation fixes a binding order, exactly as arithmetic does for powers, products and sums.
| Operator | Binds | Arithmetic analogue |
|---|---|---|
| star | tightest | exponent |
| concatenation | middle | multiplication |
| union | loosest | addition |
Union and concatenation are both associative, so a chain of them needs no internal grouping. Parentheses are only ever needed to override the order above.
\[ ab^{*} \cup c \quad \equiv \quad \big(a(b^{*})\big) \cup c \]
Pattern
Step through it
Step through Precedence: star, then concatenation, then union one row at a time. What is driving the change, and what would the row after the last one be?
Missing information
Discussion prompt
Insert every parenthesis the precedence rules imply, so that the structure is explicit.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Star is tightest, and it takes only the single symbol immediately to its left. So the star attaches to b alone.
Worked example
Insert every parenthesis the precedence rules imply, so that the structure is explicit.
\[ ab^{*} \cup cd \]
Bind the stars first
Why: Star is tightest, and it takes only the single symbol immediately to its left. So the star attaches to b alone.
\[ a(b^{*}) \cup cd \]
Bind the concatenations next
Why: Group each maximal run of adjacent factors. On the left that is a with the starred b; on the right it is c with d.
\[ \big(a(b^{*})\big) \cup (cd) \]
Bind the union last
Why: One union remains, and it now joins two fully grouped operands. The expression is completely parenthesized.
\[ \Big(\big(a(b^{*})\big)\Big) \cup \big((cd)\big) \]
Verify by testing a string against both readings
Why: Take the string abb. Under the parenthesization just derived it matches the left branch: one a, then two b's. Under the wrong reading, in which the star covered ab, the string would have to be a repetition of the block ab and would fail. The derived reading accepts it, which is the intended meaning.
\[ abb \in L\big(ab^{*} \cup cd\big) \qquad abb \notin L\big((ab)^{*}\big) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Parenthesize an expression completely", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Take the string abb. Under the parenthesization just derived it matches the left branch: one a, then two b's. Under the wrong reading, in which the star covered ab, the string would have to be a repetition of the block ab and would fail. The derived reading accepts it, which is the intended meaning.
Prediction
Predict first
Which expression is equivalent to a union b star, written without parentheses as a ∪ b*?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: a unioned with (b starred)
Why: Star binds tighter than union, so the star applies to b alone. The expression denotes the language holding the single string a together with every finite run of b's, including the empty string.
Check
No parentheses were written. Apply the binding order.
Check your understanding
Which expression is equivalent to a union b star, written without parentheses as a ∪ b*?
Answer: A
Why: Star binds tighter than union, so the star applies to b alone. The expression denotes the language holding the single string a together with every finite run of b's, including the empty string.
Pattern
When an expression is dense, resolve it in this fixed order rather than reading left to right.
If the reading you get differs from the one you intended, the fix is always the same: add parentheses. Relying on a reader to guess the intended grouping is how expression bugs get shipped.
Section
Section 3
Constraint
Discussion prompt
Run The design recipe for regular expressions with this step confiscated:
Write the fixed part of the pattern first — the symbols that must literally appear.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Designing an expression is the mirror of designing a machine in Lesson 4. Instead of asking what to remember, ask what the strings look like.
Step five is not optional. Almost every wrong expression is wrong only on the empty string or on a single-symbol string.
\[ \Sigma^{*} \;=\; \text{the 'anything' idiom} \;=\; (a \cup b \cup \cdots)^{*} \]
Edge cases
Discussion prompt
The design recipe for regular expressions works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Designing an expression is the mirror of designing a machine in Lesson 4. Instead of asking what to remember, ask what the strings look like.
Intuition
Almost every practical expression is a fixed skeleton with the starred alphabet poured into the gaps.
| Requirement | Shape of the expression |
|---|---|
| begins with the block w | w, then anything |
| ends with the block w | anything, then w |
| contains the block w | anything, then w, then anything |
| is exactly w | w alone, with no anything at all |
Notice how the placement of the anything idiom does all the work. Getting a language wrong is usually a matter of putting it on the wrong side, or forgetting it entirely.
\[ \Sigma^{*}w\Sigma^{*} \quad \text{versus} \quad \Sigma^{*}w \quad \text{versus} \quad w\Sigma^{*} \]
Step zero
Discussion prompt
Design: strings that end in 01 — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Name the fixed part
Answer:
Worked example
Over the alphabet of zeros and ones, describe every string whose last two symbols are 0 then 1.
Name the fixed part
Why: The requirement pins down the final two symbols and says nothing whatever about what comes before them.
\[ \text{fixed suffix} = 01 \]
Put the anything idiom where the freedom is
Why: Everything before the suffix is unconstrained, including being absent altogether. The starred alphabet covers exactly that, since it contains the empty string.
\[ (0 \cup 1)^{*}01 \]
Check the shortest member
Why: The string 01 itself must qualify. It does, by taking zero repetitions in the star and then the literal suffix.
\[ 01 = \varepsilon \cdot 01 \in L \]
Check a string that must be excluded
Why: The string 010 ends in 1 then 0, not 0 then 1. No split can put the required suffix at the end, so it is correctly outside.
\[ 010 \notin L \]
Verify against the Lesson 4 machine for the same language
Why: The three-state DFA built in Lesson 4 accepts exactly the strings ending in 01. Running 01, 1101 and 010 through both descriptions gives accept, accept, reject in each case — the expression and the machine agree on all three.
\[ 01,\ 1101 \in L \qquad 010,\ \varepsilon \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Design: strings that end in 01", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The string 01 itself must qualify. It does, by taking zero repetitions in the star and then the literal suffix.
Worked example
Now the block may appear anywhere at all, and the constraint is existence rather than position.
Recognize the shape
Why: This is a substring requirement, so the block needs freedom on both sides — before it and after it.
\[ \Sigma^{*} \; 001 \; \Sigma^{*} \]
Write it over the concrete alphabet
Why: Expand the anything idiom into the starred union of the two symbols, since the notation has no shorthand for a whole alphabet.
\[ (0 \cup 1)^{*}001(0 \cup 1)^{*} \]
Notice what the expression does not claim
Why: It does not say the block appears exactly once. If a string contains two occurrences, several different splits witness membership — and membership only needs one.
Verify on an edge case and a near miss
Why: The string 001 itself belongs, taking both stars empty. The string 0001 belongs too, with a single leading zero absorbed on the left. The string 0101 has no 001 anywhere, and no split can produce one, so it is correctly excluded.
\[ 001,\ 0001 \in L \qquad 0101 \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Design: strings containing 001 somewhere", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The string 001 itself belongs, taking both stars empty. The string 0001 belongs too, with a single leading zero absorbed on the left. The string 0101 has no 001 anywhere, and no split can produce one, so it is correctly excluded.
Estimation
Predict first
A counting condition, not a positional one. The trick is to write down the repeating unit.
Commit before you compute: what does Design: an even number of a's come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify on the boundary and on both parities
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The empty string has zero a's, which is even, and it matches by taking every star empty.
Worked example
A counting condition, not a positional one. The trick is to write down the repeating unit.
Find the block that keeps the count even
Why: Adding a's two at a time preserves evenness. Between and around those a's, any number of b's may appear without affecting the count.
\[ b^{*}ab^{*}ab^{*} \]
Allow that block to repeat any number of times
Why: Zero repetitions must be allowed, since a string of b's alone already has an even count of a's — namely zero.
\[ b^{*}\big(ab^{*}ab^{*}\big)^{*} \]
Read the result back
Why: A run of b's, then any number of a-pairs each padded with b's. Every a is matched with a partner, so the total is always even.
Verify on the boundary and on both parities
Why: The empty string has zero a's, which is even, and it matches by taking every star empty. The string aa matches with one repetition. The string a has one a, an odd count, and no repetition of a paired block can produce a single a, so it is correctly rejected.
\[ \varepsilon,\ aa,\ bab \in L \qquad a,\ aab \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Design: an even number of a's", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The empty string has zero a's, which is even, and it matches by taking every star empty. The string aa matches with one repetition. The string a has one a, an odd count, and no repetition of a paired block can produce a single a, so it is correctly rejected.
Ranking
Put in order
Put the moves of Design: the third symbol from the end is a 1 into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Counting from the right, position three must hold a 1, and positions two and one may hold anything.
Worked example
This is the language whose DFA needed exponentially many states in Lesson 6. As an expression it is almost trivial.
Anchor the constrained position
Why: Counting from the right, position three must hold a 1, and positions two and one may hold anything.
\[ \dots 1 \, \Sigma \, \Sigma \]
Fill in the freedom before the anchor
Why: Everything to the left of the anchored 1 is unconstrained, so the anything idiom goes there and nowhere else.
\[ (0 \cup 1)^{*}\,1\,(0 \cup 1)(0 \cup 1) \]
Note the size contrast
Why: The expression grows linearly as the position moves further from the end, while the smallest DFA doubles. Same languages, wildly different notation costs — which is one honest reason to keep both formalisms.
\[ \text{expression: } O(k) \qquad \text{smallest DFA: } 2^{k} \]
Verify on a member and a non-member of the same length
Why: In 100 the third symbol from the end is the leading 1, so it matches with the star empty. In 011 the third from the end is a 0, and no split can place the literal 1 at that position, so it is excluded.
\[ 100 \in L \qquad 011 \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Design: the third symbol from the end is a 1", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: In 100 the third symbol from the end is the leading 1, so it matches with the star empty. In 011 the third from the end is a 0, and no split can place the literal 1 at that position, so it is excluded.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Write an expression for the strings consisting of some a's followed by the same number of b's.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The strings are a run of a's followed by a run of b's, so star the a and star the b and put them side by side.
Write an expression for the strings consisting of some a's followed by the same number of b's.
Why: The strings are a run of a's followed by a run of b's, so star the a and star the b and put them side by side.
Trap
Write an expression for the strings consisting of some a's followed by the same number of b's.
Star each half of the pattern
Why: The strings are a run of a's followed by a run of b's, so star the a and star the b and put them side by side.
\[ a^{*}b^{*} \]
Claim the counts are tied
Why: Both halves repeat, so surely they repeat together and the counts stay equal.
\[ L(a^{*}b^{*}) \overset{?}{=} \{\, a^{n}b^{n} \;:\; n \ge 0 \,\} \]
Write an expression for the strings consisting of some a's followed by the same number of b's.
Notice the two stars are completely independent
Why: Each star chooses its own repetition count. Nothing in the notation lets one star see what the other did — that is the entire expressive limit of this formalism.
\[ L(a^{*}b^{*}) = \{\, a^{m}b^{n} \;:\; m, n \ge 0 \,\} \]
Exhibit a member that breaks the claim
Why: The string aab has two a's and one b, and it plainly matches the expression. So the expression denotes strictly more than the equal-count language, and the two are not the same.
\[ aab \in L(a^{*}b^{*}) \quad \text{but} \quad aab \notin \{a^{n}b^{n}\} \]
Draw the real conclusion
Why: No regular expression denotes the equal-count language at all. Lesson 11 proves it, and this trap is the first hint: the notation has no way to tie two independent repetitions together.
\[ \{\, a^{n}b^{n} \,\} \text{ is not regular} \]
Notation
Annotate
From *Trap: ab* is not 'equal numbers of a's and b's'** — read this one piece at a time. What is each part doing?
On: \( \{\, a^{n}b^{n} \,\} \text{ is not regular} \)
Commit first
Predict first
Over the alphabet of a and b, which expression denotes the strings that both begin and end with a?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: a(a ∪ b)*a ∪ a
Why: The main pattern handles every string of length two or more: a literal a, anything in the middle, a literal a. But the single string a also begins and ends with a, and the main pattern cannot produce it because it demands two separate a's. Unioning the one-symbol case on covers it.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Watch the boundary case before you answer.
Check your understanding
Over the alphabet of a and b, which expression denotes the strings that both begin and end with a?
Answer: A
Why: The main pattern handles every string of length two or more: a literal a, anything in the middle, a literal a. But the single string a also begins and ends with a, and the main pattern cannot produce it because it demands two separate a's. Unioning the one-symbol case on covers it.
Concept
Real tools add abbreviations. None of them add power — each one expands into the six core rules.
| Shorthand | Means | Expands to |
|---|---|---|
| R plus | one or more copies | R concatenated with R starred |
| R question mark | optional | R unioned with the empty-string constant |
| R to the k | exactly k copies | R written out k times |
| the alphabet symbol | any one symbol | the union of every symbol |
\[ R^{+} = RR^{*}, \qquad R^{?} = R \cup \varepsilon, \qquad R^{k} = \underbrace{RR\cdots R}_{k} \]
Keep the distinction sharp when reading proofs: a shorthand is a convenience of writing, whereas a genuinely new operator would demand a new closure theorem before it could be used at all.
Comparison
Comparison matrix
From Shorthands worth knowing: refill the Means column from what you know. The rest of the table is as it appeared.
| Shorthand | Means | Expands to |
|---|---|---|
| R plus | one or more copies | R concatenated with R starred |
| R question mark | optional | R unioned with the empty-string constant |
| R to the k | exactly k copies | R written out k times |
| the alphabet symbol | any one symbol | the union of every symbol |
Missing information
Discussion prompt
A modular condition on length, with no constraint on the symbols themselves.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Three arbitrary symbols in a row form the block whose repetition keeps the length a multiple of three.
Worked example
A modular condition on length, with no constraint on the symbols themselves.
Write the repeating unit
Why: Three arbitrary symbols in a row form the block whose repetition keeps the length a multiple of three.
\[ \Sigma\Sigma\Sigma \;=\; (0 \cup 1)(0 \cup 1)(0 \cup 1) \]
Star the block
Why: Zero repetitions gives the empty string, whose length is zero — divisible by three, so it must be included, and the star includes it automatically.
\[ \big((0 \cup 1)(0 \cup 1)(0 \cup 1)\big)^{*} \]
Compare with the machine
Why: A DFA for this language needs three states, one per remainder. The expression needs no notion of remainder at all — it simply refuses to stop except at a multiple of three.
Verify on three lengths
Why: Length 0 matches with zero repetitions and length 3 with one. Length 4 cannot match, because every repetition contributes exactly three symbols and no sum of threes equals four.
\[ |w| \in \{0, 3, 6\} \Rightarrow \text{match} \qquad |w| = 4 \Rightarrow \text{no match} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Design: length divisible by three", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Length 0 matches with zero repetitions and length 3 with one. Length 4 cannot match, because every repetition contributes exactly three symbols and no sum of threes equals four.
Section
Section 4
Concept
Different strings of symbols can name the same language. When they do, the expressions are called equivalent, and the whole algebra rests on this one definition.
equivalent expressions — Two regular expressions that denote exactly the same language, regardless of how differently they are written.
\[ R \equiv S \quad \overset{\text{def}}{\iff} \quad L(R) = L(S) \]
Equivalence is about the languages, never about the shape of the notation. So proving two expressions equal always means proving two sets equal — usually by showing each contains the other.
Concept
The two constants play the roles that zero and one play in arithmetic, and getting them straight removes most simplification errors.
| Law | Statement | Arithmetic analogue |
|---|---|---|
| union identity | R unioned with the empty set is R | adding zero |
| concatenation identity | R glued to the empty string is R | multiplying by one |
| annihilator | R glued to the empty set is the empty set | multiplying by zero |
| star of nothing | the empty set starred is the empty string | no analogue |
\[ R \cup \varnothing = R, \qquad R\varepsilon = \varepsilon R = R, \qquad R\varnothing = \varnothing R = \varnothing, \qquad \varnothing^{*} = \varepsilon \]
The last one has no arithmetic counterpart and is the one most often written down wrong. Starring the empty language does not give the empty language.
Concept
The structural laws say which parentheses may be dropped and which may not.
\[ R(S \cup T) = RS \cup RT, \qquad (S \cup T)R = SR \cup TR, \qquad R \cup R = R \]
Idempotence has no arithmetic analogue either — adding a number to itself does change it. It holds here because the union of a set with itself is that same set.
Step zero
Discussion prompt
Simplify an expression with the laws — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Distribute the concatenation over the union
Answer:
Worked example
Simplify the expression below to something obviously smaller.
\[ (a \cup \varepsilon)a^{*} \]
Distribute the concatenation over the union
Why: Concatenation distributes over union on the right, so split the product into two products.
\[ (a \cup \varepsilon)a^{*} = aa^{*} \cup \varepsilon a^{*} \]
Apply the concatenation identity
Why: The empty-string constant glued to anything leaves it unchanged, so the second term collapses.
\[ aa^{*} \cup \varepsilon a^{*} = aa^{*} \cup a^{*} \]
Recognize the containment
Why: Every string matching the first term is a nonempty run of a's, and every such run also matches the second term. So the first term contributes nothing new and the union absorbs it.
\[ L(aa^{*}) \subseteq L(a^{*}) \;\Rightarrow\; aa^{*} \cup a^{*} = a^{*} \]
Verify by comparing the two languages directly
Why: The original expression offers a choice: an a followed by any run of a's, or nothing followed by any run of a's. Either way the result is some run of a's, of any length including zero. That is precisely the language of the simplified expression.
\[ L\big((a \cup \varepsilon)a^{*}\big) = \{a^{k} : k \ge 0\} = L(a^{*}) \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Simplify an expression with the laws", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The original expression offers a choice: an a followed by any run of a's, or nothing followed by any run of a's. Either way the result is some run of a's, of any length including zero. That is precisely the language of the simplified expression.
Concept
Star has its own small set of identities, and they are the ones that make simplification of loops possible.
\[ (R^{*})^{*} = R^{*}, \qquad \varepsilon^{*} = \varepsilon, \qquad R^{*}R^{*} = R^{*} \]
\[ R^{*} = \varepsilon \cup RR^{*}, \qquad (R \cup S)^{*} = (R^{*}S^{*})^{*} \]
The fourth is the unrolling law: a run of copies is either empty, or one copy followed by a run. It is the identity that turns a star into a recursion, and Lesson 9 leans on it heavily.
Hypothesis
Predict first
Prove that starring twice adds nothing is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Prove the easy inclusion first
Why: Any language is contained in its own star, by taking exactly one copy. Applying that with the starred language in the role of the language gives one direction immediately.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Prove the first of the star laws properly, by two inclusions.
\[ (R^{*})^{*} = R^{*} \]
Prove the easy inclusion first
Why: Any language is contained in its own star, by taking exactly one copy. Applying that with the starred language in the role of the language gives one direction immediately.
\[ R^{*} \subseteq (R^{*})^{*} \]
Take an arbitrary member of the left side
Why: A string of the doubly starred language is a concatenation of finitely many pieces, each of which is itself a concatenation of finitely many copies of members of the inner language.
\[ w = u_1u_2\cdots u_m, \quad \text{each } u_i \in L(R)^{*} \]
Flatten the two levels into one
Why: Replace each piece by its own decomposition. What remains is a single concatenation of finitely many members of the inner language — because a finite sum of finite numbers is finite.
\[ w = \underbrace{x_{1}\cdots x_{k}}_{\text{all in } L(R)} \;\Rightarrow\; w \in L(R)^{*} \]
Verify that the empty string is handled on both sides
Why: The empty string lies in every star, including both of these, so the boundary case does not break either inclusion. With both inclusions established and the boundary checked, the two languages are equal.
\[ \varepsilon \in R^{*} \text{ and } \varepsilon \in (R^{*})^{*} \;\Rightarrow\; (R^{*})^{*} = R^{*} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove that starring twice adds nothing", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The empty string lies in every star, including both of these, so the boundary case does not break either inclusion. With both inclusions established and the boundary checked, the two languages are equal.
Intuition
The algebra looks like ordinary algebra, which makes the two places it differs genuinely dangerous.
| Instinct from arithmetic | Holds here? | Why |
|---|---|---|
| addition commutes | yes, union commutes | sets have no order |
| multiplication commutes | no | concatenation has an order |
| multiplication distributes | yes | splitting a choice is safe |
| adding a thing to itself doubles it | no, union is idempotent | sets absorb duplicates |
The reliable habit: whenever a proposed law feels obvious, test it on a two-symbol alphabet with the shortest strings you can find. One counterexample settles it, and finding one takes seconds.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Simplify the expression that concatenates a with b, unioned with the concatenation of b with a.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Multiplication commutes in arithmetic, so treat ab and ba as two spellings of one product and collapse the union by idempotence.
Simplify the expression that concatenates a with b, unioned with the concatenation of b with a.
Why: Multiplication commutes in arithmetic, so treat ab and ba as two spellings of one product and collapse the union by idempotence.
Trap
Simplify the expression that concatenates a with b, unioned with the concatenation of b with a.
Treat the two products as the same thing
Why: Multiplication commutes in arithmetic, so treat ab and ba as two spellings of one product and collapse the union by idempotence.
\[ ab \cup ba \overset{?}{=} ab \cup ab = ab \]
Report a one-string language
Why: With the union collapsed, the expression appears to denote a single string.
\[ |L(ab \cup ba)| \overset{?}{=} 1 \]
Simplify the expression that concatenates a with b, unioned with the concatenation of b with a.
Check whether the two products denote the same language
Why: Concatenation glues strings in a fixed order, so the first product denotes the single string a-then-b and the second denotes b-then-a. Those are different strings of length two.
\[ L(ab) = \{ab\}, \qquad L(ba) = \{ba\}, \qquad ab \neq ba \]
Apply the union honestly
Why: The union of two distinct one-string languages has two members, and no law removes either. The expression is already in simplest form.
\[ L(ab \cup ba) = \{ab, ba\}, \qquad |L(ab \cup ba)| = 2 \]
Keep the surviving rule
Why: Union commutes, so the two branches may be written in either order. Concatenation does not, so the symbols inside a branch may never be reordered.
\[ ab \cup ba = ba \cup ab \quad \text{but} \quad ab \neq ba \]
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Prediction
Predict first
Which of the following is NOT a valid law of regular expressions?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: (RS)* equals RS
Why: Take R to be a and S to be b. Then (ab)* contains ab and abab but not aab, while ab contains aab. A single string in one side and not the other refutes the law.
Check
Three of these hold for all regular expressions. One does not.
Check your understanding
Which of the following is NOT a valid law of regular expressions?
Answer: B
Why: Take R to be a and S to be b. Then (ab)* contains ab and abab but not aab, while ab contains aab. A single string in one side and not the other refutes the law.
Pattern
There is no shortcut through the notation. Equivalence is a claim about two sets, and it is proved the way set equality is always proved.
Step one is the highest-value step. Most proposed equivalences that feel plausible die on a string of length two.
Estimation
Predict first
Prove the last of the star laws for a two-symbol alphabet.
Commit before you compute: what does Prove an equivalence between two starred expressions come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify on the boundary and on a mixed string
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The empty string matches both sides with every star taken empty.
Worked example
Prove the last of the star laws for a two-symbol alphabet.
\[ (a \cup b)^{*} = (a^{*}b^{*})^{*} \]
Show the left side is contained in the right
Why: Every string on the left is a finite sequence of individual symbols. Each single symbol matches the inner expression on the right, taking one star empty and the other as a single copy. So the same sequence is a valid run of repetitions on the right.
\[ a = a^{1}b^{0}, \qquad b = a^{0}b^{1} \]
Show the right side is contained in the left
Why: Every string on the right is a finite sequence of blocks, and each block is a run of a's followed by a run of b's. Erase the block boundaries: what remains is a finite sequence of individual symbols drawn from the alphabet.
\[ (a^{*}b^{*})^{*} \subseteq \Sigma^{*} = (a \cup b)^{*} \]
Note that both sides are simply everything
Why: Over a two-symbol alphabet, the left side denotes every string at all. So the claim reduces to showing the right side is not missing any string — which the first inclusion did.
Verify on the boundary and on a mixed string
Why: The empty string matches both sides with every star taken empty. The string bab matches the left directly, and matches the right as three blocks: b, then ab, then nothing. Both inclusions plus the boundary give equality.
\[ \varepsilon,\ bab \in \text{both sides} \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Prove an equivalence between two starred expressions", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The empty string matches both sides with every star taken empty. The string bab matches the left directly, and matches the right as three blocks: b, then ab, then nothing. Both inclusions plus the boundary give equality.
Section
Section 5
Concept
The central result of the whole regular story: the machine formalism and the notation formalism describe exactly the same class of languages.
Kleene's theorem — A language is denoted by some regular expression if and only if it is recognized by some finite automaton.
\[ \exists R : L = L(R) \quad \iff \quad \exists M : L = L(M) \]
This is what finally justifies the word regular being used for both. Before this theorem there were two unrelated definitions that happened to share a name.
Matching
Match the pairs
Match each term to the definition this lesson gave it — not the one you would guess from the word.
Why: These are the working definitions of regular expression, equivalent expressions, Kleene's theorem as Regular Expressions uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.
Intuition
An if-and-only-if needs a proof in each direction, and the two halves feel nothing alike.
The first direction is easy because the notation is already inductive: there is one gadget per syntax rule, and induction does the rest. The second is harder because a graph has no inductive structure to recurse on — which is why it gets a lesson of its own.
Concept
Convert an expression to an ε-NFA by structural induction. For every subexpression, build a small machine obeying two strict invariants.
These invariants are what make the gadgets snap together. Because every fragment has a single entry and a single exit, a larger fragment can wire to it without knowing anything about its insides.
\[ \text{one entry, one exit} \;\Rightarrow\; \text{fragments compose} \]
The ε-arrows of Lesson 5 do all the wiring, which is exactly the job they were introduced for.
Picture it
Figure (svg): Automaton with states s, f
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Three atoms, three machines. Each has two states and obeys both invariants.
Concept
Three atoms, three machines. Each has two states and obeys both invariants.
Figure (svg): Automaton with states s, f
| Expression | Gadget |
|---|---|
| a single symbol | start, one arrow labelled with that symbol, accept |
| the empty-string constant | start, one ε-arrow, accept |
| the empty-set constant | start and accept, with no arrow between them at all |
The third gadget is the odd one, and it is correct precisely because nothing connects the two states, so no string can ever reach the accepting state.
Trade off
Comparison matrix
From The base gadgets: every row here is a choice with a cost. Fill the Gadget column, then say which row you would actually pick and what you give up for it.
| Expression | Gadget |
|---|---|
| a single symbol | start, one arrow labelled with that symbol, accept |
| the empty-string constant | start, one ε-arrow, accept |
| the empty-set constant | start and accept, with no arrow between them at all |
Concept
Given fragments for R and for S, build one for their union by offering both roads and then rejoining.
The fresh start makes the choice, non-deterministically, before any symbol is read. The fresh accept collects both outcomes so the result again has a single exit.
\[ \text{states}(R \cup S) = \text{states}(R) + \text{states}(S) + 2 \]
Concept
Given fragments for R and for S, run one and then the other. This is the wiring proved correct in Lesson 7.
No fresh states are needed at all. The one thing that must not be skipped is demoting the first fragment's accepting state — the trap of Lesson 7 in its original habitat.
\[ \text{states}(RS) = \text{states}(R) + \text{states}(S) \]
Picture it
Figure (svg): Automaton with states s, in, out, f
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Given a fragment for R, allow it to run any number of times, including zero.
Concept
Given a fragment for R, allow it to run any number of times, including zero.
Figure (svg): Automaton with states s, in, out, f
The fresh start state is what keeps the construction honest. It has no incoming arrows, so the loop can never be entered from outside — the exact failure the naive construction of Lesson 7 suffered.
Explain it
Discussion prompt
Explain The star gadget to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Given a fragment for R, allow it to run any number of times, including zero.
Step zero
Discussion prompt
Convert a small expression to an ε-NFA — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Build the two leaf gadgets
Answer:
Worked example
Apply Thompson's construction to the expression below, working from the leaves upward.
\[ ab^{*} \]
Build the two leaf gadgets
Why: One two-state gadget reads a, and a separate two-state gadget reads b. Four states so far, and each gadget has one entry and one exit.
\[ 2 + 2 = 4 \text{ states} \]
Star the b gadget
Why: The star gadget adds a fresh start and a fresh accept, wires the loop back and the bypass, and demotes the old accepting state. Two more states.
\[ 2 + 2 = 4 \text{ states in the starred fragment} \]
Concatenate the a gadget onto the starred fragment
Why: Add a single ε-arrow from the a gadget's accepting state to the starred fragment's start state, and demote that accepting state. Concatenation adds no states.
\[ 2 + 4 = 6 \text{ states total} \]
Trace the input a to confirm the bypass works
Why: Read a, cross into the starred fragment by ε, then take the bypass ε-arrow straight to the final accepting state. The string a is accepted, which is right, since zero copies of b is allowed.
\[ a \in L(ab^{*}) \]
Verify against the language the expression denotes
Why: The expression denotes one a followed by any number of b's. The machine accepts a, abb and abbb by looping, and rejects both the empty string and b, since the a-arrow must be taken exactly once before anything else can happen.
\[ a,\ ab,\ abb \in L \qquad \varepsilon,\ b \notin L \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Convert a small expression to an ε-NFA", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The expression denotes one a followed by any number of b's. The machine accepts a, abb and abbb by looping, and rejects both the empty string and b, since the a-arrow must be taken exactly once before anything else can happen.
Ranking
Put in order
Put the moves of Convert an expression with a union and count the states into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. One gadget for a and one for b, two states each.
Worked example
Now a fuller expression, and a count of the machine it produces.
\[ (a \cup b)^{*}a \]
Build the two leaf gadgets
Why: One gadget for a and one for b, two states each.
\[ 2 + 2 = 4 \]
Apply the union gadget
Why: A fresh start and a fresh accept are added, and both old accepting states are demoted.
\[ 4 + 2 = 6 \]
Apply the star gadget
Why: Two more fresh states, plus the loop-back and bypass arrows.
\[ 6 + 2 = 8 \]
Concatenate the final a gadget
Why: A third leaf gadget contributes two states, and the concatenation itself contributes none.
\[ 8 + 2 = 10 \text{ states} \]
Verify the count against the size rule and the language
Why: The expression has three symbol occurrences, one union and one star, so the rule predicts two states per symbol plus two per union and two per star: six plus two plus two, which is ten. Tracing a, ba and aa gives accept in each case, and the empty string is rejected because the trailing a must be read.
\[ 2 \cdot 3 + 2 + 2 = 10 \ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Convert an expression with a union and count the states", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The expression has three symbol occurrences, one union and one star, so the rule predicts two states per symbol plus two per union and two per star: six plus two plus two, which is ten. Tracing a, ba and aa gives accept in each case, and the empty string is rejected because the trailing a must be read.
Concept
Each gadget adds at most two states, and there is one gadget per node of the parse tree. So the machine never blows up.
\[ |Q| \;\le\; 2|R| \]
That bound is worth holding next to the one from Lesson 6. Converting an expression to an NFA is cheap; converting that NFA to a DFA is where the exponential lives.
| Conversion | Cost |
|---|---|
| expression to ε-NFA | linear |
| ε-NFA to DFA | exponential in the worst case |
| DFA to expression | exponential in the worst case, see Lesson 9 |
Practical tools exploit exactly this: they build the linear NFA and simulate it directly, determinizing lazily or not at all.
Comparison
Comparison matrix
From The construction is linear in the size of the expression: refill the Cost column from what you know. The rest of the table is as it appeared.
| Conversion | Cost |
|---|---|
| expression to ε-NFA | linear |
| ε-NFA to DFA | exponential in the worst case |
| DFA to expression | exponential in the worst case, see Lesson 9 |
Elimination
Eliminate the wrong options
In Thompson's construction, why must the concatenation gadget demote the first fragment's accepting state?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Reaching the end of the first fragment means the prefix matched R, not that the whole string matched RS. If that state stayed accepting, every string of L(R) would be accepted outright, so the machine would recognize a strictly larger language.
Check
Think about which gadget adds states and which does not.
Check your understanding
In Thompson's construction, why must the concatenation gadget demote the first fragment's accepting state?
Answer: A
Why: Reaching the end of the first fragment means the prefix matched R, not that the whole string matched RS. If that state stayed accepting, every string of L(R) would be accepted outright, so the machine would recognize a strictly larger language.
Ranking
Put in order
These are the steps of The conversion recipe, start to finish, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Given any regular expression, this produces an ε-NFA for it every time.
If a determinstic machine is wanted, follow with the subset construction from Lesson 6. Expression to ε-NFA to DFA is the standard pipeline, and every stage of it has now been proved correct.
Real world
Discussion prompt
Outside this lesson: where does Regular Expressions actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The conversion recipe, start to finish is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
Lesson 8 introduces the notation that describes exactly the languages finite automata recognize. Covers the six-rule inductive syntax, the semantic clauses that assign a language to every expression, the star-then-concatenation-then-union precedence order, design idioms built around the starred alphabet, the algebraic laws including the identities, distribution and the star laws, and Thompson's construction converting any expression into a linear-size epsilon-NFA.
Intuition
Going from a machine back to an expression cannot be done by recursion, because a graph has no leaves to start from and no root to finish at.
Cycles are the difficulty. A loop in the diagram means a string can revisit a state any number of times, and capturing that in a finite formula is exactly what the star operator is for — but finding which star, for which loop, needs a systematic method.
Lesson 9 supplies it: rip out one state at a time, and whenever a state is removed, relabel the arrows that used to pass through it with an expression that summarizes every path it offered. When only the start and accept remain, the surviving label is the answer.
\[ \text{state elimination} \;:\; \text{graph} \;\longrightarrow\; \text{expression} \]
Analogy
Discussion prompt
Explain Why the reverse direction needs its own lesson by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Going from a machine back to an expression cannot be done by recursion, because a graph has no leaves to start from and no root to finish at.
Concept
The theorem is not a curiosity. It means every result proved about one formalism transfers instantly to the other.
From here on, showing a language is regular means producing either a machine or an expression, whichever is easier — and the two lessons that follow show how to move between them mechanically.
\[ \text{DFA} \;\equiv\; \text{NFA} \;\equiv\; \varepsilon\text{-NFA} \;\equiv\; \text{regular expression} \]
Counterexample
Discussion prompt
The theorem is not a curiosity. It means every result proved about one formalism transfers instantly to the other.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
From here on, showing a language is regular means producing either a machine or an expression, whichever is easier — and the two lessons that follow show how to move between them mechanically.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Syntax: What an Expression Is · Semantics: The Language of an Expression · Designing Expressions · The Algebra of Expressions · Expressions and Automata Are the Same. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can now write regular languages as formulas, manipulate those formulas algebraically, and convert them into machines mechanically.
| Situation | Move |
|---|---|
| an expression to understand | parse it, then read it aloud by the recipe |
| a language to describe | fixed skeleton first, anything idiom in the gaps |
| two expressions to compare | hunt a counterexample of length two, then prove both inclusions |
| an expression to run | Thompson's construction, then the subset construction if needed |
Lesson 9 closes the loop by converting a machine back into an expression, completing the proof of Kleene's theorem.
Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.