Lesson 3 of the mathematical toolkit, and the gateway to automata. It covers alphabets and strings, length and the empty string, concatenation and its algebra, substrings, prefixes, and suffixes, and string exponentiation and reversal. It then moves to the set of all strings, languages as sets of strings, the language operations including Kleene star, and the counting results showing that strings are countable while languages are not. It targets the confusion between the empty string and the empty set, the confusion between a substring and a subsequence, the order in which a concatenation reverses, and the difference between star and plus. All computations were verified by hand.
Subject: Theory of Computation · 99 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Theory of Computation · Lesson 3
The objects every machine reads and every grammar generates — and a first look at what computers cannot do.
Objectives
Everything a computation processes is a string; every problem is a language. This lesson makes both precise. By the end you can:
Warm-up
Discussion prompt
Before we open Strings & Languages: without looking back, what was the main idea of Proof, Induction & Closures, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Lesson 2 of the math toolkit: what a proof is, direct proof and the contrapositive, proof by contradiction (irrationality of root 2, infinitude of primes), weak and strong mathematical induction, the well-ordering principle, the pigeonhole principle, recursive definitions, and computing the reflexive/symmetric/transitive closures of a relation. Targets converse-vs-contrapositive confusion, the missing base case, and the one-round transitive-closure error.
Section
Section 1
Concept
An alphabet is any finite, non-empty set of symbols. It is the raw material — the characters everything else is built from. We name it with the Greek letter sigma.
\[ \Sigma = \{a, b\} \qquad \Sigma = \{0, 1\} \]
symbol — An indivisible member of the alphabet. What counts as a single symbol is whatever the alphabet declares — a letter, a digit, or a whole token.
Counterexample
Discussion prompt
An alphabet is any finite, non-empty set of symbols. It is the raw material — the characters everything else is built from. We name it with the Greek letter sigma.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Intuition
Think of the alphabet as the set of keys on a keyboard. It is finite and fixed in advance, and every message you can ever produce is some sequence of those keys.
Change the keyboard and you change what can be written. A binary alphabet has two keys; the ASCII alphabet has many. The theory works the same for any of them.
Analogy
Discussion prompt
Explain The alphabet is your keyboard by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Think of the alphabet as the set of keys on a keyboard. It is finite and fixed in advance, and every message you can ever produce is some sequence of those keys.
Concept
A string (or word) over an alphabet is a finite sequence of its symbols, written with no separators. Order matters and symbols may repeat.
\[ w = abba \quad \text{over} \quad \Sigma = \{a,b\} \]
Unlike a set, a string remembers position and repetition. The string 'ab' is different from 'ba', and 'aa' is a perfectly good two-symbol string.
Explain it
Discussion prompt
Explain A string is a finite sequence of symbols to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A string (or word) over an alphabet is a finite sequence of its symbols, written with no separators. Order matters and symbols may repeat.
Concept
The length of a string is the number of symbol occurrences in it — counting repeats. It is written with vertical bars.
\[ |abba| = 4 \qquad |aaa| = 3 \]
Length counts positions, not distinct symbols: the string 'aaa' has length 3 even though it uses only one symbol.
Concept
There is one special string with no symbols at all: the empty string, written with the Greek letter epsilon. Its length is zero.
\[ \varepsilon \quad\text{with}\quad |\varepsilon| = 0 \]
It is a genuine string — the sequence of length zero — and it will play the role that zero plays in arithmetic and the empty set plays in set theory.
Ranking
Put in order
Put the moves of Read off symbols, strings, and lengths into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Scan w: its symbols are 1, 0, 1, 1, 0 — all drawn from the alphabet {0,1}, so w is a legal string over this alphabet.
Worked example
Work over the binary alphabet and answer three questions about the string below.
\[ \Sigma = \{0,1\}, \qquad w = 10110 \]
Confirm every symbol is in the alphabet
Why: Scan w: its symbols are 1, 0, 1, 1, 0 — all drawn from the alphabet {0,1}, so w is a legal string over this alphabet.
Count the length
Why: There are five symbol occurrences, so the length is 5. Repeats of 1 and 0 each count.
\[ |w| = |10110| = 5 \]
Verify against the empty string
Why: Deleting all five symbols would leave the empty string of length 0. Since w has 5 symbols, it is certainly not empty — the length count is consistent.
\[ |w| = 5 \neq 0 = |\varepsilon|\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Read off symbols, strings, and lengths", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Deleting all five symbols would leave the empty string of length 0. Since w has 5 symbols, it is certainly not empty — the length count is consistent.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student treats the empty string, a blank space, and the empty set as interchangeable.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Confuses a zero-length STRING with a zero-element SET, and with a space CHARACTER.
Keep the three apart — different types entirely.
Why: Confuses a zero-length STRING with a zero-element SET, and with a space CHARACTER. All three are different kinds of object.
Trap
A student treats the empty string, a blank space, and the empty set as interchangeable.
\[ \varepsilon \overset{?}{=} \varnothing \overset{?}{=} \text{' '} \]
Claim ε is 'nothing', so it equals the empty set
Why: Confuses a zero-length STRING with a zero-element SET, and with a space CHARACTER. All three are different kinds of object.
\[ |\varepsilon| = 0 \ \text{but}\ \varepsilon \neq \varnothing \]
Keep the three apart — different types entirely.
\[ \varepsilon,\quad \varnothing,\quad \text{space} \]
ε is a string; ∅ is a set
Why: The empty string is a sequence of length 0 — an element of the strings. The empty set is a set with no members. A string is not a set.
\[ \varepsilon \in \Sigma^{*}, \qquad \varnothing \subseteq \Sigma^{*} \]
A space is a symbol of length one
Why: If the alphabet includes a blank character, that blank is an ordinary symbol and a string of it has length 1 — unlike ε, which has length 0.
\[ |\text{' '}| = 1 \neq 0 = |\varepsilon| \]
Notation
Annotate
From Trap: the empty string is not the empty set — read this one piece at a time. What is each part doing?
On: \( \varepsilon \overset{?}{=} \varnothing \overset{?}{=} \text{' '} \)
Concept
Collect every finite string over the alphabet — of every length, including the empty string — and you get Σ-star.
\[ \Sigma^{*} = \{\, \text{all finite strings over } \Sigma \,\} \]
Drop the empty string and you get Σ-plus, the non-empty strings. They differ by exactly one element.
\[ \Sigma^{+} = \Sigma^{*} \setminus \{\varepsilon\} \]
Intuition
Σ-star is the universe of this course. Every input to every machine, every program, every proof written in the alphabet lives inside it.
It is infinite — there is no longest string — but every individual member is finite. That combination, infinitely many finite things, is exactly what makes it countable, as we will see at the end.
Pattern
1. Fix the alphabet first
Why: Every string and language is 'over' some alphabet; know its symbols before anything else.
2. Read a string left to right, counting positions
Why: Length is the number of positions, repeats included; the empty string has none.
3. Know which universe a symbol lives in
Why: Elements of Σ are symbols; elements of Σ-star are strings; subsets of Σ-star are languages. Keep the three levels straight.
Check
Work over the alphabet below.
\[ \Sigma = \{a,b\} \]
Check your understanding
Which statement is TRUE?
Answer: A
Why: The empty string is a genuine member of Σ*, the set of ALL finite strings, and its length is 0 by definition. Both parts of the statement are correct.
Section
Section 2
Concept
The one fundamental operation on strings is concatenation: write the first string, then the second, with nothing between.
\[ u = ab,\ v = ba \ \Rightarrow\ uv = abba \]
The length of a concatenation is the sum of the lengths — no symbols are created or lost by joining.
\[ |uv| = |u| + |v| \]
Intuition
Picture each string as a printed ribbon. Concatenation glues the end of one to the start of the next, making a single longer ribbon. Nothing is inserted at the seam.
Because you can always glue on another ribbon, concatenation is how every string is built up from single symbols — and how machines consume input one symbol at a time.
Concept
Concatenation is associative — regrouping does not change the result — so we drop parentheses freely.
\[ (uv)w = u(vw) \]
The empty string is its identity: gluing on nothing changes nothing. But concatenation is not commutative — order matters.
\[ \varepsilon w = w \varepsilon = w, \qquad uv \neq vu \text{ in general} \]
Step zero
Discussion prompt
Concatenate and check the length — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Write u then v with no gap
Answer:
Worked example
Concatenate the two strings, in order, and confirm the length rule.
\[ u = abb, \qquad v = ba \]
Write u then v with no gap
Why: Lay down abb, then ba, joined directly: a-b-b-b-a.
\[ uv = abb \cdot ba = abbba \]
Note that order matters
Why: The reverse concatenation vu = ba·abb = baabb is a different string, confirming concatenation is not commutative.
\[ vu = baabb \neq abbba = uv \]
Verify the length adds
Why: u has length 3 and v has length 2; the result abbba has length 5, matching the sum. The length rule holds.
\[ |uv| = 5 = 3 + 2 = |u| + |v|\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Concatenate and check the length", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: u has length 3 and v has length 2; the result abbba has length 5, matching the sum. The length rule holds.
Concept
A substring is any block of consecutive symbols sitting inside a string. A prefix is a substring at the very start; a suffix is one at the very end.
\[ w = abc: \quad \text{prefixes } \varepsilon, a, ab, abc \]
The empty string and the whole string count as prefixes, suffixes, and substrings. A string of length n has exactly n+1 prefixes and n+1 suffixes.
Ranking
Put in order
Put the moves of List every prefix and suffix into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Take the first 0, 1, 2, then 3 symbols.
Worked example
Find all prefixes and all suffixes of the string below.
\[ w = abc, \qquad |w| = 3 \]
Prefixes: cut after each position, from 0 to 3
Why: Take the first 0, 1, 2, then 3 symbols. That gives four prefixes including the empty string and the whole word.
\[ \varepsilon,\ a,\ ab,\ abc \]
Suffixes: cut before each position
Why: Take the last 0, 1, 2, then 3 symbols, giving four suffixes.
\[ \varepsilon,\ c,\ bc,\ abc \]
Verify the count
Why: Length 3 predicts 3 + 1 = 4 prefixes and 4 suffixes, exactly the numbers found. The count checks out.
\[ n + 1 = 3 + 1 = 4\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "List every prefix and suffix", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Length 3 predicts 3 + 1 = 4 prefixes and 4 suffixes, exactly the numbers found. The count checks out.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Asked for a substring of 'abcd', a student offers 'acd'.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Confuses substring with subsequence.
A substring must be a contiguous block; a subsequence may skip around but must keep order.
Why: Confuses substring with subsequence. 'acd' skips b, so its symbols are not consecutive in w — that makes it a subsequence, not a substring.
Trap
Asked for a substring of 'abcd', a student offers 'acd'.
\[ w = abcd, \quad \text{'is } acd \text{ a substring?'} \]
Accept acd as a substring
Why: Confuses substring with subsequence. 'acd' skips b, so its symbols are not consecutive in w — that makes it a subsequence, not a substring.
\[ acd \ \text{skips } b \ \Rightarrow \ \text{not consecutive} \]
A substring must be a contiguous block; a subsequence may skip around but must keep order.
\[ w = abcd \]
Substrings are consecutive
Why: Legitimate substrings of abcd are blocks like bc and abc — no gaps. 'acd' is not among them.
\[ \text{substrings include } bc,\ abc,\ bcd \]
Subsequences may have gaps
Why: 'acd' is a valid subsequence because it keeps the left-to-right order while dropping b. Every substring is a subsequence, but not the reverse.
\[ acd \ \text{is a subsequence, not a substring} \]
Translation
\( \text{substrings include } bc,\ abc,\ bcd \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Concept
Writing a string to a power means concatenating it with itself that many times. Power zero is defined as the empty string.
\[ w^{0} = \varepsilon, \qquad w^{n} = w\,w^{\,n-1} \]
The length multiplies: repeating a string n times gives n copies of its symbols.
\[ |w^{n}| = n \cdot |w| \]
Estimation
Predict first
Evaluate the string 'ab' raised to the third power.
Commit before you compute: what does Compute a string power come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the length multiplies
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. w has length 2, so w cubed should have length 3 times 2 = 6.
Worked example
Evaluate the string 'ab' raised to the third power.
\[ w = ab, \qquad w^{3} = ? \]
Concatenate three copies
Why: Glue ab to ab to ab in a row.
\[ w^{3} = ab \cdot ab \cdot ab = ababab \]
Verify the length multiplies
Why: w has length 2, so w cubed should have length 3 times 2 = 6. The result ababab has six symbols — consistent.
\[ |w^{3}| = 6 = 3 \cdot 2 = 3 \cdot |w|\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Compute a string power", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: w has length 2, so w cubed should have length 3 times 2 = 6. The result ababab has six symbols — consistent.
Pattern
1. Concatenate by writing in order
Why: First string then second, no separator; the length is the sum of lengths.
2. Substrings are contiguous; subsequences may skip
Why: For prefixes and suffixes, cut at each position — a length-n string has n+1 of each.
3. A power repeats the whole string
Why: w to the n is n glued copies; its length is n times |w|, and w to the 0 is the empty string.
Check
Let the two strings be as below.
\[ u = aba, \qquad v = bb \]
Check your understanding
What is uv, and what is its length?
Answer: A
Why: Concatenation writes u then v: aba followed by bb gives ababb. Its length is |u| + |v| = 3 + 2 = 5.
Section
Section 3
Concept
The reversal of a string writes its symbols in the opposite order. It is marked with a superscript R.
\[ w = abc \ \Rightarrow \ w^{R} = cba \]
Reversal preserves length — the same symbols appear, only reordered — and the empty string is its own reversal.
\[ |w^{R}| = |w|, \qquad \varepsilon^{R} = \varepsilon \]
Intuition
If a string is a printed ribbon, its reversal is the same ribbon read from the far end. No symbol is added or removed — the order is simply mirrored.
Concept
Reversing a concatenation reverses each piece and swaps their order. The outside becomes the inside.
\[ (uv)^{R} = v^{R} u^{R} \]
This mirrors putting on socks then shoes: to undo, you take off shoes first, then socks. The last part attached is the first part to appear when reversed.
Sorting
Sort into buckets
These are the pieces of Strings & Languages, out of order. Put each one back under the part of the lesson it belongs to.
Step zero
Discussion prompt
Reverse a concatenation — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Concatenate first, then reverse the whole thing
Answer:
Worked example
Reverse the concatenation of the two strings, and confirm the swap rule.
\[ u = ab, \qquad v = cd \]
Concatenate first, then reverse the whole thing
Why: uv = abcd; reading it backward gives dcba.
\[ (uv)^{R} = (abcd)^{R} = dcba \]
Now reverse the pieces and swap
Why: Reverse each: u reversed is ba, v reversed is dc. Swap the order — v reversed first — to get dc followed by ba.
\[ v^{R} u^{R} = dc \cdot ba = dcba \]
Verify both routes agree
Why: Reversing the whole and swapping the reversed parts both give dcba. Note u reversed then v reversed would give badc — the wrong answer, which is why order swaps.
\[ (uv)^{R} = dcba = v^{R} u^{R}\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Reverse a concatenation", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Reversing the whole and swapping the reversed parts both give dcba. Note u reversed then v reversed would give badc — the wrong answer, which is why order swaps.
Concept
A palindrome is a string that equals its own reversal — the same forward and backward.
\[ w \text{ is a palindrome} \iff w = w^{R} \]
The empty string and every single symbol are trivially palindromes. Palindromes will return as a classic example of a language no finite-memory machine can recognize.
Worked example
Decide whether the string below is a palindrome.
\[ w = abba \]
Compute the reversal
Why: Read abba backward: a, b, b, a — which is abba again.
\[ w^{R} = (abba)^{R} = abba \]
Verify by comparing to the original
Why: The reversal equals the original string, so by definition w is a palindrome. Contrast abb, whose reversal bba differs — not a palindrome.
\[ w = abba = w^{R} \Rightarrow \text{palindrome}\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Test a string for the palindrome property", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The reversal equals the original string, so by definition w is a palindrome. Contrast abb, whose reversal bba differs — not a palindrome.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student reverses each part of a concatenation but leaves the parts in the same order.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Reverses the pieces but forgets to swap them.
Reverse each piece AND reverse their order.
Why: Reverses the pieces but forgets to swap them. This yields badc, which is not the reversal of abcd.
Trap
A student reverses each part of a concatenation but leaves the parts in the same order.
\[ u = ab,\ v = cd, \quad (uv)^{R} \overset{?}{=} u^{R} v^{R} \]
Write u^R v^R = ba·dc = badc
Why: Reverses the pieces but forgets to swap them. This yields badc, which is not the reversal of abcd.
\[ u^{R} v^{R} = badc \neq dcba \]
Reverse each piece AND reverse their order.
\[ (uv)^{R} = v^{R} u^{R} \]
Swap the reversed pieces
Why: Put v reversed first, then u reversed: dc followed by ba gives dcba, which truly is abcd read backward.
\[ v^{R} u^{R} = dc \cdot ba = dcba \]
Sanity-check against the direct reversal
Why: Reading abcd straight backward gives dcba — matching the swapped form and refuting the un-swapped badc.
\[ (abcd)^{R} = dcba\ \checkmark \]
Notation
Annotate
From Trap: reversing a concatenation keeps the order — read this one piece at a time. What is each part doing?
On: \( (abcd)^{R} = dcba\ \checkmark \)
Explain it to yourself
Discussion prompt
In Working with reversal this move is made:
2. Reverse a concatenation by swap-and-reverse
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
Reverse each part and reverse their order: (uv) reversed is v-reversed then u-reversed.
Pattern
1. Reverse a single string by mirroring symbols
Why: Last symbol becomes first; length is unchanged.
2. Reverse a concatenation by swap-and-reverse
Why: Reverse each part and reverse their order: (uv) reversed is v-reversed then u-reversed.
3. Test a palindrome by comparing to its reversal
Why: The string is a palindrome exactly when it equals its own reversal.
Check
Take the two strings below.
\[ u = go, \qquad v = ld \]
Check your understanding
What is (uv)ᴿ?
Answer: A
Why: First uv = go·ld = gold. Reversing gold reads it backward: d-l-o-g = dlog. Equivalently, vᴿuᴿ = dl·og = dlog.
Section
Section 4
Concept
A language over an alphabet is simply any set of strings over it — that is, any subset of Σ-star.
\[ L \subseteq \Sigma^{*} \]
That is the entire definition. 'The set of binary strings with an even number of ones' is a language; so is 'the set of valid programs'. Both are just subsets of Σ-star.
Intuition
Every decision problem is secretly a language: the language is exactly the set of inputs whose answer is 'yes'. 'Is this number prime?' becomes the language of all strings that encode primes.
So when a machine 'recognizes a language', it is answering a yes/no question about its input. This is why languages, not numbers, are the central objects of computation theory.
Concept
Three small languages look alike and are constantly mixed up. Learn them cold.
\[ \varnothing, \qquad \{\varepsilon\}, \qquad \Sigma^{*} \]
The first has zero members; the second has exactly one member; the third has infinitely many. They could not be more different.
Estimation
Predict first
Over the binary alphabet, take the language of all strings of even length. List its members up to length 2.
Commit before you compute: what does Describe a language and list short members come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the exclusions
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The length-1 strings 0 and 1 are odd, hence not in L, and no even-length string up to 2 was missed.
Worked example
Over the binary alphabet, take the language of all strings of even length. List its members up to length 2.
\[ L = \{\, w \in \{0,1\}^{*} \mid |w| \text{ is even} \,\} \]
Length 0
Why: Only the empty string has length 0, and 0 is even, so ε is in L.
\[ \varepsilon \in L \]
Length 2
Why: All four two-symbol strings have even length and belong to L. Odd length 1 strings are excluded.
\[ 00,\ 01,\ 10,\ 11 \in L \]
Verify the exclusions
Why: The length-1 strings 0 and 1 are odd, hence not in L, and no even-length string up to 2 was missed. The membership rule is applied consistently.
\[ 0 \notin L,\ 1 \notin L\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Describe a language and list short members", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The length-1 strings 0 and 1 are odd, hence not in L, and no even-length string up to 2 was missed. The membership rule is applied consistently.
Concept
Languages are sets, so every operation from Lesson 1 works: union, intersection, difference, and complement (against the universe Σ-star).
\[ \overline{L} = \Sigma^{*} \setminus L \]
The complement of a language is 'all the strings it rejects' — flipping every yes to a no and vice versa.
Concept
Languages get a new operation strings gave them: concatenation. Take every string from the first language glued to every string from the second.
\[ L_1 L_2 = \{\, xy \mid x \in L_1,\ y \in L_2 \,\} \]
It is the language-level version of gluing ribbons — but now over all combinations, like a Cartesian product that concatenates each pair.
Fill the middle
Fill in the blanks
From Concatenate two languages — finish the line. Write what belongs on the right of the equals sign before you look.
L_1 L_2 = \{ab,\ ac,\ abb,\ abc\}
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Two choices from L1 times two from L2 gives four gluings: a·b, a·c, ab·b, ab·c.
Worked example
Concatenate the two small languages below.
\[ L_1 = \{a, ab\}, \qquad L_2 = \{b, c\} \]
Pair every x with every y and glue
Why: Two choices from L1 times two from L2 gives four gluings: a·b, a·c, ab·b, ab·c.
\[ L_1 L_2 = \{ab,\ ac,\ abb,\ abc\} \]
Verify the count and check for collisions
Why: Four pairings produced four distinct strings — none coincided — so the result has four members, matching two times two. If two gluings had matched, the set would be smaller.
\[ |L_1 L_2| = 4 = |L_1| \cdot |L_2|\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Concatenate two languages", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Four pairings produced four distinct strings — none coincided — so the result has four members, matching two times two. If two gluings had matched, the set would be smaller.
Concept
The Kleene star of a language is every string you can form by concatenating zero or more members of it. Zero copies gives the empty string, so ε is always in the star.
\[ L^{*} = \bigcup_{i \geq 0} L^{i} = L^{0} \cup L^{1} \cup L^{2} \cup \cdots \]
The plus version demands one or more copies, so it omits the empty string unless the empty string was already in L.
\[ L^{+} = \bigcup_{i \geq 1} L^{i} \]
Intuition
The star is the theory's loop. It says: pick a member, then maybe another, then maybe another — stop whenever you want, including immediately (that is the empty string).
Applied to a whole alphabet treated as one-symbol strings, this is exactly how Σ-star got its name: zero or more symbols, glued — every possible string.
Ranking
Put in order
Put the moves of Compute a Kleene star into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Zero copies is the empty string; one copy is ab; two copies is abab.
Worked example
Find the members of the star of the one-string language below, up to length 4.
\[ L = \{ab\} \]
Zero copies, then one, then two
Why: Zero copies is the empty string; one copy is ab; two copies is abab. Each copy adds the block ab.
\[ L^{0} = \{\varepsilon\},\ L^{1} = \{ab\},\ L^{2} = \{abab\} \]
Collect the star
Why: Union over all counts gives every repetition of ab, including the empty string.
\[ L^{*} = \{\varepsilon,\ ab,\ abab,\ ababab,\ \ldots\} \]
Verify the empty string is included
Why: The zero-copy case puts ε in every Kleene star. In contrast, L-plus starts at one copy, so its shortest member is ab, not ε — confirming the star/plus difference.
\[ \varepsilon \in L^{*}, \qquad \varepsilon \notin L^{+}\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Compute a Kleene star", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The zero-copy case puts ε in every Kleene star. In contrast, L-plus starts at one copy, so its shortest member is ab, not ε — confirming the star/plus difference.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student writes that the star of the empty language is empty.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Forgets the zero-copy case. Concatenating zero strings gives the empty string regardless of L, so the star can never be truly empty.
The star always contains the empty string from its zero-copy term.
Why: Forgets the zero-copy case. Concatenating zero strings gives the empty string regardless of L, so the star can never be truly empty.
Trap
A student writes that the star of the empty language is empty.
\[ \varnothing^{*} \overset{?}{=} \varnothing \]
Claim ∅* = ∅
Why: Forgets the zero-copy case. Concatenating zero strings gives the empty string regardless of L, so the star can never be truly empty.
\[ \varnothing^{*} \neq \varnothing \]
The star always contains the empty string from its zero-copy term.
\[ L^{0} = \{\varepsilon\} \subseteq L^{*} \]
Star of the empty language is {ε}
Why: You can select zero strings from the empty language (that needs no members), producing ε. So the star is the one-element language containing the empty string.
\[ \varnothing^{*} = \{\varepsilon\} \]
Keep the three straight
Why: The empty language has no strings; its star has exactly one, the empty string; and that is still a far cry from all of Σ-star.
\[ \varnothing \neq \{\varepsilon\} \neq \Sigma^{*} \]
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Pattern
1. Set operations act member-wise
Why: Union, intersection, difference, complement treat languages as the sets of strings they are; complement is taken against Σ-star.
2. Concatenation glues every pair
Why: L1L2 is all xy with x from L1 and y from L2 — combine across the two, then discard duplicate strings.
3. Star = zero-or-more copies; plus = one-or-more
Why: The star always includes ε (zero copies); the plus starts at one copy. Even the empty language's star is {ε}.
Check
Consider the language below over the binary alphabet.
\[ L = \{0, 1\} \]
Check your understanding
Which string is NOT in L*?
Answer: A
Why: L contains both single symbols 0 and 1, so concatenating zero or more of them produces every possible binary string, plus the empty string. Thus L* equals {0,1}* and excludes nothing over this alphabet.
Section
Section 5
Concept
Over an alphabet of size k, a string of length n is n independent choices of symbol, so the count is k to the n — the same multiplication principle as tuples in Lesson 1.
\[ \#\{\, w : |w| = n \,\} = |\Sigma|^{\,n} \]
Each length gives a finite pile of strings; stacking the piles for all lengths gives the infinite Σ-star.
Estimation
Predict first
How many strings of length 3 are there over the binary alphabet?
Commit before you compute: what does Count the strings of a fixed length come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify by listing them all
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The eight strings are 000, 001, 010, 011, 100, 101, 110, 111 — exactly eight, matching the formula.
Worked example
How many strings of length 3 are there over the binary alphabet?
\[ \Sigma = \{0,1\}, \quad n = 3 \]
Apply the formula
Why: Alphabet size 2 raised to length 3 gives 2 times 2 times 2 = 8.
\[ |\Sigma|^{n} = 2^{3} = 8 \]
Verify by listing them all
Why: The eight strings are 000, 001, 010, 011, 100, 101, 110, 111 — exactly eight, matching the formula.
\[ 000,001,010,011,100,101,110,111 \ (8)\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Count the strings of a fixed length", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The eight strings are 000, 001, 010, 011, 100, 101, 110, 111 — exactly eight, matching the formula.
Concept
A set is countable when its members can be arranged in one infinite list, indexed by the natural numbers. Σ-star can.
\[ \Sigma^{*} = \{ w_0, w_1, w_2, w_3, \ldots \} \]
The trick is to order strings first by length, then alphabetically within each length — the shortlex order — so every string gets a definite position.
Intuition
Order by length, breaking ties alphabetically: ε, then 0, 1, then 00, 01, 10, 11, then the length-3 strings, and so on. Each finite pile is listed completely before the next begins.
Because every string is finite, it sits in some pile and therefore appears at a specific numbered spot. Nothing is skipped and nothing waits forever — that is exactly what countable means.
Concept
A language is a subset of Σ-star, so the set of all languages is the power set of Σ-star. And the power set of an infinite countable set is uncountable — strictly bigger.
\[ \{\text{all languages}\} = \mathcal{P}(\Sigma^{*}) \]
No infinite list can contain every language. Any proposed listing must miss some language — a fact proved by Cantor's diagonal argument.
Explain it
Discussion prompt
Explain The set of all languages is uncountable to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A language is a subset of Σ-star, so the set of all languages is the power set of Σ-star. And the power set of an infinite countable set is uncountable — strictly bigger.
Intuition
Suppose someone lists languages L0, L1, L2, and claims the list is complete. Line up the strings w0, w1, w2 from the countable Σ-star along the top.
Build a new language D by disagreeing on the diagonal: put wᵢ into D exactly when Lᵢ leaves it out. Then D differs from every Lᵢ on the string wᵢ, so D is nowhere on the list. The 'complete' list was not complete.
\[ w_i \in D \iff w_i \notin L_i \]
Analogy
Discussion prompt
Explain The diagonal escape by analogy to something with no Theory of Computation in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Suppose someone lists languages L0, L1, L2, and claims the list is complete. Line up the strings w0, w1, w2 from the countable Σ-star along the top.
Concept
Programs are finite strings over a finite alphabet, so there are only countably many programs. But there are uncountably many languages.
\[ \#\text{programs} = \aleph_0 \ < \ \#\text{languages} \]
Countably many programs cannot cover uncountably many languages. So most languages are not recognized by any program — undecidability is not a rare bug, it is the overwhelming majority. That is where this course is headed.
Counterexample
Discussion prompt
Programs are finite strings over a finite alphabet, so there are only countably many programs. But there are uncountably many languages.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Pattern
1. Strings of length n number |Σ| to the n
Why: Independent symbol choices multiply; each length is a finite pile.
2. Σ-star is countable via shortlex
Why: List by length then alphabetically; every finite string gets a numbered slot.
3. Languages are the power set — uncountable
Why: Cantor's diagonal builds a language missing from any list, so no enumeration is complete.
4. Conclude: programs (countable) cannot cover languages (uncountable)
Why: Sizes force the gap — most languages are unrecognizable, the seed of every impossibility result to come.
Real world
Discussion prompt
Outside this lesson: where does Strings & Languages actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The counting argument in one breath is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
Lesson 3 of the math toolkit and the gateway to automata: alphabets and strings, length and the empty string, concatenation and its algebra, substrings/prefixes/suffixes, string exponentiation and reversal, the set of all strings, languages as sets of strings, language operations including Kleene star, and the counting results that show strings are countable but languages are not. Targets the empty-string-vs-empty-set confusion, substring-vs-subsequence, the reversal-of-concatenation order, and star-vs-plus.
Check
Fix a finite alphabet Σ.
Check your understanding
Which statement is correct?
Answer: A
Why: Σ* can be listed in shortlex order, so it is countably infinite. The set of all languages is its power set, and the power set of a countably infinite set is uncountable by Cantor's diagonal argument. This size gap is why most languages have no recognizing program.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Alphabets & Strings · Concatenation & Structure · Reversal & Palindromes · Languages · Counting: Strings vs Languages. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can now speak the language of languages — the objects the rest of automata theory manipulates.
| Object | Key fact |
|---|---|
| ε (empty string) | Length 0; identity for concatenation |
| (uv)ᴿ | Equals vᴿuᴿ — order swaps |
| L* (Kleene star) | Zero or more copies; always contains ε |
| All languages | P(Σ*) — uncountable |
Want this taught 1-on-1? Alexander tutors Theory of Computation — $55/session, free consultation.