Session 1 - What to Expect & the Questions That Repeat

This first session maps the technical interview itself and then the small set of patterns most questions are built from. It covers what a 45-minute round actually contains and the four axes it is scored on, the six-step loop to run on any unseen problem, and how to read a problem's constraints as a statement of the complexity you are expected to hit. It then works the eight patterns that carry most interview questions - hash map, two pointers, sliding window, binary search, stack, BFS/DFS, heap, and dynamic programming - each with a runnable Python solution and a real execution trace, including Two Sum, valid palindrome, longest substring without repeating characters, daily temperatures, number of islands, climbing stairs, and merge intervals. It closes with the twenty most frequently asked problems mapped to their patterns, a narration and self-testing script for the round itself, and a ten-week practice plan. Traps target the mistakes that actually cost rounds: coding before clarifying, the accidental quadratic from a list membership test, rebuilding a window instead of sliding it, sorting away the indices the question asked for, and handing verification back to the interviewer.

Subject: DSA Interview Prep · 85 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. What to Expect, and the Questions That Repeat

Title

DSA Interview Prep ยท Session 1

A map of the interview itself, then the handful of patterns that most questions are wearing a costume over. Python throughout.

2. What you will be able to do

Objectives

Interview questions look infinite and are not. This first session builds the map: what a round actually contains, and which shapes keep coming back. By the end you can:

  1. Describe the structure of a technical round and name what the interviewer is scoring besides a working answer.
  2. Run the six-step loop - clarify, examples, brute force, optimize, code, test - on a problem you have never seen.
  1. Read a problem's constraints and name the complexity you are being asked for before writing a line.
  2. Recognize the eight patterns that cover most interview questions, and give the signature clue that triggers each one.
  3. Recall the most frequently asked problems and say which pattern each belongs to rather than which solution you memorized.

3. What You Are Actually Walking Into

Section

Section 1

4. Estimate first: how big is the surface really?

Estimation

Before we look at any list. LeetCode alone has well over 3,000 problems, and companies pull from all of them.

Predict first

Roughly how many distinct problem patterns would you have to know to have a workable approach to the large majority of interview questions?

  • Around 300 - basically one per problem
  • Around 60
  • Around 15
  • Around 3

Correct: Around 15 - and about 8 of them carry most of the weight.

  • Hash map / counting
  • Two pointers
  • Sliding window
  • Binary search
  • Stack (including monotonic)
  • Tree & graph traversal (BFS / DFS)
  • Heap / top-K
  • Dynamic programming

Why: Interviewers reuse a small set of structural ideas: hash lookup, two pointers, sliding window, binary search, stack, tree/graph traversal, heap, and dynamic programming. Problems differ in story and in edge cases far more than in structure, which is exactly why memorizing solutions scales badly and recognizing patterns scales well.

5. The rounds, in the order you meet them

Concept

Almost every company runs some version of this ladder. The DSA content is concentrated in the middle two.

Recruiter screen
30 min, no code. Background and logistics.
Online assessment
60-90 min, auto-graded. Hidden tests, no human to talk to.
Technical phone screen
45-60 min, 1-2 problems, shared editor, human listening.
Onsite loop
3-5 rounds: DSA, sometimes system design, always behavioral.

The online assessment is scored only on passing tests - so speed and edge cases decide it. Every other round is scored by a person who is listening to your reasoning, which is a different game and the one we mostly train for.

6. What a 45-minute round is made of

Concept

A round is not 45 minutes of coding. Here is where the time actually goes in a round that goes well.

minuteswhat is happeningwho is talking
0-3Intro, problem statementthem
3-8You clarify, restate, walk an exampleyou
8-12Brute force stated out loud, complexity givenyou
12-18Optimize: the pattern is named and justifiedboth
18-35Code ityou, narrating
35-42Dry run on a real example, edge casesyou
42-45Their questions, your questionsboth

Notice that coding is under half of it. Candidates who spend minutes 3-12 silently typing are the ones who get 'weak communication' on the feedback form even when the code works.

7. Four scores, and only one of them is the answer

Intuition

Interview feedback forms at most large companies rate roughly four things separately. This is why a working solution can still fail a round.

Problem solving
Did you get to a reasonable approach, and how much help did it take?
Coding
Is it clean, correct, and does it compile in your head?
Communication
Could they follow your reasoning without asking?
Verification
Did you test it yourself, or wait to be told it was broken?

Two candidates can produce the same final code and get opposite decisions. The difference is almost always communication and verification - the two you can improve fastest, and the two we drill every session.

McDowell, Cracking the Coding Interview, 6th ed. Ch. VI — the same four axes, described from the interviewer's side

8. Two of these are true. One is folklore.

Two truths and a lie

Sort each claim. Commit before you scroll - the wrong one is wrong for a reason worth having.

Sort into buckets

Which of these hold up, and which is folklore?

Holds up
Asking a clarifying question early is expected, not a weakness.; You can say 'I don't know the optimal approach yet' and still pass.; Stating a brute force before optimizing counts in your favour.
Folklore
Silence while you think is fine - they only score the final answer.
true
Everything the interviewer can observe is evidence. Clarifying, admitting uncertainty out loud, and offering a brute force all generate observable reasoning, which is what the form asks them to rate.
false
The interviewer cannot score what they cannot see. Long silence reads as being stuck, and it also blocks the hint they were ready to give you - hints are usually offered in response to a stated wrong idea, not to a quiet screen.

9. Three families of question

Concept

Every DSA question you get will sit in one of three families, and the family tells you how to spend the first five minutes.

familywhat it testsexample promptyour first move
ImplementationCareful coding, edge casesMerge two sorted listsClarify the edge cases, then just write it
Pattern recognitionDo you see the structureLongest substring without repeating charactersName the pattern out loud before coding
Open / design-flavouredTrade-offs and judgementDesign a data structure with O(1) insert, delete, getRandomAsk about scale and operations, then choose structures

Most of a 2-3 month prep window should go to the middle family - it is the biggest, and it is the one that transfers to problems you have not seen.

10. Which family is each of these?

Sorting

Take thirty seconds per prompt. You are not solving them.

Sort into buckets

Sort each prompt into its family.

Implementation
Reverse a linked list.; Determine whether a string of brackets is valid.
Pattern recognition
Longest substring without repeating characters.; Count the islands in a grid of 1s and 0s.
Open / design
Design an LRU cache.; Insert, delete, and get a random element, all in O(1).
impl
The approach is obvious the moment you read it; the difficulty is pointer discipline and edge cases (empty input, single element).
pat
The naive approach is obvious and too slow. The work is spotting which structure collapses the cost - a window, a visited set, a traversal.
open
There is no single right answer. You are being scored on which structures you combine and whether you can state the trade-off.

Valid parentheses is borderline - it needs a stack, but the stack is the first thing anyone thinks of. When a pattern is that obvious, the question has become an implementation question.

11. The problem statement is deliberately incomplete

Missing information

Interviewers under-specify on purpose. The unasked question is the trap.

Prompt: "Given an array of numbers, return the two that add up to a target."

Discussion prompt

Write down every question you would ask before touching the keyboard.

Hint: Think about the input's shape, its contents, and what 'return' means.

Answer:

  • Is the array sorted? Sorted unlocks two pointers and O(1) space.
  • Return indices or values? Indices forbid sorting in place.
  • Can numbers be negative, or repeat? Both break naive assumptions.
  • Is there exactly one answer, or should I return all pairs?
  • What if there is no answer - empty list, None, or raise?
  • How large can the array get? That sets the complexity target.

Six questions, about forty seconds, and every one of them changes the code you would write. This is the cheapest signal you can send.

12. Trap: starting to code because the silence felt bad

Trap

The trap

Interviewer: "Find the two numbers that sum to the target."

Immediately type for i in range(len(nums)):

Why: The silence felt like failure, so typing felt like progress.

Four minutes later, discover the array was sorted and indices were wanted

Why: Now the nested loop has to be thrown away and the clock has moved.

What the interviewer wrote down: jumped to code, did not scope the problem, needed a hint to find the sorted-array optimization.

The fix

Interviewer: "Find the two numbers that sum to the target."

Say: "Let me restate it and ask two things first."

Why: This buys thinking time and reads as method, not hesitation.

Ask about sorting and about indices-versus-values, then state the brute force and its cost out loud

Why: Both answers redirect the solution before any code exists, and the brute force gives you a working fallback you can always retreat to.

What the interviewer wrote down: scoped the problem, identified the quadratic baseline unprompted, chose the right structure.

13. The six-step loop - run it on every problem

Pattern

This is the procedure. It does not change with the problem, and it is what we will rehearse until it is automatic.

  1. Clarify. Restate the problem in your own words. Ask about input shape, duplicates, empties, and what to return when there is no answer.
  2. Example. Write one small concrete input and its expected output by hand. This catches misunderstandings that words hide.
  3. Brute force. State the obvious solution and its time and space cost out loud. Do not code it. You now have a floor and a working fallback.
  1. Optimize. Name the wasted work in the brute force, then name the structure that removes it. This is where the pattern gets chosen, and said aloud.
  2. Code. Write it while narrating what each block is for. Meaningful names, no clever one-liners.
  3. Test. Dry-run your own code on the example from step 2, then on the edges: empty, one element, all-duplicates, no answer.

Steps 1-4 should take about a third of the round. If you have said nothing for two minutes at any point, you are off the loop.

14. Why bother stating a brute force you will not write?

Explain it to yourself

Discussion prompt

Step 3 costs you about ninety seconds and produces code you are about to throw away. Argue for it - why is it worth the time?

Hint: Think about what happens if you run out of clock at minute 40.

Answer:

  • It proves the problem is understood, which is the thing the interviewer is least sure of at minute 8.
  • It gives a guaranteed fallback: a slow correct answer beats an unfinished fast one, every time.
  • It creates the optimization target. 'This is O(n squared) because I re-scan the rest of the array each time' names the waste, and naming the waste is what suggests the fix.
  • It converts a silent stretch into observable reasoning.

A candidate who says "brute force is quadratic, and the waste is the repeated inner scan" has already half-derived the hash map.

15. Check yourself: the first three minutes

Check

Answer before you look. This one is about method, not code.

Check your understanding

The interviewer finishes reading a problem you have never seen. What is the highest-value use of your next three minutes?

  • A. Start coding the first approach that comes to mind, and refine it as you go.
  • B. Restate the problem, ask about input shape and return value, and work one small example by hand. (correct)
  • C. Silently think until you have the optimal solution, then explain it all at once.
  • D. Ask the interviewer which algorithm they want you to use.

Answer: B

Why: Restating plus one hand-worked example catches almost every misunderstanding, and it does so while the clock is cheap. It also generates the test case you will dry-run at minute 38, so the three minutes are spent twice.

Why A tempts people
Coding before scoping is the single most common way rounds go wrong - you commit to a shape before knowing whether the input is sorted, can repeat, or can be empty.
Why C tempts people
Thinking is necessary, silence is not. The interviewer scores reasoning they can observe, and long silence also blocks the hint they were waiting to offer.
Why D tempts people
This hands the problem-solving score back to them. Asking about constraints is expected; asking for the method is asking for the answer.

16. Complexity: The Language of the Round

Section

Section 2

17. Before we build on it - what do you already have?

Warm-up

Discussion prompt

Without looking anything up, write the time complexity of each: (a) looking up a key in a Python dict, (b) x in my_list, (c) my_list.sort(), (d) my_list.insert(0, x), (e) my_set.add(x).

Hint: Two of these are the ones candidates get wrong under pressure.

Answer:

operationaverage costwhy
d[k]O(1)hash straight to a bucket
x in my_listO(n)linear scan, no index
my_list.sort()O(n log n)Timsort
my_list.insert(0, x)O(n)every later element shifts right
my_set.add(x)O(1)same hashing as dict

x in my_list versus x in my_set is the difference between a quadratic and a linear solution, and it is one character of Python. This is the most common accidental O(n squared) in interviews.

Python Wiki - TimeComplexity (per-operation costs for list, dict, set) list / dict / set — the authoritative per-operation table

18. Big-O measures growth, not speed

Concept

Big-O answers exactly one question: when the input gets 10 times bigger, what happens to the work? It says nothing about milliseconds on your laptop.

O(f(n)) — The shape of the growth curve as n gets large, with constant factors and lower-order terms dropped - because those stop mattering exactly when n stops being small.

complexityn = 10n = 1,000n = 1,000,000verdict at n = 10^6
O(log n)31020instant
O(n)101,0001,000,000fine
O(n log n)3310,00020,000,000fine
O(n^2)1001,000,00010^12hopeless
O(2^n)1,024astronomicalastronomicalhopeless past n = 30

Read the bottom two rows again. The gap between n log n and n^2 is not a tuning problem - at a million elements it is 2 x 10^7 operations against 10^12, which at a realistic ten million operations per second is a couple of seconds against more than a day.

19. Watch the brute force die

Scale up

count_pairs compares every element with every later element. Step through what happens as the input grows - these counts are exact, not estimates.

Step through it

At which step does this stop being a viable interview answer?

  1. n = 4 gives 6 pairs.
  2. n = 10 gives 45 pairs.
  3. n = 100 gives 4,950 pairs - 10x the input, 100x the work.
  4. n = 1,000 gives 499,500 pairs.
  5. n = 100,000 gives about 5 billion pairs, and the clock runs out.

The shape to remember: the input grew 10x, the work grew 100x. That squaring is what O(n^2) means, and it is why the constraint line in the problem statement is the most important line in it.

20. Decode it symbol by symbol

Notation

Candidates say 'oh of en log en' fluently and cannot say where the log came from. Read each piece.

Annotate

On: \( T(n) = O(n \log n) \)

  • T(n) is the running time as a function of the input size n - not a number of seconds.
  • The O means 'grows no faster than', an upper bound. It is a ceiling on the shape, not a measurement.
  • The log n factor almost always means the work halves each round - a binary search, or the depth of a balanced divide-and-conquer.
  • The n multiplying it means that halving happens for every element, or that each of the log n levels does n work. Sorting is the canonical example: log n levels of merging, n work per level.

If you can say where the log comes from, you can derive the complexity of a solution you have never analysed. If you cannot, you are reciting.

21. The constraints tell you the answer's shape

Concept

Every well-posed problem states a bound on n. That bound is the interviewer telling you which complexity they will accept.

n up tobudgettarget complexitytypical shape
10-20anythingO(2^n) or O(n!)backtracking, permutations
~500generousO(n^3)triple loop, some DP
~5,000moderateO(n^2)DP table, all pairs
~10^5tightO(n log n) or O(n)sort, heap, window, hash
~10^9extremeO(log n) or O(1)binary search, math, formula

Read the table right to left in the interview: you are given n up to 10^5, so you need at most n log n, so sorting is affordable but a nested loop is not. That reasoning, said out loud, is worth real points.

22. Pick the target, do not solve anything

Discrimination

For each constraint, choose the complexity you should be aiming for. No code.

Sort into buckets

Sort each constraint into the complexity you should target.

O(n log n) or better
n up to 100,000; you may return the answer in any order.
O(n^2) is acceptable
n up to 2,000; count pairs satisfying a condition.
Exponential is expected
n up to 16; find the best subset of items.
O(log n)
The array is sorted and n can reach 10^9 conceptually.
nlogn
At 10^5, a quadratic solution is 10^10 operations. Sorting, hashing, a heap or a sliding window all fit inside the budget.
n2
At n = 2,000 a quadratic solution is 4 million operations - comfortably fast. This bound is usually a hint that a DP table or an all-pairs scan is the intended answer.
exp
n up to about 20 is the universal signal for backtracking, subset enumeration or bitmask DP. 2^16 is only 65,536.
log
A huge sorted domain with no way to touch every element means the answer is binary search - on the array, or on the answer itself.

23. Mark where the time actually goes

Cost model

The trap is not the loop you can see. Read this and find the expensive line.

Annotate

  • Line 3: one pass over n elements. This is the O(n) everyone sees.
  • Line 4 is the real cost. n in seen scans a list linearly, so it is O(n) by itself - inside an O(n) loop. Total: O(n^2).
  • Line 6 is O(1) amortized. Appending is not the problem.
  • The fix is on line 2: seen = set(). Membership becomes O(1) and the whole function becomes O(n). One word.

This function is the single most common accidental quadratic in Python interviews, and it is invisible unless you cost the operations inside the loop rather than counting the loops.

24. Cost a function line by line

Worked example

Interviewers ask 'what is the complexity of what you just wrote?' every single round. Here is the procedure, on a function with two loops that are not nested and one that is.

def summarize(nums):
    total = 0
    for n in nums:            # A
        total += n
    nums_sorted = sorted(nums)  # B
    pairs = 0
    for i in range(len(nums)):  # C
        for j in range(i + 1, len(nums)):
            pairs += 1
    return total, nums_sorted[0], pairs

Cost each region on its own, in terms of n

Why: Regions that run one after another are added, not multiplied - the sequencing is what most candidates get wrong first.

regionwhat it doescost
Aone pass, O(1) work per elementO(n)
BTimsortO(n log n)
Cevery pair (i, j) with j > iO(n^2)

Add them, then keep only the fastest-growing term

Why: O(n) + O(n log n) + O(n^2) = O(n^2). Lower-order terms vanish because at large n the square dwarfs the rest.

Check the space separately

Why: sorted() allocates a new list of n elements, so space is O(n) even though total and pairs are O(1). Interviewers ask for both and candidates volunteer only time.

Verify against the real counts

Why: Region C on n = 1,000 executes 499,500 times, which is n(n-1)/2 - exactly the quadratic the analysis predicted.

25. Trap: two loops means quadratic

Trap

The trap

for n in nums:
    total += n
for n in nums:
    biggest = max(biggest, n)

Say "two loops over nums, so it's O(n^2)"

Why: Counting loops instead of counting the work each element causes.

The real cost is O(n) + O(n) = O(2n) = O(n). Sequential loops add.

The fix

for i in range(len(nums)):
    for j in range(len(nums)):
        check(nums[i], nums[j])

Ask: for each element of the outer loop, how much work happens inside?

Why: n elements, each causing n units of inner work, so the costs multiply: O(n^2). Nesting multiplies, sequencing adds.

shapen = 1,000 iterationscost
two loops in sequence1,000 + 1,000 = 2,000O(n)
one loop inside another1,000 x 1,000 = 1,000,000O(n^2)

The test is not how many for keywords are on screen - it is whether one loop runs inside another.

26. Space complexity, the half nobody volunteers

Concept

Every complexity answer has two halves. Saying only the time half is a missed point in almost every round.

what you allocatespacenote
A few counters and pointersO(1)loop variables are free
A hash map of every elementO(n)the usual price of a hash solution
A sorted copy via sorted()O(n)list.sort() in place is O(1) extra
A recursion stack of depth dO(d)for a tree, that is the height
A DP table of n by mO(n*m)often reducible to one row

Say both halves without being asked: "O(n) time, O(n) extra space - and I can get space down to O(1) if I'm allowed to sort in place." That sentence is a trade-off, which is what the design half of the score is looking for.

27. Check yourself: cost this function

Check

Read it carefully. The answer is not the number of loops.

Check your understanding

What is the time complexity of first_duplicate, where seen is a list and the membership test is if n in seen?

  • A. O(n) - it is a single pass over the input.
  • B. O(n^2) - the linear membership test runs inside the loop. (correct)
  • C. O(n log n) - membership testing is logarithmic.
  • D. O(1) - it returns early on the first duplicate.

Answer: B

Why: n in seen on a list is a linear scan, so each of the n iterations can do up to n comparisons: O(n^2). Changing seen to a set makes membership O(1) on average and the whole function O(n) - a one-word fix worth naming out loud in a round.

Why A tempts people
Counts the visible loop but not the hidden scan inside it. Costing the operations in the body, not just the loops, is the whole skill here.
Why C tempts people
Logarithmic membership would require a sorted structure and a binary search. Python lists give you neither for free.
Why D tempts people
Early return improves the best case only. Complexity questions in interviews mean the worst case unless stated otherwise - and the worst case here has no duplicate at all.

28. The Eight Patterns Behind Most Questions

Section

Section 3

29. The clue-to-pattern reflex

Concept

Interview problems announce themselves. This table is the reflex we are building - phrase on the left, structure on the right.

the prompt says...reach fortypical cost
"has it appeared before", "count of each"hash map / setO(n)
"sorted array", "pair that sums to"two pointersO(n)
"substring", "subarray", "consecutive"sliding windowO(n)
"sorted", "find the boundary", "minimum such that"binary searchO(log n)
"matching", "nesting", "next greater"stackO(n)
"shortest path in steps", "connected region"BFS / DFSO(V + E)
"top K", "K largest", "running median"heapO(n log k)
"number of ways", "maximum over choices"dynamic programmingO(n) to O(n*m)

Saying "the word substring plus a constraint on characters makes me think sliding window" is not guessing. It is the reasoning interviewers most want to hear, because it is what transfers to the next problem.

30. Match the clue to the structure

Matching

No solving. Just pair each phrase with the structure it should trigger.

Match the pairs

Match each phrase from a problem statement to the pattern it signals.

  • l1. "...the array is sorted..."
  • l2. "...longest substring such that..."
  • l3. "...the k most frequent..."
  • l4. "...every valid combination..."
  • l5. "...has this value appeared before..."
  • r1. Two pointers or binary search
  • r2. Sliding window
  • r3. Heap of size k
  • r4. Backtracking
  • r5. Hash set

Why: "Sorted" is the loudest word in any prompt - it means you can move two pointers inward or halve the search space, and it is free information you paid nothing for. "Longest ... such that" is a window growing and shrinking. "K most" is a heap of size k, not a full sort. "Every combination" is exhaustive search with pruning. "Appeared before" is membership, which is a set.

31. Pattern 1 - hash map: buy time with memory

Concept

The hash map is the most common single idea in interview questions. It turns "have I seen this?" from a scan into a lookup.

# The whole pattern, in three lines
seen = {}                 # value -> where I saw it
for i, n in enumerate(nums):
    if n in seen:         # O(1) average, not a scan
        return seen[n], i
    seen[n] = i
questionwithout a hash mapwith a hash map
Is x in the collection?O(n) scanO(1) average
How many times does x occur?O(n) per queryO(1) after one O(n) pass
Where did I last see x?O(n) scanO(1) average
Total for n queriesO(n^2)O(n)

The cost is O(n) extra space. That is the trade you should say out loud: "I'll spend O(n) memory to drop the time from quadratic to linear."

Python Wiki - TimeComplexity (per-operation costs for list, dict, set) dict — average-case O(1) get and set

32. Predict the output before you read the code twice

Prediction

Commit to an answer. This is the shape of a dozen real questions.

def first_repeat(nums):
    seen = set()
    for n in nums:
        if n in seen:
            return n
        seen.add(n)
    return None

print(first_repeat([5, 3, 9, 3, 5]))

Predict first

What does this print - and how many elements does it look at before it stops?

  • 5, after looking at all 5 elements
  • 3, after looking at 4 elements
  • 3, after looking at all 5 elements
  • None

Correct: 3, after looking at 4 elements.

It prints 3, and it never reads the last element.

Why: The loop returns on the first value that has already been seen, not on the value that repeats earliest in the array. 5 repeats too, but its second occurrence is at index 4, while 3's second occurrence is at index 3 - so the function stops there and never reads the final 5. Interviewers love this distinction because 'first duplicate' is ambiguous until you ask.

indexnseen before this stepaction
05{}add 5
13{5}add 3
29{5, 3}add 9
33{5, 3, 9}return 3 - index 4 is never reached

"First duplicate" is ambiguous: first by second occurrence (what this code does) or first by first occurrence (which would be 5). Ask. This is a real clarifying question, not a pedantic one.

33. The most-asked question there is: Two Sum

Worked example

Given nums and a target, return the indices of the two numbers that add to the target. This is the single most commonly assigned interview problem, and its one-pass solution is the template for a dozen others.

State the brute force and its cost first

Why: Every pair is O(n^2) time and O(1) space. It is correct, and it is the fallback you keep in your pocket.

Name the waste: the inner loop re-asks a question the outer loop already answered

Why: For each n we are searching the rest of the array for target - n. That is a membership question, and membership is what a hash map does in O(1).

def two_sum(nums, target):
    seen = {}                     # value -> index
    for i, n in enumerate(nums):
        need = target - n
        if need in seen:
            return [seen[need], i]
        seen[n] = i
    return []

Trace it on nums = [2, 7, 11, 15], target = 26

Why: The table below is the actual execution - each row is one iteration, showing what seen held before the check.

inneed = 26 - nneed in seen?seen before this step
0224no{}
1719no{2: 0}
21115no{2: 0, 7: 1}
31511yes{2: 0, 7: 1, 11: 2}

Verify the returned indices against the input

Why: It returns [2, 3]; nums[2] + nums[3] = 11 + 15 = 26. Correct, in one pass, O(n) time and O(n) space.

The move worth stealing: store what you have seen keyed by what a future element would need. That single idea also solves subarray-sum-equals-k, pairs-with-difference-k, and contains-duplicate-within-distance-k.

34. Trap: sorting first when the question wants indices

Trap

The trap

"The array is unsorted, so let me sort it and use two pointers."

nums.sort()          # destroys the original positions
lo, hi = 0, len(nums) - 1
# ... two-pointer scan finds the values 11 and 15

Return the two indices found after sorting

Why: Those indices point into the sorted array. The caller asked about the original one, so the answer is wrong even though the values are right.

originalafter sort
array[15, 2, 11, 7][2, 7, 11, 15]
index of 1122
index of 1503

The fix

Ask at minute 2: "Do you want the indices or the values?"

If indices: use the hash map, leave the array untouched

Why: O(n) time, O(n) space, positions preserved - and it beats the sorted approach's O(n log n) anyway.

If values, and the array is already sorted: two pointers

Why: O(n) time and O(1) space, which is strictly better than hashing - so the right answer genuinely depends on the clarifying question.

asked forarray statebest approachcost
indicesunsortedhash mapO(n) time, O(n) space
valuessortedtwo pointersO(n) time, O(1) space
valuesunsortedsort, then two pointersO(n log n) time

35. Pattern 2 - two pointers: walk from both ends

Concept

When the data is sorted, or the answer depends on a pair, two indices moving toward each other replace a nested loop with a single pass.

lo, hi = 0, len(a) - 1
while lo < hi:
    total = a[lo] + a[hi]
    if total == target:
        return lo, hi
    if total < target:
        lo += 1          # need more: only the left end can grow
    else:
        hi -= 1          # need less: only the right end can shrink
variantpointers startthey moveexample question
convergingboth endsinwardtwo-sum on a sorted array
same directionboth at 0fast one leadsremove duplicates in place
fast / slowboth at headfast moves 2xlinked-list cycle detection
read / writeboth at 0write lags readcompact an array in place

The reason it is O(n) and not O(n^2): each pointer only ever moves in one direction, so together they take at most n steps in total.

36. What stays true the whole time?

Invariant

Step through a palindrome check on "Race car!" and watch for the thing that never changes. The cleaned string is racecar.

Step through it

What is true about the characters OUTSIDE lo and hi at every single step?

  1. Outside the pointers: nothing checked yet, and nothing has failed.
  2. Outside the pointers: r matches r.
  3. Outside the pointers: ra matches ar.
  4. Outside the pointers: rac matches car - the whole string is a palindrome.

The invariant: everything outside lo and hi has already been verified to match. When the pointers meet, 'outside' is the whole string, so the answer is True. That sentence is the proof of correctness, and being able to state it is what separates a memorized loop from an understood one.

37. Valid palindrome, two pointers

Worked example

Given a string, ignore case and non-alphanumeric characters, and decide whether it reads the same forwards and backwards.

Clarify before coding

Why: What counts as a character? Are digits included? Is the empty string a palindrome? (Conventionally yes - and saying so is worth a point.)

def is_palindrome(s):
    t = [c.lower() for c in s if c.isalnum()]
    lo, hi = 0, len(t) - 1
    while lo < hi:
        if t[lo] != t[hi]:
            return False
        lo += 1
        hi -= 1
    return True

Trace it on "Race car!"

Why: After cleaning, t = racecar. Each row is one pass of the while loop, taken from the real run.

passlot[lo]hit[hi]equal?
10r6ryes
21a5ayes
32c4cyes
-3-3-lo is not < hi, loop ends

Verify: the loop ended without returning False, so the result is True

Why: is_palindrome("Race car!") returns True and is_palindrome("hello") returns False on the first comparison, h against o.

Cost: O(n) time, and O(n) space because of the cleaned list. If asked for O(1) space, skip non-alphanumeric characters in place with two while loops inside the main one - a natural follow-up, so have the answer ready.

38. Pattern 3 - a window you resize, not rebuild

Intuition

The words substring, subarray, and consecutive all mean the same thing structurally: a contiguous stretch with two ends.

The naive solution rebuilds every stretch from scratch: n starting points times n lengths, and O(n) work to evaluate each - cubic, or quadratic if you are careful. The window keeps one stretch and edits its ends.

approachwork per new stretchtotal
rebuild each substringO(n) to re-scan itO(n^2) or worse
slide the windowO(1) to add one char and drop someO(n)

Mental image: you are not printing a new photo each time, you are dragging the edges of a crop box. Everything already inside the box stays valid.

39. Two flavours of window

Concept

Which one you need is decided by a single question: is the window's size given to you, or is it whatever the constraint allows?

# fixed size k: one in, one out
window = sum(nums[:k])
best = window
for i in range(k, len(nums)):
    window += nums[i] - nums[i - k]
    best = max(best, window)

# variable size: grow always, shrink while the rule is broken
start = 0
for end in range(len(nums)):
    add(nums[end])
    while broken():
        remove(nums[start])
        start += 1
clue in the promptflavourexample
"...of size k", "of length k"fixedmax sum subarray of size k
"longest ... such that"variable, maximizelongest substring with no repeats
"shortest ... containing"variable, minimizeminimum window substring
"at most k distinct"variable with a counterlongest substring with k distinct characters

Both are O(n) because every index enters the window once and leaves at most once - the inner while does not make it quadratic, and being able to say why is a frequent follow-up question.

40. The line that makes it a window

Fill the middle

Everything here is standard except one line. Fill it in - this is the line that decides whether the window is correct.

Fill in the blanks

Complete the two blanks.

def longest_unique(s):
last = last[ch] + 1 # char -> most recent index
start = 0 # left edge of the window
best = 0
for i, ch in enumerate(s):
if ch in last and last[ch] >= start:
start = i - start + 1
last[ch] = i
best = max(best, ___)
return best

Why: start = last[ch] + 1 jumps the left edge to just past the previous copy of this character, which is what makes the whole scan O(n) - creeping forward one index at a time would re-check characters and make it quadratic. The guard last[ch] >= start matters just as much: without it, a character last seen before the window would drag start backwards and shrink the answer. And the window length is i - start + 1 because both ends are inclusive - the off-by-one here is the most common bug in this pattern.

41. Longest substring without repeating characters

Worked example

A top-five most-asked question, and the cleanest example of a variable window. Input: "abcabcbb".

Clarify first

Why: Is the input ASCII or unicode? Do we return the length or the substring itself? Is an empty string valid input?

Brute force, out loud, not typed

Why: Check every substring for uniqueness: O(n^2) substrings, O(n) to check each, so O(n^3). The waste is re-checking characters that were already known unique.

def longest_unique(s):
    last = {}
    start = 0
    best = 0
    for i, ch in enumerate(s):
        if ch in last and last[ch] >= start:
            start = last[ch] + 1
        last[ch] = i
        best = max(best, i - start + 1)
    return best

Trace every index on "abcabcbb" - this is the real execution

Why: Watch start jump rather than creep, and watch best only ever grow.

ichstart afterwindowlengthbest
0a0a11
1b0ab22
2c0abc33
3a1bca33
4b2cab33
5c3abc33
6b5cb23
7b7b13

Verify against the edge cases

Why: "abcabcbb" gives 3, "bbbbb" gives 1, "pwwkew" gives 3 (wke, not the non-contiguous pwke). All three match the real output.

Cost: O(n) time - each index is visited once - and O(min(n, alphabet)) space for the map. Say the alphabet bound; it is the kind of precision that reads as experience.

42. Trap: recomputing the window instead of updating it

Trap

The trap

best = 0
for i in range(len(nums) - k + 1):
    best = max(best, sum(nums[i:i + k]))

Call sum() on a fresh slice for every window

Why: It looks like one loop, so it looks linear. But nums[i:i+k] copies k elements and sum adds k of them, inside a loop that runs n times.

nkoperations
1,000100about 100,000
100,0001,000about 100,000,000 - times out

The fix

window = sum(nums[:k])
best = window
for i in range(k, len(nums)):
    window += nums[i] - nums[i - k]
    best = max(best, window)

Add the entering element, subtract the leaving one

Why: The window's value is repaired in O(1) instead of rebuilt in O(k). This is the whole point of the pattern.

nkoperations
1,000100about 1,000
100,0001,000about 100,000 - instant

43. Pattern 4 - binary search: throw away half, every time

Concept

Any time the data is sorted, or the answer is monotonic ("if 7 works then 8 works"), you can halve the space instead of scanning it.

def bsearch(a, x):
    lo, hi = 0, len(a) - 1
    while lo <= hi:               # inclusive range [lo, hi]
        mid = (lo + hi) // 2
        if a[mid] == x:
            return mid
        if a[mid] < x:
            lo = mid + 1          # x must be to the right
        else:
            hi = mid - 1          # x must be to the left
    return -1

Trace it: find 23 in [2, 5, 8, 12, 16, 23, 38, 56]

Why: Eight elements, and it finishes in two comparisons - that is the log.

passlohimida[mid]decision
10731212 < 23, go right: lo = 4
247523found at index 5

The version worth memorizing is not this one but its cousin: the smallest index where a condition first becomes true. Most real binary-search interview questions - rotated array, first bad version, minimum capacity - are that boundary search wearing a costume.

44. This binary search hangs. Find the line.

Error analysis

It compiles, it passes the happy path, and on some inputs it never returns. Read it as if reviewing a colleague's pull request.

Annotate

  • Line 8 is the bug: lo = mid does not shrink the range when mid == lo. With lo = 0, hi = 1, mid is 0, and if a[0] < x we set lo = 0 again - the loop spins forever.
  • The fix is lo = mid + 1. Since a[mid] has already been compared and rejected, excluding it is safe, and it guarantees the range strictly shrinks.
  • hi = len(a) with while lo < hi is a half-open range, which is fine on its own - but it must be paired consistently. Mixing hi = len(a) - 1 with lo < hi silently drops the last element.
  • hi = mid is correct in the half-open convention because hi is exclusive. Pick one convention per problem and say which you are using out loud.

The rule that prevents every version of this bug: each iteration must make the range strictly smaller. Before you write the loop, say which of lo/hi is inclusive - most binary-search bugs are convention drift, not logic errors.

45. Now write it with the scaffolding removed

Faded example

The boundary search - find the first index where the condition holds. Fill the three blanks.

Fill in the blanks

Complete the boundary search.

def first_true(a, condition):
lo, hi = 0, len(a) - 1
answer = -1
while lo <= hi:
mid = (lo + hi) // 2
if condition(a[mid]):
answer = mid
hi = mid - 1 # a better answer may still be left of mid
else:
lo = mid + 1
return answer

Why: This template answers 'first bad version', 'minimum eating speed', 'smallest divisor', 'search insert position' and every rotated-array variant - which is most of the binary-search questions asked. The key difference from plain search is that finding a match does not return: it records the match and keeps shrinking toward the left edge, so the loop always ends at the first index that works.

46. Pattern 5 - stack: the most recent unfinished thing

Concept

A stack is the right structure whenever the thing you need next is the thing you saw most recently: nesting, matching, undo, and 'next greater element'.

def valid(s):
    pairs = {')': '(', ']': '[', '}': '{'}
    stack = []
    for ch in s:
        if ch in '([{':
            stack.append(ch)
        else:
            if not stack or stack.pop() != pairs[ch]:
                return False
    return not stack

Trace it on "([{}])"

Why: Every closer must match the most recent unmatched opener - which is exactly the top of the stack.

charactionstack after
(push['(']
[push['(', '[']
{push['(', '[', '{']
}pop { - matches['(', '[']
]pop [ - matches['(']
)pop ( - matches[]

Two edge cases carry the whole question, and both are on line 8 and line 10: a closer arriving when the stack is empty ("())"), and leftovers on the stack at the end ("(("). Candidates who only check the middle case get one of the hidden tests wrong.

47. Commit: which of these are valid?

Prediction

Four inputs to the bracket checker above. Decide all four before revealing.

Predict first

Which of these return True: "([{}])", "(]", "((", "())("?

  • Only the first
  • The first and the last
  • The first and third
  • All but the second

Correct: Only the first.

Why: "([{}])" is properly nested. "(]" pops ( and compares it with the [ that ] requires - mismatch. "((" never fails during the loop but ends with two unmatched openers, so the final return not stack catches it. "())(" fails on the third character, where a closer arrives with an empty stack. Those last two are exactly the cases a rushed implementation misses.

inputresultwhich guard catches it
([{}])True-
(]Falsestack.pop() != pairs[ch]
((Falsereturn not stack at the end
())(Falsenot stack when a closer arrives

48. Monotonic stack: Daily Temperatures

Worked example

Given daily temperatures, return for each day how many days until a warmer one. This is the standard 'next greater element' question, and it is the one that separates candidates who know a stack from candidates who know when.

Brute force first, out loud

Why: For each day, scan forward until a warmer day: O(n^2). The waste is that a cold day gets re-scanned by every earlier day.

The insight: keep the days still waiting for an answer on a stack

Why: When today is warmer than the day on top of the stack, today is that day's answer - and we can pop it, permanently.

def daily(temps):
    out = [0] * len(temps)
    stack = []                       # indices, temperatures decreasing
    for i, t in enumerate(temps):
        while stack and temps[stack[-1]] < t:
            j = stack.pop()
            out[j] = i - j
        stack.append(i)
    return out

Trace it on [73, 74, 75, 71, 69, 72, 76, 73]

Why: The stack always holds indices whose temperatures decrease from bottom to top - that is the invariant that makes the pops safe.

itempindices poppedstack afterout so far
073-[0][0,0,0,0,0,0,0,0]
1740[1][1,0,0,0,0,0,0,0]
2751[2][1,1,0,0,0,0,0,0]
371-[2,3]unchanged
469-[2,3,4]unchanged
5724, 3[2,5][1,1,0,2,1,0,0,0]
6765, 2[6][1,1,4,2,1,1,0,0]
773-[6,7]final: [1,1,4,2,1,1,0,0]

Verify the interesting entries by hand

Why: Day 2 (75 degrees) waits until day 6 (76): 6 - 2 = 4. Correct. Days 6 and 7 never get a warmer day, so they keep 0 - which is why the output array is initialised with zeros rather than left empty.

Why it is O(n) despite the inner while: every index is pushed once and popped at most once, so the total number of pops across the whole run is at most n. That amortized argument is the follow-up question, every time.

49. Pattern 6 - BFS and DFS are the same code with one line changed

Concept

Every tree question and every grid question is one of these two walks. The only structural difference is which end of the pending collection you take from.

from collections import deque

def bfs(start, neighbors):
    seen = {start}
    q = deque([start])
    while q:
        node = q.popleft()        # FIFO -> level by level
        for nxt in neighbors(node):
            if nxt not in seen:
                seen.add(nxt)
                q.append(nxt)

def dfs(node, neighbors, seen):
    seen.add(node)
    for nxt in neighbors(node):   # LIFO via the call stack
        if nxt not in seen:
            dfs(nxt, neighbors, seen)
BFSDFS
structurequeue (deque.popleft)recursion or an explicit stack
visitsnearest first, level by levelone branch to the bottom first
memoryO(width of the level)O(depth)
shortest path?yes, on unweighted graphsno
typical usemin steps, level orderconnected regions, cycles, paths

The seen set is not optional and is not an optimization - without it a graph with any cycle loops forever. Forgetting it on a grid problem, where every cell has a path back, is the most common graph mistake in interviews.

50. BFS or DFS? Fill in the table.

Trade off

Complete the blanks from what you just read. Both are correct answers to different questions.

Comparison matrix

questionBFSDFS
Fewest moves on an unweighted graphyesno - the first path found may be long
Extra memory usedO(width)O(depth)
Deep, narrow tree (a linked list, say)safe - the level is tinyrisky - recursion depth
Count connected regions in a gridworks fineworks fine, usually shorter code

The one that decides real rounds: "shortest" or "fewest" plus "unweighted" means BFS. If edges have weights, neither is right and the answer is Dijkstra - a good sentence to have ready.

51. Number of Islands - a grid is just a graph

Worked example

Given a grid of "1" (land) and "0" (water), count the connected regions of land. Grids show up constantly, and the trick is seeing that a cell's neighbours are its four adjacent cells.

Clarify

Why: Is diagonal adjacency connected? (Usually no.) May I modify the input grid? That answer decides whether you need a separate visited set.

def num_islands(grid):
    rows, cols = len(grid), len(grid[0])
    count = 0

    def sink(r, c):
        if r < 0 or c < 0 or r >= rows or c >= cols or grid[r][c] != '1':
            return
        grid[r][c] = '0'                 # mark visited by sinking it
        sink(r + 1, c); sink(r - 1, c)
        sink(r, c + 1); sink(r, c - 1)

    for r in range(rows):
        for c in range(cols):
            if grid[r][c] == '1':
                count += 1               # a new island starts here
                sink(r, c)
    return count

Trace the outer scan on a 4x5 grid

Why: Rows: 11000, 11000, 00100, 00011. The scan only ever starts a flood fill on land it has not already sunk.

cell reached by the scangrid value thereactioncount
(0,0)1new island, sink 4 cells1
(0,1) ... (1,1)0 (already sunk)skip1
(2,2)1new island, sink 1 cell2
(3,3)1new island, sink 2 cells3
everything else0skip3

Verify by eye and by cost

Why: Three regions: the 2x2 block, the single cell, and the pair - the function returns 3. Every cell is visited a constant number of times, so it is O(rows x cols) time and O(rows x cols) worst-case recursion depth.

Mention the recursion depth unprompted: on a 1000x1000 all-land grid, recursive DFS can exceed Python's stack limit, and the fix is an explicit stack or BFS. Naming a failure mode of your own solution scores well.

52. Order the moves for any grid problem

Ranking

These five steps solve nearly every grid question. Put them in the order you would actually perform them.

Put in order

Order the steps for attacking a grid traversal problem.

  1. Ask whether diagonals count and whether the grid may be modified.
  2. Decide BFS or DFS from whether the question asks for a shortest path.
  3. Write the bounds check and the visited check as the first lines of the traversal.
  4. Walk the traversal on a 3x3 example by hand.
  5. State the cost as O(rows x cols) and name the recursion-depth risk.

Why: Clarifying comes first because the answers change the code. The BFS/DFS choice comes next because it decides the skeleton. The bounds-and-visited guard is written before the recursive calls - writing it last is how infinite recursion gets shipped. Hand-walking a small example catches sign errors in the neighbour offsets before the interviewer does, and the cost statement closes the round with the analysis they were going to ask for anyway.

53. Pattern 7 - DP is just refusing to solve the same thing twice

Intuition

Dynamic programming has a fearsome name and a small idea: some recursive problems ask the same sub-question over and over, so you write the answer down the first time.

Naive fib(40) makes about 205 million calls, and roughly 15 million of them are recomputing fib(5) alone. Storing each answer once turns an exponential tree into a linear walk - the algorithm did not get cleverer, it just stopped forgetting.

plain recursionwith memobottom-up
calls for fib(40)about 205 million4040
timeO(2^n)O(n)O(n)
spaceO(n) stackO(n) + stackO(1) if you keep two values

Two questions identify a DP problem: is the answer built from answers to smaller versions of the same problem, and do those smaller versions repeat? If yes to both, write the recurrence before any code.

54. Two ways to write the same DP

Concept

Top-down memoization and bottom-up tabulation compute identical values. Write whichever makes the recurrence obvious - then say the other exists.

from functools import cache

@cache                              # top-down: recursion + a memo
def climb(n):
    if n <= 2:
        return n
    return climb(n - 1) + climb(n - 2)

def climb_iter(n):                  # bottom-up: no recursion at all
    a, b = 1, 1
    for _ in range(n - 1):
        a, b = b, a + b
    return b
top-down (memo)bottom-up (table)
how you write itthe recurrence, literallya loop filling an array
riskrecursion depth on large nnone
spaceO(n) memo + O(n) stackO(n), often reducible to O(1)
best whenthe state space is sparseevery state is needed anyway

In an interview, write the memoized version first - it is closer to the recurrence you just said aloud, so it is easier to justify - then offer the iterative rewrite as the optimization. That sequence is itself a signal.

55. Climbing Stairs, and the recurrence behind it

Worked example

You can climb 1 or 2 steps at a time. How many distinct ways to reach step n? This is the DP everyone is asked first, and its recurrence is the one to be able to derive cold.

Derive the recurrence instead of recalling it

Why: The last move onto step n was either a 1-step from n-1 or a 2-step from n-2. Those two sets of paths are disjoint and cover everything, so ways(n) = ways(n-1) + ways(n-2).

Pin the base cases with tiny hand cases, not intuition

Why: ways(1) = 1 (one step). ways(2) = 2 (1+1, or 2). Getting these wrong shifts the whole sequence, and it is the most common error on this problem.

def climb(n):
    a, b = 1, 1        # a = ways(n-2), b = ways(n-1)
    for _ in range(n - 1):
        a, b = b, a + b
    return b

Trace it to n = 6

Why: Each row is one pass of the loop, from the real execution.

passabmeaning
start11ways(1) = 1
112ways(2) = 2
223ways(3) = 3
335ways(4) = 5
458ways(5) = 8
5813ways(6) = 13

Verify against a hand count

Why: For n = 4 the paths are 1111, 112, 121, 211, 22 - five of them, matching the table. The sequence 1, 2, 3, 5, 8, 13 is Fibonacci shifted by one, which is a good sanity check to mention.

Cost: O(n) time and O(1) space because only two values are kept. House Robber, Min Cost Climbing Stairs, and Decode Ways are all this same loop with a different line inside.

56. Given the recurrence, name the problem

Reverse engineer

Interviewers often hand you a recurrence and ask what it computes. Work backwards from this one.

Fill in the blanks

Fill in the missing half of the recurrence, then name the problem.

best[i] = max(best[i - 1], best[i - 2] + nums[i]) -> the problem is House Robber

Why: At each index there are exactly two choices: skip i, keeping best[i-1]; or take i, which forbids i-1 and so builds on best[i-2]. The max of those two is the answer at i. Recognising a recurrence by its shape - a choice between skip and take - is far more transferable than remembering the story about houses, and the same shape appears in Delete and Earn and in Maximum Sum of Non-Adjacent Elements.

57. Pattern 8 - heap: when you want the top K, not the order

Concept

Sorting to answer "the k largest" costs O(n log n). A heap of size k costs O(n log k), and when k is small that is effectively linear.

import heapq
from collections import Counter

def top_k(nums, k):
    counts = Counter(nums)                 # O(n)
    return [n for n, _ in heapq.nlargest(
        k, counts.items(), key=lambda p: p[1])]

print(top_k([1, 1, 1, 2, 2, 3], 2))        # -> [1, 2]
approachtimewhen to prefer it
sort everythingO(n log n)you need the full order anyway
heap of size kO(n log k)k is much smaller than n
bucket by frequencyO(n)counts are bounded by n - the optimal answer
heapq.heappush/heappopO(log n) eacha stream, where n is unknown

Two Python facts worth knowing cold: heapq is a min-heap, so push negated values for a max-heap; and Counter(nums).most_common(k) exists, but interviewers usually want to hear the heap or bucket reasoning behind it.

Python 3 documentation - heapq heapq.nlargest — the size-k heap used here

58. The whole toolkit on one card - fill the blanks

Comparison

Complete this from memory. This card is the thing to be able to reproduce before your first real interview.

Comparison matrix

patternsignature cluetypical cost
Hash map"seen before", "count of each"O(n)
Two pointerssorted array, or a pairO(n)
Sliding window"substring", "consecutive"O(n)
Binary searchsorted, or monotonic conditionO(log n)
Stacknesting, or next greaterO(n)
BFS / DFSgrid, tree, "connected", "fewest steps"O(V + E)
Heap"top k", "k largest"O(n log k)
Dynamic programming"number of ways", "maximum over choices"O(n) to O(n*m)

59. Every option is plausible. Knock them out.

Elimination

Prompt: "Given an unsorted array of up to 100,000 integers, return the length of the longest run of consecutive integers - for example [100, 4, 200, 1, 3, 2] gives 4, for 1,2,3,4."

Eliminate the wrong options

Which approach survives?

  • A. Sort the array, then scan for the longest consecutive run.
  • B. Put everything in a set; for each value that has no predecessor in the set, walk upward counting.
  • C. For each element, scan the array for element+1, then element+2, and so on.
  • D. Dynamic programming over the array indices.

Survives elimination: B

Why: The set solution is O(n): each value is the start of a run only if value - 1 is absent, and the upward walk from each start visits each value at most once across the whole run, so the total work is linear despite the nested loop. The reasoning to say out loud is 'n up to 100,000 means I want linear or n log n; sorting gives me n log n immediately, and a hash set gets me to linear by replacing the order with membership.'

60. The Questions That Keep Coming Back

Section

Section 4

61. The short list, part one

Concept

These are the problems that appear again and again across company question banks and every widely used study list. Learn them as instances of patterns, in the third column - that is the part that transfers.

problempatterntarget cost
Two Sumhash mapO(n)
Contains Duplicatehash setO(n)
Valid Anagramcounting with a hash mapO(n)
Group Anagramshash map keyed by sorted wordO(n k log k)
Best Time to Buy and Sell Stockone pass, running minimumO(n)
Maximum Subarray (Kadane)one pass, running bestO(n)
Valid ParenthesesstackO(n)
Merge Two Sorted Liststwo pointersO(n + m)

NeetCode - pattern-grouped practice list — a widely used pattern-grouped list, if you want a practice order

62. The short list, part two

Concept

problempatterntarget cost
Longest Substring Without Repeating Characterssliding windowO(n)
Binary Search / Search Insert Positionbinary searchO(log n)
Reverse Linked Listthree pointersO(n)
Linked List Cyclefast and slow pointersO(n)
Invert / Maximum Depth of Binary TreeDFSO(n)
Binary Tree Level Order TraversalBFSO(n)
Number of IslandsDFS or BFS on a gridO(rows x cols)
Course Scheduletopological sort / cycle detectionO(V + E)
Top K Frequent Elementsheap or bucket by countO(n log k)
Climbing Stairs / House Robberdynamic programmingO(n)
Coin Changedynamic programmingO(n x amount)
Merge Intervalssort, then sweepO(n log n)

Twenty problems, eight patterns. Working these until the pattern is instant - not the code - is most of what a 2-3 month runway is for.

63. Name the pattern, do not solve it

Pattern

Four prompts you have not seen in this deck. Thirty seconds each, pattern only.

Predict first

Which pattern does each of the four prompts signal?

Correct: Sliding window (minimum window substring); hash set or XOR (single number); binary search (rotated array minimum); dynamic programming (stock with fee).

promptpatternclue word
shortest substring containingsliding windowsubstring
every element twice except onehash set / XORappears
rotated sorted arraybinary searchsorted
maximize profit with a feedynamic programmingmaximize

Why: Prompt 1 says 'shortest substring containing' - contiguous plus a constraint is a variable window. Prompt 2 is a membership question, so a set works in O(n) space, and XOR does it in O(1) if you spot that a ^ a = 0. Prompt 3 still has a monotonic structure despite the rotation, so the halving argument survives. Prompt 4 has a state you carry forward (holding versus not holding), which is the signature of DP rather than a greedy scan.

64. How the same question gets disguised

Concept

Interviewers know the famous problems are famous. The defence is to change the story, not the structure - which is precisely why pattern recognition beats memorization.

the classicthe disguise you will actually get
Two Sum"Find two transactions that reconcile to this amount."
Number of Islands"Count the distinct clusters in this sensor grid."
Valid Parentheses"Validate that these HTML tags are properly nested."
Climbing Stairs"How many ways can a robot cross n tiles moving 1 or 2 at a time?"
Top K Frequent"Show the k most-viewed products from this event log."

When a prompt has an unusual story, strip the nouns: what is the input shape, what is the output, and what would the brute force be? The pattern usually appears the moment the story is gone.

65. Strip the story yourself

Real world

Prompt: "Our logging service receives a stream of request IDs. Report the first ID that we have already handled, so we can drop the retry."

Discussion prompt

Rewrite this as an abstract problem in one sentence, then name the pattern and the cost.

Hint: What is the input's type? What single question is being asked of it?

Answer:

  • Abstract form: given a sequence of values, return the first value that has already occurred.
  • Pattern: hash set membership.
  • Cost: O(n) time, O(n) space - and this is the first_repeat function from earlier in this deck, word for word.
  • The clarifying question the story hides: is the stream bounded? If it never ends, O(n) space is not acceptable and the real answer involves a bounded structure such as a sliding window of recent IDs.

That last point is where a story-wrapped question earns its keep: the context introduces a constraint the abstract version does not have.

66. Break this plausible claim

Counterexample

A confident-sounding statement. It is false. Find the case that kills it.

Claim: "If the array is sorted, two pointers always beat a hash map, so you should never hash a sorted array."

Discussion prompt

Produce a counterexample, or state the condition under which the claim fails.

Hint: Two pointers need the pair to be found by moving inward. What if the question is not about a pair?

Answer:

  • Counterexample 1 - indices are wanted. Two pointers on a sorted copy lose the original positions, so a question asking for original indices needs the map even though the array is sorted.
  • Counterexample 2 - it is not a pair question. "Does any value appear more than twice?" is a counting question; two pointers have nothing to converge on.
  • Counterexample 3 - many queries. For repeated lookups on the same array, an O(n) hash build then O(1) per query beats O(log n) per query once the query count is large.

Sorted data makes two pointers available, not automatically correct. The test is what the question asks for, not what the input looks like.

67. Will it pass? Commit before you compute.

Hypothesis

An online assessment gives you n up to 200,000 and a 2-second limit. Your candidate solution sorts the array, then for each element does a binary search.

Predict first

Will this pass, and what is its complexity? State your prediction, then check it against the constraint table from Section 2.

  • No - it is O(n^2)
  • Yes - it is O(n log n)
  • Yes - it is O(n)
  • No - it is O(n log n) but that is too slow at 200,000

Correct: Yes - it is O(n log n), which is comfortable at n = 200,000.

partcostat n = 200,000
sortO(n log n)~3.6 million
n binary searchesO(n log n)~3.6 million
totalO(n log n)~7 million - fine

Why: The sort is O(n log n). The loop runs n times and does an O(log n) binary search each time, which is also O(n log n). Added, the whole thing is O(n log n) - roughly 200,000 x 18, about 3.6 million operations, far inside a 2-second budget even in Python. The habit worth building is doing this arithmetic before coding: it takes fifteen seconds and it tells you whether to keep going or find a better approach.

68. Merge Intervals - the sort-then-sweep shape

Worked example

Given a list of intervals, merge all the overlapping ones. It is the canonical member of a whole family: insert interval, meeting rooms, non-overlapping intervals.

The insight comes from sorting, and it is worth stating explicitly

Why: Once the intervals are sorted by start, any interval can only overlap the one immediately before it in the output - so a single sweep suffices, and no pairwise comparison is needed.

def merge(intervals):
    intervals.sort(key=lambda p: p[0])
    out = []
    for start, end in intervals:
        if out and start <= out[-1][1]:
            out[-1][1] = max(out[-1][1], end)
        else:
            out.append([start, end])
    return out

Trace it on [[1,3], [8,10], [2,6], [15,18]]

Why: After sorting by start: [1,3], [2,6], [8,10], [15,18].

intervalout[-1] beforeoverlap?out after
[1,3]-no (empty)[[1,3]]
[2,6][1,3]2 <= 3, yes[[1,6]]
[8,10][1,6]8 > 6, no[[1,6],[8,10]]
[15,18][8,10]15 > 10, no[[1,6],[8,10],[15,18]]

Verify the nesting case separately - it is the one that breaks naive code

Why: merge([[1,4],[2,3]]) returns [[1,4]], not [[1,3]]. That is what max(...) on line 6 is for: the second interval ends earlier, so overwriting the end blindly would shrink the merged interval.

Cost: O(n log n), dominated entirely by the sort. When a solution's cost is all sorting, say so - it tells the interviewer you know the sweep itself is linear.

69. Check yourself: pick the pattern

Check

A prompt you have not seen. Pattern only - do not solve it.

Check your understanding

"Given an array of integers and an integer k, return the maximum sum of any contiguous subarray of exactly k elements. n can be 100,000." Which approach should you propose?

  • A. Sort the array and add the k largest values.
  • B. A fixed-size sliding window, adding the entering element and subtracting the leaving one. (correct)
  • C. A hash map of prefix sums.
  • D. Dynamic programming over every (start, length) pair.

Answer: B

Why: 'Contiguous' plus 'exactly k elements' is the fixed-size window, which is O(n) time and O(1) space. Each step repairs the running sum in constant time instead of re-adding k elements, which is exactly the trap slide from Section 3.

Why A tempts people
Sorting destroys adjacency, and the question is about a contiguous stretch - the k largest values are usually scattered.
Why C tempts people
Prefix sums do solve it in O(n), and mentioning them is not wrong, but the hash map adds nothing here: the window's length is fixed, so no lookup for a previous prefix is needed.
Why D tempts people
Enumerating every (start, length) pair is O(n^2) at best - about 10 billion operations at n = 100,000, and the constraint rules it out before you write it.

70. Teach it to someone a year behind you

Explain it

Discussion prompt

In four sentences, and without using the words 'window', 'pointer' or 'index', explain to a first-year student why the sliding window is faster than checking every substring.

Hint: What work does the fast version refuse to redo?

Answer:

A model answer: "Checking every stretch means re-reading characters you have already read, over and over. Instead, keep one stretch and edit its ends: when you take in a new character on the right, you only have to drop characters from the left until the stretch is legal again. Every character gets taken in once and dropped at most once, so the total work is proportional to the length of the string. The naive version does that much work for each starting position."

If you cannot explain it without the jargon, you are relying on the words rather than the idea - and an interviewer's follow-up will find that out. This is the single most useful drill between sessions.

71. Answer, and rate how sure you are

Commit first

Rate your confidence honestly. Confident-and-wrong is the state worth finding now rather than in a real round.

Predict first

You need the k largest elements of an unsorted array of 1,000,000 numbers, with k = 10. What is the best complexity achievable, and with what structure?

  • O(n log n) with a sort
  • O(n log k) with a min-heap of size k
  • O(n) with a hash map
  • O(k log n) with a max-heap of everything

Correct: O(n log k) with a min-heap of size k - about 1,000,000 x 3 comparisons rather than 1,000,000 x 20.

Why: Keep a min-heap holding the k largest seen so far: push each element, and if the heap exceeds size k, pop the smallest. Each operation is O(log k) with k = 10, so log k is about 3 while log n is about 20 - a real speedup, not a theoretical one. A hash map cannot help because the question is about order, not membership; and building a max-heap of all n elements is O(n) to heapify but then O(k log n) to extract, which is fine too but keeps O(n) memory instead of O(k).

72. Saying It Out Loud, and Testing It Yourself

Section

Section 5

73. The narration script

Concept

Communication is scored separately, and it is the easiest score to raise. These are sentences to have ready - they work on any problem.

Notice how many of these state a trade-off. That is the sentence type that separates a mid-level from a junior read, and it costs nothing to add.

74. Plan in English before you plan in Python

Step zero

Prompt: "Given a list of words, group together the ones that are anagrams of each other."

Discussion prompt

Write your plan in plain English - no code, no Python. Three or four sentences.

Hint: What single fact makes two words anagrams? Could that fact be a dictionary key?

Answer:

  • Two words are anagrams exactly when their sorted letters are identical.
  • So build a dictionary whose key is the sorted word and whose value is the list of words that produced it.
  • One pass over the words, sorting each: O(n k log k) for n words of length k.
  • Return the dictionary's values. Clarify first whether the output order matters, and whether case and spaces are significant.

Every line of that plan becomes one or two lines of code. Candidates who write this plan finish; candidates who start typing usually discover the sorted-key idea at minute 30.

75. Test your own code, before they ask you to

Pattern

Verification is a separate score. This checklist is what to run at minute 35, out loud, without being prompted.

  1. Dry-run the example from step 2. Walk the actual values through the actual lines, saying each variable's value. Do not re-read the code - execute it.
  2. Empty input. Does nums = [] crash on nums[0] or on len(nums) - 1?
  3. One element. Two-pointer and window loops often never execute here.
  4. All identical values. The classic breaker for dedup and window logic.
  1. No valid answer exists. What do you return - -1, None, an empty list? Whatever you clarified at minute 3.
  2. The boundary of the constraint. k = n, k = 0, target smaller than every element.
  3. Off-by-one on every range. Is the last index reachable? Is range(n) or range(n + 1) correct here?

Say what you are testing as you test it: "empty input - line 3 would index into an empty list, so I need a guard." Finding your own bug scores higher than never having one, because it is evidence of a habit.

76. The output is wrong. Diagnose it.

Anomaly

This function is supposed to remove every even number. It does not, and it raises no error - the worst kind of bug.

def remove_evens(nums):
    for n in nums:
        if n % 2 == 0:
            nums.remove(n)
    return nums

print(remove_evens([1, 2, 4, 5]))   # -> [1, 4, 5]

Predict first

Why does 4 survive? Be specific about what the loop is doing when the list changes underneath it.

Correct: Removing 2 shifts every later element left by one, and the loop's internal index has already advanced - so 4 slides into the position the loop just left, and is skipped.

The list shrank under the iterator, so one element was stepped over.

Why: Python's list iterator walks by index. At index 1 the value is 2, which is removed, so the list becomes [1, 4, 5] and 4 now sits at index 1 - but the iterator moves on to index 2, which holds 5. The element is never examined. The fix is to build a new list ([n for n in nums if n % 2]) or to iterate over a copy (for n in nums[:]). Mutating a collection while iterating over it is the most common silent bug in Python interviews, and it shows up in linked-list and grid problems too.

iterator indexlist at that momentvalue seenaction
0[1, 2, 4, 5]1odd, keep
1[1, 2, 4, 5]2remove -> [1, 4, 5]
2[1, 4, 5]5odd, keep - 4 was skipped entirely

77. Same problem, one tool removed

Constraint

You have solved 'find the first repeated value' with a hash set: O(n) time, O(n) space.

Discussion prompt

Now the interviewer says: "Nice. Now do it in O(1) extra space." What do you ask, and what do you propose?

Hint: What extra information about the input would make an in-place approach possible? And what does sorting cost you?

Answer:

  • Ask first: may I modify the input array? If yes, sort it in place and scan for adjacent equals: O(n log n) time, O(1) extra space. State the trade explicitly - you bought space with time.
  • Ask second: are the values bounded, for example all in 1..n? If so, the array can encode 'seen' in itself (negate the value at index |v| - 1), giving O(n) time and O(1) extra space.
  • If neither: say so. "With arbitrary values and an immutable input, O(1) space forces O(n^2) time - which do you prefer?" Naming the impossibility is a correct answer.

This kind of follow-up is not a trap. It is testing whether you know why your solution costs what it costs, and whether you will ask before assuming.

78. Trap: "I think that's it" without a dry run

Trap

The trap

Minute 34. The code is written. You say: "I think that's it - should I run it?"

Hand the verification back to the interviewer

Why: They now have to find your bugs, which converts the verification score into a hint - and hints are recorded.

what happens nexthow it is scored
They spot an empty-input crashneeded a hint on correctness
They ask 'what about duplicates?'did not test independently
Code was actually finestill no evidence of a testing habit

The fix

Minute 34. The code is written. You say: "Let me trace it on our example and then check three edge cases."

Walk the real values through the real lines, out loud

Why: This catches the majority of bugs, and it is observable - which is what the verification score is measuring.

what happens nexthow it is scored
You find your own off-by-onestrong: self-corrected
You note the empty case and add a guardstrong: anticipates edges
Everything passesstrong: verified independently

79. Check yourself: which case breaks it?

Check

A fixed-size window that looks correct. One of these inputs makes it silently wrong - no exception, just a wrong number.

def max_sum_k(nums, k):
    window = sum(nums[:k])
    best = window
    for i in range(k, len(nums)):
        window += nums[i] - nums[i - k]
        best = max(best, window)
    return best
candidate inputwhat is unusual about it
nums = [1, 2, 3], k = 2nothing - the ordinary case
nums = [1, 2, 3], k = 5k is larger than the list
nums = [-5, -2, -9], k = 1every value is negative
nums = [4, 4, 4, 4], k = 4k equals the length

Check your understanding

Which input makes max_sum_k return a wrong answer with no error?

  • A. nums = [1, 2, 3], k = 2
  • B. nums = [1, 2, 3], k = 5 (correct)
  • C. nums = [-5, -2, -9], k = 1
  • D. nums = [4, 4, 4, 4], k = 4

Answer: B

Why: With k larger than the list, nums[:5] silently returns the whole 3-element list rather than raising, so window is the sum of all three, the loop body never runs, and the function confidently returns a sum over a window that does not exist. Python's forgiving slicing is what hides it - a guard such as if k > len(nums): return None (after clarifying what the caller wants) is the fix.

Why A tempts people
A normal case: windows [1,2] and [2,3], returning 5. Nothing special happens here.
Why C tempts people
Negative numbers are handled correctly because best starts at the first window rather than at 0 - which is the bug this code deliberately avoids.
Why D tempts people
k equal to the length is the boundary and it works: the initial window is the whole array and the loop simply does not run.

80. Your Runway: Ten Weeks From Here

Section

Section 6

81. A 10-week plan for a 2-3 month window

Concept

You said interviews start in two to three months and that you want the solution to click on unseen problems. That means pattern reps, spaced out - not a race through a problem list.

weeksfocusvolumethe test that you are ready
1-2Arrays, hashing, two pointers3-4 problems/weekYou name the pattern before reading the constraints
3-4Sliding window, stack, binary search4 problems/weekYou write the boundary search from memory
5-6Trees and graphs (BFS/DFS)4 problems/weekYou write BFS without looking up deque
7-8Dynamic programming, heaps3-4 problems/weekYou derive a recurrence before coding
9-10Mixed review, timed, out loud2 mocks/week45 minutes, narrated, no pausing

Two rules that matter more than the schedule: always talk while you solve, even alone; and redo a problem you got wrong three days later rather than moving on. Volume without recall is the most common way a 10-week plan produces no improvement.

82. How to practise a problem so it transfers

Concept

The difference between 400 problems and 120 problems is much smaller than the difference between these two habits.

instead ofdo thiswhy
Reading the solution when stuckSitting with it 20 minutes, then reading only the hintThe struggle is what builds retrieval
Moving on once it passesWriting the pattern and clue in a one-line logThe log becomes your revision deck
Solving silentlyNarrating as if someone is watchingCommunication is a separate score, and it needs reps
Never repeating a problemRedoing wrong ones after 3 days, then 10Spaced repetition is what makes it automatic

Your one-line log for Two Sum: "unsorted + indices wanted -> hash map of value to index, one pass, O(n)/O(n)." Twenty of those lines is the revision sheet you read the morning of an interview.

83. Draw the map before you close the deck

Connect it up

Do this on paper too. Reproducing it from memory tomorrow is the actual homework.

Draw it

Draw the eight patterns as boxes. For each one, write (a) the clue phrase that triggers it, (b) its typical complexity, and (c) one named problem from the short list. Then draw an arrow between any two patterns that can solve the same problem - for example two pointers and hash map on Two Sum - and label the arrow with what decides between them.

The arrows are the part that matters. Interviewers rarely ask "do you know BFS" - they ask "why did you choose that over the alternative", and the arrows are where those answers live.

84. Where is your weakest link right now?

Exit ticket

Predict first

Of the eight patterns, which one could you not write from scratch in ten minutes today? Name it, and name what specifically is missing - the structure, the loop condition, or the complexity argument.

Correct: Whatever you named is the first thing we work in session two, and the answer is almost always a loop condition rather than a whole structure.

Why: Being able to say 'I can write BFS but I always fumble whether to mark seen on push or on pop' is a far more useful self-assessment than 'I'm bad at graphs'. Specific gaps get fixed in a single session; vague ones turn into months of unfocused practice. Send me your answer before we meet and I will build the next session around it.

85. What you now have

Recap

You have the map. Sessions from here are reps on it - narrated, timed, and with the weak patterns first.

before next sessionwhy
Redo Two Sum and Valid Palindrome, narrated aloudTwo patterns, one session's worth of reps
Write the eight-pattern card from memoryRecall, not recognition
Send me your exit-ticket answerSo session two starts on your weak pattern

Sources

  1. Python Wiki - TimeComplexity (per-operation costs for list, dict, set)
  2. Python 3 documentation - heapq
  3. Python 3 documentation - collections (Counter, deque)
  4. Cormen, Leiserson, Rivest & Stein, Introduction to Algorithms, 4th ed. — MIT Press, 2022 - Ch. 3 (growth of functions), Ch. 6 (heaps), Ch. 14 (dynamic programming), Ch. 20 (elementary graph algorithms)
  5. McDowell, Cracking the Coding Interview, 6th ed. — CareerCup, 2015 - Ch. VI (Big O) and the interview walkthrough
  6. NeetCode - pattern-grouped practice list

Want this taught 1-on-1? Alexander tutors DSA Interview Prep — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108