Lesson 47: Decision Trees — CART, Information Gain, Gini, and Pruning

USAAIO Lesson 47, from Phase 2. It covers information gain from entropy, Gini impurity, the CART binary-split induction algorithm, the stopping criteria, and cost-complexity pruning through ccp_alpha. You build a DecisionTreeClassifier from raw splits up to a pruned final model, on iris and on make_classification. All the trace values were verified with scikit-learn and numpy in June 2026. The lesson runs to 32 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Decision Trees: CART, Info Gain & Pruning

Title

USAAIO · Lesson 47 · Phase 2

From a greedy split criterion to a full tree, then back: how entropy, Gini, and cost-complexity pruning interact — and why the greedy optimum is never globally optimal.

2. By the end of this lesson you can

Objectives

  1. Derive information gain from entropy and apply it to a binary split
  2. Compute Gini impurity and explain when it differs from entropy
  3. Trace the CART induction algorithm — recursive binary splits with stopping criteria
  4. Apply cost-complexity pruning (ccp_alpha) and read the pruning path
  5. Fit, visualize, and prune a DecisionTreeClassifier on real data

3. What survived from Batch Normalization?

Warm-up

Discussion prompt

Before we open Lesson 47: Decision Trees — CART, Information Gain, Gini, and Pruning: without looking back, what was the main idea of Batch Normalization, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

batch normalization forward pass (mu, sigma^2, x_hat, gamma/beta), the full backward pass (dL/dgamma, dL/dbeta, dL/dx), running mean/variance for inference, LayerNorm (feature-dim normalization for transformers), and GroupNorm for small batches. Implement BatchNorm1d from scratch and verify against nn.BatchNorm1d.

4. Impurity: entropy & information gain

Section

Part 1 of 4

5. Entropy measures node impurity

Concept

A node is pure when all samples share one class (H = 0). It is maximally impure when all classes are equally likely (H = log₂ K for K classes).

\[ H(S) = -\sum_{k} p_k \log_2 p_k \]

distributionH (bits)
all one class [1,0,0]0.0000
binary uniform [0.5,0.5]1.0000
binary skewed [0.25,0.75]0.8113
3-class uniform [⅓,⅓,⅓]1.5850

6. Fill in: H (bits) for Entropy measures node impurity

Comparison

Comparison matrix

From Entropy measures node impurity: refill the H (bits) column from what you know. The rest of the table is as it appeared.

distributionH (bits)
all one class [1,0,0]0.0000
binary uniform [0.5,0.5]1.0000
binary skewed [0.25,0.75]0.8113
3-class uniform [⅓,⅓,⅓]1.5850

7. Information gain: IG = H(parent) − weighted child entropy

Concept

\[ \mathrm{IG}(S,\, A) = H(S) - \sum_{v} \frac{|S_v|}{|S|}\, H(S_v) \]

CART picks the feature and threshold that maximises IG at each node. A perfect split sends all positive examples left and all negatives right — yielding maximum IG.

IG is always ≥ 0 (a split never increases total impurity). It equals 0 when the child distributions are identical to the parent.

8. Break it if you can: Information gain: IG = H(parent) − weighted child…

Counterexample

Discussion prompt

CART picks the feature and threshold that maximises IG at each node. A perfect split sends all positive examples left and all negatives right — yielding maximum IG.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

IG is always ≥ 0 (a split never increases total impurity). It equals 0 when the child distributions are identical to the parent.

9. Guess the shape of the answer: Computing IG on a toy split

Estimation

Predict first

Parent: 10 samples, [4 pos, 6 neg]. Split A: left=[3+,1−] (n=4), right=[1+,5−] (n=6).

Commit before you compute: what does Computing IG on a toy split come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: H(parent)=0.9710, H(left)=0.8113, H(right)=0.6500, IG=0.2564

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The split reduces weighted child entropy by 0.2564 bits vs the parent node — this is the information we gain about the label by knowing which side a sample falls on.

10. Computing IG on a toy split

Worked example

Parent: 10 samples, [4 pos, 6 neg]. Split A: left=[3+,1−] (n=4), right=[1+,5−] (n=6).

import numpy as np
from collections import Counter

def entropy(labels):
    n = len(labels)
    counts = Counter(labels)
    return -sum((c/n)*np.log2(c/n) for c in counts.values() if c>0)

parent = [1,1,1,1, 0,0,0,0,0,0]
left   = [1,1,1,0]
right  = [1,0,0,0,0,0]
n = len(parent)

IG = entropy(parent) - (4/10)*entropy(left) - (6/10)*entropy(right)
print(f'{entropy(parent):.4f}  {entropy(left):.4f}  {entropy(right):.4f}  IG={IG:.4f}')

H(parent)=0.9710, H(left)=0.8113, H(right)=0.6500, IG=0.2564

Why: The split reduces weighted child entropy by 0.2564 bits vs the parent node — this is the information we gain about the label by knowing which side a sample falls on.

quantityvalue (bits)
H(parent) [4+, 6−]0.9710
H(left) [3+, 1−]0.8113
H(right) [1+, 5−]0.6500
(4/10)·H(left) + (6/10)·H(right)0.7145
IG = 0.9710 − 0.71450.2564

11. What each one costs: Computing IG on a toy split

Trade off

Comparison matrix

From Computing IG on a toy split: every row here is a choice with a cost. Fill the value (bits) column, then say which row you would actually pick and what you give up for it.

quantityvalue (bits)
H(parent) [4+, 6−]0.9710
H(left) [3+, 1−]0.8113
H(right) [1+, 5−]0.6500
(4/10)·H(left) + (6/10)·H(right)0.7145
IG = 0.9710 − 0.71450.2564

12. Gini impurity

Section

Part 2 of 4

13. Gini: 1 − Σ p²

Concept

\[ G(S) = 1 - \sum_{k} p_k^2 \]

Gini = expected error rate if you labelled a random sample with the class distribution. It peaks at 1 − 1/K for K classes and reaches 0 at a pure node.

distributionGiniEntropy
pure [1,0]0.00000.0000
binary uniform [0.5,0.5]0.50001.0000
3-class [0.5,0.3,0.2]0.62001.4855
binary skewed [0.25,0.75]0.37500.8113

14. By analogy: Gini: 1 − Σ p²

Analogy

Discussion prompt

Explain Gini: 1 − Σ p² by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Gini = expected error rate if you labelled a random sample with the class distribution. It peaks at 1 − 1/K for K classes and reaches 0 at a pure node.

15. Gini vs entropy in practice

Concept

Both criteria yield nearly identical tree structures on most datasets. Entropy is more sensitive to rare classes because log₂ grows faster than the quadratic term near p = 0.

distributionentropy diffGini diff
[0.49,0.49,0.02] vs [0.50,0.50,0.00]0.12140.0194
ratio6.3×—

Gini is slightly cheaper to compute (no log) and is scikit-learn's default. Use entropy when the rare-class penalty matters (e.g. fraud detection, medical diagnosis).

16. Teach it back: Gini vs entropy in practice

Explain it

Discussion prompt

Explain Gini vs entropy in practice to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Both criteria yield nearly identical tree structures on most datasets. Entropy is more sensitive to rare classes because log₂ grows faster than the quadratic term near p = 0.

17. Something is wrong here: using Gini drop vs IG interchangeably

Anomaly

Predict first

A student writes this, and it looks reasonable:

Gini drop and information gain (IG) are both 'impurity reductions,' so they always pick the same best split.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633.

Gini drop and IG are correlated but not equivalent; they can rank splits differently.

Why: On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633. The rankings can diverge on rare-class nodes, leading to structurally different trees.

18. Trap: using Gini drop vs IG interchangeably

Trap

The trap

Gini drop and information gain (IG) are both 'impurity reductions,' so they always pick the same best split.

Pick the split maximising Gini drop, claim it matches IG exactly

Why: On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633. The rankings can diverge on rare-class nodes, leading to structurally different trees.

The fix

Gini drop and IG are correlated but not equivalent; they can rank splits differently.

Toy split: IG=0.2564 bits, Gini drop=0.1633 — same split here, but not guaranteed

Why: Entropy's log₂ is more sensitive near p≈0 (6.3× larger diff on rare class) — so trees trained on imbalanced data can differ between criteria. Know which your library uses.

19. Break it on purpose: using Gini drop vs IG interchangeably

Break the constraint

Discussion prompt

The rule this trap just fixed:

Entropy's log₂ is more sensitive near p≈0 (6.3× larger diff on rare class) — so trees trained on imbalanced data can differ between criteria. Know which your library uses.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633. The rankings can diverge on rare-class nodes, leading to structurally different trees.

20. CART induction algorithm

Section

Part 3 of 4

21. CART: recursive binary splits

Concept

CART (Classification And Regression Trees) grows a tree greedily: at each node, exhaustively search every feature and every threshold for the split that maximises IG (or Gini drop).

  1. For each feature j and each candidate threshold t, compute weighted child impurity
  2. Pick (j, t) = argmin of weighted child impurity
  3. Split the node; recurse on left and right subsets
  4. Stop when a stopping criterion fires

CART always creates binary splits, even for multi-class or continuous features. Categorical features are encoded before splitting.

22. Stopping criteria

Concept

parametereffectdefault
max_depthstop if tree depth reaches thisNone (grow full)
min_samples_splitstop if node has fewer samples than this2
min_impurity_decreasestop if best IG < this threshold0.0
max_leaf_nodeslimit total leaf countNone

No stopping criterion → the tree grows until every leaf is pure (training accuracy = 1.0) or all features are exhausted. Pre-pruning via stopping criteria is fast but tends to underfit; post-pruning is preferred.

23. Guess the shape of the answer: CART root split on iris (verified)

Estimation

Predict first

Iris has 3 classes and 4 features. Manually verify the root split that CART selects: petal length ≤ 2.45.

Commit before you compute: what does CART root split on iris (verified) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Root: petal length ≤ 2.45 → left 40 samples all class 0 (pure, H=0); right 80 samples, IG=0.9183

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Petal length at 2.45 perfectly isolates setosa (class 0).

24. CART root split on iris (verified)

Worked example

Iris has 3 classes and 4 features. Manually verify the root split that CART selects: petal length ≤ 2.45.

from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

iris = load_iris()
X, y = iris.data, iris.target
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)

clf = DecisionTreeClassifier(criterion='entropy', max_depth=3, random_state=42)
clf.fit(X_tr, y_tr)
print(export_text(clf, feature_names=list(iris.feature_names), max_depth=2))
print('test acc:', accuracy_score(y_te, clf.predict(X_te)))

Root: petal length ≤ 2.45 → left 40 samples all class 0 (pure, H=0); right 80 samples, IG=0.9183

Why: Petal length at 2.45 perfectly isolates setosa (class 0). This is the single highest-IG split over all four features and all candidate thresholds.

split nodefeaturethresholdleft nright nIG
rootpetal length≤ 2.4540 (all class 0)80 (class 1+2)0.9183
right childpetal length≤ 4.7541 (class 1)39 (class 2)—
depth-2 rightpetal width≤ 1.75variesvaries—

25. Work backwards from the answer: CART root split on iris (verified)

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Root: petal length ≤ 2.45 → left 40 samples all class 0 (pure, H=0); right 80 samples, IG=0.9183

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Iris has 3 classes and 4 features. Manually verify the root split that CART selects: petal length ≤ 2.45.

26. Why greedy ≠ globally optimal

Concept

CART's greedy criterion maximises IG at each individual node. A split that looks poor locally can enable two much better downstream splits.

Finding the globally optimal tree is NP-hard (exponential search over all possible subtrees). The greedy algorithm is a practical approximation — it works well in practice but has no optimality guarantee.

Proof sketch (USAAIO homework): construct a 2-feature dataset where the globally optimal depth-2 tree uses a 'worse' root split (lower IG) than the greedy choice, but produces zero error while the greedy tree errors on ≥1 sample.

27. Teach it back: Why greedy ≠ globally optimal

Explain it

Discussion prompt

Explain Why greedy ≠ globally optimal to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

CART's greedy criterion maximises IG at each individual node. A split that looks poor locally can enable two much better downstream splits.

28. Cost-complexity pruning

Section

Part 4 of 4

29. Post-growth pruning: the cost-complexity objective

Concept

Grow the full tree first (overfit), then collapse subtrees that aren't earning their complexity cost.

\[ R_\alpha(T) = R(T) + \alpha \cdot |T| \]

R(T) is the misclassification rate of tree T; |T| is the number of leaves; α is the complexity penalty (ccp_alpha in scikit-learn). Larger α → fewer leaves.

α valueeffect
0no penalty — keeps the full grown tree
small (0.01–0.03)collapses a few high-error twigs
large (0.05+)aggressive collapse toward a stump

30. Fill in: effect for Post-growth pruning: the cost-complexity…

Comparison

Comparison matrix

From Post-growth pruning: the cost-complexity objective: refill the effect column from what you know. The rest of the table is as it appeared.

α valueeffect
0no penalty — keeps the full grown tree
small (0.01–0.03)collapses a few high-error twigs
large (0.05+)aggressive collapse toward a stump

31. The weakest-link criterion

Concept

To find the pruning order, compute each internal node's effective α: the penalty at which collapsing its subtree to a leaf becomes neutral.

\[ \alpha_{\text{eff}}(t) = \frac{R(t) - R(T_t)}{|T_t| - 1} \]

R(t) = leaf-error if we collapse to t; R(T_t) = subtree error; |T_t| − 1 = leaves gained by collapsing. Collapse the node with the smallest α_eff first (weakest link).

quantityvalue
R(parent as leaf)5/14 = 0.3571
R(subtree, weighted)2/8 + 6/14·(1/6) = 0.2143
leaves gained2 − 1 = 1
α_eff(0.3571 − 0.2143) / 1 = 0.1429

32. By analogy: The weakest-link criterion

Analogy

Discussion prompt

Explain The weakest-link criterion by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

To find the pruning order, compute each internal node's effective α: the penalty at which collapsing its subtree to a leaf becomes neutral.

33. Guess the shape of the answer: Pruning path on make_classification

Estimation

Predict first

On a 300-sample, 10-feature synthetic dataset, the full tree (depth 8) memorises training data. The pruning path reveals exactly where test accuracy peaks.

Commit before you compute: what does Pruning path on make_classification come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: α=0.050 (depth 3, 5 leaves) gives test acc 0.8444, outperforming the overfitting full tree (0.7556)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The full tree memorises noise; pruning to 5 leaves cuts variance more than it adds bias, recovering 9 percentage points on the held-out test set.

34. Pruning path on make_classification

Worked example

On a 300-sample, 10-feature synthetic dataset, the full tree (depth 8) memorises training data. The pruning path reveals exactly where test accuracy peaks.

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

X, y = make_classification(n_samples=300, n_features=10, n_informative=5,
                           n_redundant=2, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)

full = DecisionTreeClassifier(criterion='entropy', random_state=0)
full.fit(X_tr, y_tr)
path = full.cost_complexity_pruning_path(X_tr, y_tr)
for a in [0.0, 0.028, 0.050, 0.071]:
    c = DecisionTreeClassifier(criterion='entropy', ccp_alpha=a, random_state=0).fit(X_tr, y_tr)
    print(f'a={a:.3f} depth={c.get_depth()} leaves={c.get_n_leaves()} test={accuracy_score(y_te, c.predict(X_te)):.4f}')

α=0.050 (depth 3, 5 leaves) gives test acc 0.8444, outperforming the overfitting full tree (0.7556)

Why: The full tree memorises noise; pruning to 5 leaves cuts variance more than it adds bias, recovering 9 percentage points on the held-out test set.

ccp_alphadepthleavestrain acctest acc
0.0008241.00000.7556
0.0286120.93330.8111
0.050350.83810.8444
0.071120.81430.8444

35. Which is which, by test acc

Discrimination

Sort into buckets

Sort these by test acc, from memory, without looking back at Pruning path on make_classification. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.7556
0.000
0.8111
0.028
0.8444
0.050; 0.071
g1
test acc is "0.7556" for 0.000 — that is what the table on "Pruning path on make_classification" records, and it is the single property separating this group from the rest.
g2
test acc is "0.8111" for 0.028 — that is what the table on "Pruning path on make_classification" records, and it is the single property separating this group from the rest.
g3
test acc is "0.8444" for 0.050, 0.071 — that is what the table on "Pruning path on make_classification" records, and it is the single property separating this group from the rest.

36. Something is wrong here: choosing ccp_alpha on the test set

Anomaly

Predict first

A student writes this, and it looks reasonable:

Scan the pruning path, pick the alpha with the best test accuracy, and report that accuracy as the model's performance.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: You've now used the test set twice (once for selection, once for evaluation).

Select alpha on a validation set (or via cross-validation), then evaluate once on the held-out test set.

Why: You've now used the test set twice (once for selection, once for evaluation). The reported 84.4% is optimistically biased — the model was chosen to maximize it on that split.

37. Trap: choosing ccp_alpha on the test set

Trap

The trap

Scan the pruning path, pick the alpha with the best test accuracy, and report that accuracy as the model's performance.

Use the test set to select alpha=0.050 → report 84.4%

Why: You've now used the test set twice (once for selection, once for evaluation). The reported 84.4% is optimistically biased — the model was chosen to maximize it on that split.

The fix

Select alpha on a validation set (or via cross-validation), then evaluate once on the held-out test set.

Cross-validate alpha → pick best → retrain on full train → evaluate test ONCE

Why: DecisionTreeClassifier.cost_complexity_pruning_path returns the training path; pair it with GridSearchCV(cv=5) to pick alpha without touching the test set at all.

38. Which of these survive contact with Lesson 47: Decision Trees — CART…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A node is pure when all samples share one class (H = 0). It is maximally impure when all classes are equally likely (H = log₂ K for K classes).; CART picks the feature and threshold that maximises IG at each node. A perfect split sends all positive examples left and all negatives right — yielding maximum IG.; Gini = expected error rate if you labelled a random sample with the class distribution. It peaks at 1 − 1/K for K classes and reaches 0 at a pure node.
Breaks
Gini drop and information gain (IG) are both 'impurity reductions,' so they always pick the same best split.; Scan the pruning path, pick the alpha with the best test accuracy, and report that accuracy as the model's performance.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 47: Decision Trees — CART, Information Gain, Gini, and Pruning puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

39. Rebuild the recipe: Decision tree recipe: grow → verify → prune

Ranking

Put in order

These are the steps of Decision tree recipe: grow → verify → prune, scrambled. Put them back in order before the next slide shows you.

  1. Impurity: pick criterion='entropy' or 'gini'; entropy for rare-class sensitivity
  2. Induction: CART exhaustively searches (feature, threshold) pairs at each node, picks max IG
  3. Stopping: set max_depth / min_samples_split to pre-prune; otherwise grow full
  4. Pruning path: clf.cost_complexity_pruning_path(X_tr, y_tr) → scan ccp_alphas
  5. Select alpha: cross-validate on train split; evaluate final model ONCE on test
  6. Inspect: export_text(clf) to verify splits; check depth / leaves align with alpha

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

40. Decision tree recipe: grow → verify → prune

Pattern

  1. Impurity: pick criterion='entropy' or 'gini'; entropy for rare-class sensitivity
  2. Induction: CART exhaustively searches (feature, threshold) pairs at each node, picks max IG
  3. Stopping: set max_depth / min_samples_split to pre-prune; otherwise grow full
  4. Pruning path: clf.cost_complexity_pruning_path(X_tr, y_tr) → scan ccp_alphas
  5. Select alpha: cross-validate on train split; evaluate final model ONCE on test
  6. Inspect: export_text(clf) to verify splits; check depth / leaves align with alpha

41. Where does it stop working: Decision tree recipe: grow → verify → prune

Edge cases

Discussion prompt

Decision tree recipe: grow → verify → prune works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Impurity: pick criterion='entropy' or 'gini'; entropy for rare-class sensitivity
  2. Induction: CART exhaustively searches (feature, threshold) pairs at each node, picks max IG
  3. Stopping: set max_depth / min_samples_split to pre-prune; otherwise grow full
  4. Pruning path: clf.cost_complexity_pruning_path(X_tr, y_tr) → scan ccp_alphas
  5. Select alpha: cross-validate on train split; evaluate final model ONCE on test
  6. Inspect: export_text(clf) to verify splits; check depth / leaves align with alpha

42. Rule out three: Check yourself — information gain

Elimination

Eliminate the wrong options

Parent node: 10 samples [4+, 6−], H = 0.971. Split: left [3+,1−] (n=4), right [1+,5−] (n=6), H(left)=0.811, H(right)=0.650. What is IG to three decimal places?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0.256
  • B. 0.483
  • C. 0.971
  • D. 0.163

Survives elimination: A

Why: IG = 0.971 − (4/10)·0.811 − (6/10)·0.650 = 0.971 − 0.324 − 0.390 = 0.256 (≈0.2564 full precision).

43. Check yourself — information gain

Check

Work this out on paper before clicking.

Check your understanding

Parent node: 10 samples [4+, 6−], H = 0.971. Split: left [3+,1−] (n=4), right [1+,5−] (n=6), H(left)=0.811, H(right)=0.650. What is IG to three decimal places?

  • A. 0.256 (correct)
  • B. 0.483
  • C. 0.971
  • D. 0.163

Answer: A

Why: IG = 0.971 − (4/10)·0.811 − (6/10)·0.650 = 0.971 − 0.324 − 0.390 = 0.256 (≈0.2564 full precision).

Why B tempts people
Computed H(parent) − H(left) alone, forgetting to weight the children and include H(right).
Why C tempts people
Reported H(parent) directly; IG is the reduction in entropy, not the parent entropy itself.
Why D tempts people
Computed the Gini drop (0.1633), not IG. Gini and entropy are correlated but not equal.

44. Answer it before you see the options: Check yourself — Gini impurity

Prediction

Predict first

Compute Gini impurity for a 3-class node: [10, 6, 4] (20 total).

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0.620

Why: G = 1 − (0.5² + 0.3² + 0.2²) = 1 − (0.25 + 0.09 + 0.04) = 1 − 0.38 = 0.62.

45. Check yourself — Gini impurity

Check

Three-class node: 20 samples — 10 class A, 6 class B, 4 class C.

Check your understanding

Compute Gini impurity for a 3-class node: [10, 6, 4] (20 total).

  • A. 0.620 (correct)
  • B. 0.380
  • C. 0.500
  • D. 1.486

Answer: A

Why: G = 1 − (0.5² + 0.3² + 0.2²) = 1 − (0.25 + 0.09 + 0.04) = 1 − 0.38 = 0.62.

Why B tempts people
Computed Σp² = 0.38 but stopped there — Gini is 1 − Σp², not Σp² itself.
Why C tempts people
Binary-uniform Gini (0.5) applied to a 3-class node; the formula still requires 1 − Σp_k².
Why D tempts people
Computed entropy (1.486 bits) instead of Gini — the formulas are H = −Σp log₂ p vs G = 1 − Σp².

46. Rule out three: Check yourself — cost-complexity pruning

Elimination

Eliminate the wrong options

A full decision tree has train acc = 1.00 and test acc = 0.76. After setting ccp_alpha = 0.05 (depth 3, 5 leaves), train = 0.84 and test = 0.84. Why does test accuracy increase even though train accuracy dropped?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Pruning reduces overfitting: removing high-variance leaf splits lowers test error more than it raises bias
  • B. ccp_alpha directly maximises test accuracy during training
  • C. A shallower tree always has higher test accuracy than a deeper one
  • D. The alpha penalty forces the tree to use more informative features

Survives elimination: A

Why: The full tree (depth 8, 24 leaves) memorised training noise — high variance. Pruning to 5 leaves collapses low-information twigs, recovering 8+ percentage points on the test set by trading a small bias increase for a large variance decrease.

47. Check yourself — cost-complexity pruning

Check

Reason about the alpha parameter.

Check your understanding

A full decision tree has train acc = 1.00 and test acc = 0.76. After setting ccp_alpha = 0.05 (depth 3, 5 leaves), train = 0.84 and test = 0.84. Why does test accuracy increase even though train accuracy dropped?

  • A. Pruning reduces overfitting: removing high-variance leaf splits lowers test error more than it raises bias (correct)
  • B. ccp_alpha directly maximises test accuracy during training
  • C. A shallower tree always has higher test accuracy than a deeper one
  • D. The alpha penalty forces the tree to use more informative features

Answer: A

Why: The full tree (depth 8, 24 leaves) memorised training noise — high variance. Pruning to 5 leaves collapses low-information twigs, recovering 8+ percentage points on the test set by trading a small bias increase for a large variance decrease.

Why B tempts people
ccp_alpha penalises tree complexity on the training objective only; it has no access to test labels.
Why C tempts people
Shallow trees can also underfit (a stump at depth 1 gives 0.8144 here — same as depth 3 but worse in general). The relationship is bias–variance, not monotone.
Why D tempts people
ccp_alpha is a leaf-count penalty; it does not reweight or change which features are considered at any split.

48. Your turn: build & prune

Section

Project

49. Project: DecisionTreeClassifier — from scratch to pruned

Concept

Fit a full DecisionTreeClassifier on iris (entropy), visualize it, then prune it with ccp_alpha. Implement a minimal scratch version of the impurity functions.

#requirementtool
1entropy + IG functions from scratchnumpy, Counter
2full + depth-3 tree on irisDecisionTreeClassifier, export_text
3pruning path + cross-validate alphacost_complexity_pruning_path

Build rules: verify your entropy function against H(0.5,0.5)=1.0 and H(1.0)=0.0 before calling it inside any tree logic.

50. Break it if you can: Project: DecisionTreeClassifier — from scratch to…

Counterexample

Discussion prompt

Fit a full DecisionTreeClassifier on iris (entropy), visualize it, then prune it with ccp_alpha. Implement a minimal scratch version of the impurity functions.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: verify your entropy function against H(0.5,0.5)=1.0 and H(1.0)=0.0 before calling it inside any tree logic.

51. Milestone 1 — entropy & IG from scratch

Worked example

Your turn: write entropy(labels) and compute IG on the toy split [4+,6−] → {[3+,1−], [1+,5−]}. Predict IG.

Hint: use Counter(labels) to get class counts; guard against p=0 to avoid log2(0).

import numpy as np
from collections import Counter

def entropy(labels):
    n = len(labels)
    counts = Counter(labels)
    return -sum((c/n)*np.log2(c/n) for c in counts.values() if c > 0)

parent = [1,1,1,1,0,0,0,0,0,0]
left   = [1,1,1,0]
right  = [1,0,0,0,0,0]
IG = entropy(parent) - (4/10)*entropy(left) - (6/10)*entropy(right)
print(f'IG = {IG:.4f}')
quantityvalue
H(parent)0.9710
H(left)0.8113
H(right)0.6500
IG0.2564

52. Watch it run: Milestone 1 — entropy & IG from scratch

Pattern

Step through it

Step through Milestone 1 — entropy & IG from scratch one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is H(parent)
  2. Step 2: quantity is H(left)
  3. Step 3: quantity is H(right)
  4. Step 4: quantity is IG

53. Milestone 2 — grow and inspect the iris tree

Worked example

Your turn: fit a full tree and a depth-3 tree on iris. Predict which split appears at the root and what test accuracy depth-3 achieves.

Hint: export_text(clf, feature_names=list(iris.feature_names)) prints the full decision path.

from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

iris = load_iris()
X, y = iris.data, iris.target
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)

clf = DecisionTreeClassifier(criterion='entropy', max_depth=3, random_state=42)
clf.fit(X_tr, y_tr)
print(export_text(clf, feature_names=list(iris.feature_names), max_depth=2))
print('depth-3 test acc:', accuracy_score(y_te, clf.predict(X_te)))
treedepthleavestest acc
full (no limit)6101.0000
max_depth=3351.0000
root splitpetal length ≤ 2.45——

54. What each one costs: Milestone 2 — grow and inspect the iris tree

Trade off

Comparison matrix

From Milestone 2 — grow and inspect the iris tree: every row here is a choice with a cost. Fill the depth column, then say which row you would actually pick and what you give up for it.

treedepthleavestest acc
full (no limit)6101.0000
max_depth=3351.0000
root splitpetal length ≤ 2.45——

55. Milestone 3 — pruning path and alpha selection

Worked example

Your turn: on make_classification (n=300), grow a full tree, extract the pruning path, and find the alpha where test accuracy peaks.

Hint: clf.cost_complexity_pruning_path(X_tr, y_tr) returns .ccp_alphas; iterate and score each.

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

X, y = make_classification(n_samples=300, n_features=10,
                           n_informative=5, n_redundant=2, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)
full = DecisionTreeClassifier(criterion='entropy', random_state=0).fit(X_tr, y_tr)
path = full.cost_complexity_pruning_path(X_tr, y_tr)
for a in [0.0, 0.028, 0.050, 0.071]:
    c = DecisionTreeClassifier(criterion='entropy', ccp_alpha=a, random_state=0).fit(X_tr, y_tr)
    print(f'a={a:.3f} depth={c.get_depth()} test={accuracy_score(y_te,c.predict(X_te)):.4f}')
ccp_alphadepthleavestest acc
0.0008240.7556
0.0286120.8111
0.050350.8444
0.071120.8444

56. Which is which, by test acc

Discrimination

Sort into buckets

Sort these by test acc, from memory, without looking back at Milestone 3 — pruning path and alpha selection. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.7556
0.000
0.8111
0.028
0.8444
0.050; 0.071
g1
test acc is "0.7556" for 0.000 — that is what the table on "Milestone 3 — pruning path and alpha…" records, and it is the single property separating this group from the rest.
g2
test acc is "0.8111" for 0.028 — that is what the table on "Milestone 3 — pruning path and alpha…" records, and it is the single property separating this group from the rest.
g3
test acc is "0.8444" for 0.050, 0.071 — that is what the table on "Milestone 3 — pruning path and alpha…" records, and it is the single property separating this group from the rest.

57. The full program

Concept

import numpy as np
from collections import Counter
from sklearn.datasets import load_iris, make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.metrics import accuracy_score

# --- Impurity functions ---
def entropy(labels):
    n = len(labels); counts = Counter(labels)
    return -sum((c/n)*np.log2(c/n) for c in counts.values() if c > 0)

parent=[1,1,1,1,0,0,0,0,0,0]; left=[1,1,1,0]; right=[1,0,0,0,0,0]
print('IG:', round(entropy(parent)-(4/10)*entropy(left)-(6/10)*entropy(right),4))

# --- Iris tree ---
iris = load_iris(); X,y = iris.data, iris.target
Xr,Xe,yr,ye = train_test_split(X,y,test_size=0.2,random_state=42)
clf = DecisionTreeClassifier(criterion='entropy',max_depth=3,random_state=42).fit(Xr,yr)
print('iris d3 acc:', accuracy_score(ye,clf.predict(Xe)))

# --- Pruning ---
X2,y2 = make_classification(n_samples=300,n_features=10,n_informative=5,random_state=0)
Xr2,Xe2,yr2,ye2 = train_test_split(X2,y2,test_size=0.3,random_state=0)
full = DecisionTreeClassifier(criterion='entropy',random_state=0).fit(Xr2,yr2)
for a in [0.0,0.050]:
    c = DecisionTreeClassifier(criterion='entropy',ccp_alpha=a,random_state=0).fit(Xr2,yr2)
    print(f'alpha={a}: depth={c.get_depth()} test={accuracy_score(ye2,c.predict(Xe2)):.4f}')
output linevalue
IG:0.2564
iris d3 acc:1.0
alpha=0.0: depth=8 test=0.7556
alpha=0.05: depth=3 test=0.8444

All four numbers match prior milestones. If your IG or test accuracy disagrees, check: (1) random_state=42/0, (2) test_size=0.2/0.3, (3) n_informative=5 in make_classification.

58. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

output linevalue
IG:0.2564
iris d3 acc:1.0
alpha=0.0: depth=8 test=0.7556
alpha=0.05: depth=3 test=0.8444

59. Show it off

Concept

Out loud, slides closed: (1) derive IG from entropy for the toy split [4+,6−] → {[3+,1−],[1+,5−]}, (2) explain why CART's greedy split never guarantees a globally optimal tree, (3) describe what ccp_alpha penalises and how to select it without leaking the test set.

Stretch (USAAIO homework): (a) implement a DecisionTreeClassifier.fit from scratch — just binary Gini splits with max_depth — on iris; (b) prove that entropy is more sensitive to rare classes than Gini by constructing a two-class dataset where they pick different root splits; (c) implement cost-complexity pruning manually (weakest-link pass).

60. Connect it up: Lesson 47: Decision Trees — CART, Information Gain, Gini, and Pruning

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Impurity: entropy & information gain · Gini impurity · CART induction algorithm · Cost-complexity pruning · Your turn: build & prune. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. What you can do now

Recap

ideathe one thing to remember
entropyH = −Σp log₂p; pure node → 0, uniform → log₂K
information gainIG = H(parent) − Σ(|child|/|parent|)·H(child)
GiniG = 1 − Σp²; cheaper, less sensitive to rare classes
CARTexhaustive greedy search; binary; NP-hard globally
ccp_alphaleaf-count penalty; larger α = shallower tree; pick by CV

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 47 (Decision Trees, Information Gain, Gini, CART, Pruning) — Barron · USAAIO Round 2 Preparation, 2026
  2. scikit-learn DecisionTreeClassifier — entropy, gini, ccp_alpha, verified on iris and make_classification — scikit-learn 1.x + numpy 2.2.6, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108