USAAIO Lesson 47, from Phase 2. It covers information gain from entropy, Gini impurity, the CART binary-split induction algorithm, the stopping criteria, and cost-complexity pruning through ccp_alpha. You build a DecisionTreeClassifier from raw splits up to a pruned final model, on iris and on make_classification. All the trace values were verified with scikit-learn and numpy in June 2026. The lesson runs to 32 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 47 · Phase 2
From a greedy split criterion to a full tree, then back: how entropy, Gini, and cost-complexity pruning interact — and why the greedy optimum is never globally optimal.
Objectives
ccp_alpha) and read the pruning pathDecisionTreeClassifier on real dataWarm-up
Discussion prompt
Before we open Lesson 47: Decision Trees — CART, Information Gain, Gini, and Pruning: without looking back, what was the main idea of Batch Normalization, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
batch normalization forward pass (mu, sigma^2, x_hat, gamma/beta), the full backward pass (dL/dgamma, dL/dbeta, dL/dx), running mean/variance for inference, LayerNorm (feature-dim normalization for transformers), and GroupNorm for small batches. Implement BatchNorm1d from scratch and verify against nn.BatchNorm1d.
Section
Part 1 of 4
Concept
A node is pure when all samples share one class (H = 0). It is maximally impure when all classes are equally likely (H = log₂ K for K classes).
\[ H(S) = -\sum_{k} p_k \log_2 p_k \]
| distribution | H (bits) |
|---|---|
| all one class [1,0,0] | 0.0000 |
| binary uniform [0.5,0.5] | 1.0000 |
| binary skewed [0.25,0.75] | 0.8113 |
| 3-class uniform [⅓,⅓,⅓] | 1.5850 |
Comparison
Comparison matrix
From Entropy measures node impurity: refill the H (bits) column from what you know. The rest of the table is as it appeared.
| distribution | H (bits) |
|---|---|
| all one class [1,0,0] | 0.0000 |
| binary uniform [0.5,0.5] | 1.0000 |
| binary skewed [0.25,0.75] | 0.8113 |
| 3-class uniform [⅓,⅓,⅓] | 1.5850 |
Concept
\[ \mathrm{IG}(S,\, A) = H(S) - \sum_{v} \frac{|S_v|}{|S|}\, H(S_v) \]
CART picks the feature and threshold that maximises IG at each node. A perfect split sends all positive examples left and all negatives right — yielding maximum IG.
IG is always ≥ 0 (a split never increases total impurity). It equals 0 when the child distributions are identical to the parent.
Counterexample
Discussion prompt
CART picks the feature and threshold that maximises IG at each node. A perfect split sends all positive examples left and all negatives right — yielding maximum IG.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
IG is always ≥ 0 (a split never increases total impurity). It equals 0 when the child distributions are identical to the parent.
Estimation
Predict first
Parent: 10 samples, [4 pos, 6 neg]. Split A: left=[3+,1−] (n=4), right=[1+,5−] (n=6).
Commit before you compute: what does Computing IG on a toy split come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: H(parent)=0.9710, H(left)=0.8113, H(right)=0.6500, IG=0.2564
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The split reduces weighted child entropy by 0.2564 bits vs the parent node — this is the information we gain about the label by knowing which side a sample falls on.
Worked example
Parent: 10 samples, [4 pos, 6 neg]. Split A: left=[3+,1−] (n=4), right=[1+,5−] (n=6).
import numpy as np
from collections import Counter
def entropy(labels):
n = len(labels)
counts = Counter(labels)
return -sum((c/n)*np.log2(c/n) for c in counts.values() if c>0)
parent = [1,1,1,1, 0,0,0,0,0,0]
left = [1,1,1,0]
right = [1,0,0,0,0,0]
n = len(parent)
IG = entropy(parent) - (4/10)*entropy(left) - (6/10)*entropy(right)
print(f'{entropy(parent):.4f} {entropy(left):.4f} {entropy(right):.4f} IG={IG:.4f}')H(parent)=0.9710, H(left)=0.8113, H(right)=0.6500, IG=0.2564
Why: The split reduces weighted child entropy by 0.2564 bits vs the parent node — this is the information we gain about the label by knowing which side a sample falls on.
| quantity | value (bits) |
|---|---|
| H(parent) [4+, 6−] | 0.9710 |
| H(left) [3+, 1−] | 0.8113 |
| H(right) [1+, 5−] | 0.6500 |
| (4/10)·H(left) + (6/10)·H(right) | 0.7145 |
| IG = 0.9710 − 0.7145 | 0.2564 |
Trade off
Comparison matrix
From Computing IG on a toy split: every row here is a choice with a cost. Fill the value (bits) column, then say which row you would actually pick and what you give up for it.
| quantity | value (bits) |
|---|---|
| H(parent) [4+, 6−] | 0.9710 |
| H(left) [3+, 1−] | 0.8113 |
| H(right) [1+, 5−] | 0.6500 |
| (4/10)·H(left) + (6/10)·H(right) | 0.7145 |
| IG = 0.9710 − 0.7145 | 0.2564 |
Section
Part 2 of 4
Concept
\[ G(S) = 1 - \sum_{k} p_k^2 \]
Gini = expected error rate if you labelled a random sample with the class distribution. It peaks at 1 − 1/K for K classes and reaches 0 at a pure node.
| distribution | Gini | Entropy |
|---|---|---|
| pure [1,0] | 0.0000 | 0.0000 |
| binary uniform [0.5,0.5] | 0.5000 | 1.0000 |
| 3-class [0.5,0.3,0.2] | 0.6200 | 1.4855 |
| binary skewed [0.25,0.75] | 0.3750 | 0.8113 |
Analogy
Discussion prompt
Explain Gini: 1 − Σ p² by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Gini = expected error rate if you labelled a random sample with the class distribution. It peaks at 1 − 1/K for K classes and reaches 0 at a pure node.
Concept
Both criteria yield nearly identical tree structures on most datasets. Entropy is more sensitive to rare classes because log₂ grows faster than the quadratic term near p = 0.
| distribution | entropy diff | Gini diff |
|---|---|---|
| [0.49,0.49,0.02] vs [0.50,0.50,0.00] | 0.1214 | 0.0194 |
| ratio | 6.3× | — |
Gini is slightly cheaper to compute (no log) and is scikit-learn's default. Use entropy when the rare-class penalty matters (e.g. fraud detection, medical diagnosis).
Explain it
Discussion prompt
Explain Gini vs entropy in practice to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Both criteria yield nearly identical tree structures on most datasets. Entropy is more sensitive to rare classes because log₂ grows faster than the quadratic term near p = 0.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Gini drop and information gain (IG) are both 'impurity reductions,' so they always pick the same best split.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633.
Gini drop and IG are correlated but not equivalent; they can rank splits differently.
Why: On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633. The rankings can diverge on rare-class nodes, leading to structurally different trees.
Trap
Gini drop and information gain (IG) are both 'impurity reductions,' so they always pick the same best split.
Pick the split maximising Gini drop, claim it matches IG exactly
Why: On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633. The rankings can diverge on rare-class nodes, leading to structurally different trees.
Gini drop and IG are correlated but not equivalent; they can rank splits differently.
Toy split: IG=0.2564 bits, Gini drop=0.1633 — same split here, but not guaranteed
Why: Entropy's log₂ is more sensitive near p≈0 (6.3× larger diff on rare class) — so trees trained on imbalanced data can differ between criteria. Know which your library uses.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Entropy's log₂ is more sensitive near p≈0 (6.3× larger diff on rare class) — so trees trained on imbalanced data can differ between criteria. Know which your library uses.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
On the toy example: IG selects the split with the bigger entropy reduction (0.2564); Gini drop for the same split = 0.1633. The rankings can diverge on rare-class nodes, leading to structurally different trees.
Section
Part 3 of 4
Concept
CART (Classification And Regression Trees) grows a tree greedily: at each node, exhaustively search every feature and every threshold for the split that maximises IG (or Gini drop).
CART always creates binary splits, even for multi-class or continuous features. Categorical features are encoded before splitting.
Concept
| parameter | effect | default |
|---|---|---|
max_depth | stop if tree depth reaches this | None (grow full) |
min_samples_split | stop if node has fewer samples than this | 2 |
min_impurity_decrease | stop if best IG < this threshold | 0.0 |
max_leaf_nodes | limit total leaf count | None |
No stopping criterion → the tree grows until every leaf is pure (training accuracy = 1.0) or all features are exhausted. Pre-pruning via stopping criteria is fast but tends to underfit; post-pruning is preferred.
Estimation
Predict first
Iris has 3 classes and 4 features. Manually verify the root split that CART selects: petal length ≤ 2.45.
Commit before you compute: what does CART root split on iris (verified) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Root: petal length ≤ 2.45 → left 40 samples all class 0 (pure, H=0); right 80 samples, IG=0.9183
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Petal length at 2.45 perfectly isolates setosa (class 0).
Worked example
Iris has 3 classes and 4 features. Manually verify the root split that CART selects: petal length ≤ 2.45.
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
iris = load_iris()
X, y = iris.data, iris.target
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
clf = DecisionTreeClassifier(criterion='entropy', max_depth=3, random_state=42)
clf.fit(X_tr, y_tr)
print(export_text(clf, feature_names=list(iris.feature_names), max_depth=2))
print('test acc:', accuracy_score(y_te, clf.predict(X_te)))Root: petal length ≤ 2.45 → left 40 samples all class 0 (pure, H=0); right 80 samples, IG=0.9183
Why: Petal length at 2.45 perfectly isolates setosa (class 0). This is the single highest-IG split over all four features and all candidate thresholds.
| split node | feature | threshold | left n | right n | IG |
|---|---|---|---|---|---|
| root | petal length | ≤ 2.45 | 40 (all class 0) | 80 (class 1+2) | 0.9183 |
| right child | petal length | ≤ 4.75 | 41 (class 1) | 39 (class 2) | — |
| depth-2 right | petal width | ≤ 1.75 | varies | varies | — |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Root: petal length ≤ 2.45 → left 40 samples all class 0 (pure, H=0); right 80 samples, IG=0.9183
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Iris has 3 classes and 4 features. Manually verify the root split that CART selects: petal length ≤ 2.45.
Concept
CART's greedy criterion maximises IG at each individual node. A split that looks poor locally can enable two much better downstream splits.
Finding the globally optimal tree is NP-hard (exponential search over all possible subtrees). The greedy algorithm is a practical approximation — it works well in practice but has no optimality guarantee.
Proof sketch (USAAIO homework): construct a 2-feature dataset where the globally optimal depth-2 tree uses a 'worse' root split (lower IG) than the greedy choice, but produces zero error while the greedy tree errors on ≥1 sample.
Explain it
Discussion prompt
Explain Why greedy ≠ globally optimal to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
CART's greedy criterion maximises IG at each individual node. A split that looks poor locally can enable two much better downstream splits.
Section
Part 4 of 4
Concept
Grow the full tree first (overfit), then collapse subtrees that aren't earning their complexity cost.
\[ R_\alpha(T) = R(T) + \alpha \cdot |T| \]
R(T) is the misclassification rate of tree T; |T| is the number of leaves; α is the complexity penalty (ccp_alpha in scikit-learn). Larger α → fewer leaves.
| α value | effect |
|---|---|
| 0 | no penalty — keeps the full grown tree |
| small (0.01–0.03) | collapses a few high-error twigs |
| large (0.05+) | aggressive collapse toward a stump |
Comparison
Comparison matrix
From Post-growth pruning: the cost-complexity objective: refill the effect column from what you know. The rest of the table is as it appeared.
| α value | effect |
|---|---|
| 0 | no penalty — keeps the full grown tree |
| small (0.01–0.03) | collapses a few high-error twigs |
| large (0.05+) | aggressive collapse toward a stump |
Concept
To find the pruning order, compute each internal node's effective α: the penalty at which collapsing its subtree to a leaf becomes neutral.
\[ \alpha_{\text{eff}}(t) = \frac{R(t) - R(T_t)}{|T_t| - 1} \]
R(t) = leaf-error if we collapse to t; R(T_t) = subtree error; |T_t| − 1 = leaves gained by collapsing. Collapse the node with the smallest α_eff first (weakest link).
| quantity | value |
|---|---|
| R(parent as leaf) | 5/14 = 0.3571 |
| R(subtree, weighted) | 2/8 + 6/14·(1/6) = 0.2143 |
| leaves gained | 2 − 1 = 1 |
| α_eff | (0.3571 − 0.2143) / 1 = 0.1429 |
Analogy
Discussion prompt
Explain The weakest-link criterion by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
To find the pruning order, compute each internal node's effective α: the penalty at which collapsing its subtree to a leaf becomes neutral.
Estimation
Predict first
On a 300-sample, 10-feature synthetic dataset, the full tree (depth 8) memorises training data. The pruning path reveals exactly where test accuracy peaks.
Commit before you compute: what does Pruning path on make_classification come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: α=0.050 (depth 3, 5 leaves) gives test acc 0.8444, outperforming the overfitting full tree (0.7556)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The full tree memorises noise; pruning to 5 leaves cuts variance more than it adds bias, recovering 9 percentage points on the held-out test set.
Worked example
On a 300-sample, 10-feature synthetic dataset, the full tree (depth 8) memorises training data. The pruning path reveals exactly where test accuracy peaks.
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=300, n_features=10, n_informative=5,
n_redundant=2, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)
full = DecisionTreeClassifier(criterion='entropy', random_state=0)
full.fit(X_tr, y_tr)
path = full.cost_complexity_pruning_path(X_tr, y_tr)
for a in [0.0, 0.028, 0.050, 0.071]:
c = DecisionTreeClassifier(criterion='entropy', ccp_alpha=a, random_state=0).fit(X_tr, y_tr)
print(f'a={a:.3f} depth={c.get_depth()} leaves={c.get_n_leaves()} test={accuracy_score(y_te, c.predict(X_te)):.4f}')α=0.050 (depth 3, 5 leaves) gives test acc 0.8444, outperforming the overfitting full tree (0.7556)
Why: The full tree memorises noise; pruning to 5 leaves cuts variance more than it adds bias, recovering 9 percentage points on the held-out test set.
| ccp_alpha | depth | leaves | train acc | test acc |
|---|---|---|---|---|
| 0.000 | 8 | 24 | 1.0000 | 0.7556 |
| 0.028 | 6 | 12 | 0.9333 | 0.8111 |
| 0.050 | 3 | 5 | 0.8381 | 0.8444 |
| 0.071 | 1 | 2 | 0.8143 | 0.8444 |
Discrimination
Sort into buckets
Sort these by test acc, from memory, without looking back at Pruning path on make_classification. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Scan the pruning path, pick the alpha with the best test accuracy, and report that accuracy as the model's performance.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: You've now used the test set twice (once for selection, once for evaluation).
Select alpha on a validation set (or via cross-validation), then evaluate once on the held-out test set.
Why: You've now used the test set twice (once for selection, once for evaluation). The reported 84.4% is optimistically biased — the model was chosen to maximize it on that split.
Trap
Scan the pruning path, pick the alpha with the best test accuracy, and report that accuracy as the model's performance.
Use the test set to select alpha=0.050 → report 84.4%
Why: You've now used the test set twice (once for selection, once for evaluation). The reported 84.4% is optimistically biased — the model was chosen to maximize it on that split.
Select alpha on a validation set (or via cross-validation), then evaluate once on the held-out test set.
Cross-validate alpha → pick best → retrain on full train → evaluate test ONCE
Why: DecisionTreeClassifier.cost_complexity_pruning_path returns the training path; pair it with GridSearchCV(cv=5) to pick alpha without touching the test set at all.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Ranking
Put in order
These are the steps of Decision tree recipe: grow → verify → prune, scrambled. Put them back in order before the next slide shows you.
criterion='entropy' or 'gini'; entropy for rare-class sensitivitymax_depth / min_samples_split to pre-prune; otherwise grow fullclf.cost_complexity_pruning_path(X_tr, y_tr) → scan ccp_alphasexport_text(clf) to verify splits; check depth / leaves align with alphaWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
criterion='entropy' or 'gini'; entropy for rare-class sensitivitymax_depth / min_samples_split to pre-prune; otherwise grow fullclf.cost_complexity_pruning_path(X_tr, y_tr) → scan ccp_alphasexport_text(clf) to verify splits; check depth / leaves align with alphaEdge cases
Discussion prompt
Decision tree recipe: grow → verify → prune works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
criterion='entropy' or 'gini'; entropy for rare-class sensitivitymax_depth / min_samples_split to pre-prune; otherwise grow fullclf.cost_complexity_pruning_path(X_tr, y_tr) → scan ccp_alphasexport_text(clf) to verify splits; check depth / leaves align with alphaElimination
Eliminate the wrong options
Parent node: 10 samples [4+, 6−], H = 0.971. Split: left [3+,1−] (n=4), right [1+,5−] (n=6), H(left)=0.811, H(right)=0.650. What is IG to three decimal places?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: IG = 0.971 − (4/10)·0.811 − (6/10)·0.650 = 0.971 − 0.324 − 0.390 = 0.256 (≈0.2564 full precision).
Check
Work this out on paper before clicking.
Check your understanding
Parent node: 10 samples [4+, 6−], H = 0.971. Split: left [3+,1−] (n=4), right [1+,5−] (n=6), H(left)=0.811, H(right)=0.650. What is IG to three decimal places?
Answer: A
Why: IG = 0.971 − (4/10)·0.811 − (6/10)·0.650 = 0.971 − 0.324 − 0.390 = 0.256 (≈0.2564 full precision).
Prediction
Predict first
Compute Gini impurity for a 3-class node: [10, 6, 4] (20 total).
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 0.620
Why: G = 1 − (0.5² + 0.3² + 0.2²) = 1 − (0.25 + 0.09 + 0.04) = 1 − 0.38 = 0.62.
Check
Three-class node: 20 samples — 10 class A, 6 class B, 4 class C.
Check your understanding
Compute Gini impurity for a 3-class node: [10, 6, 4] (20 total).
Answer: A
Why: G = 1 − (0.5² + 0.3² + 0.2²) = 1 − (0.25 + 0.09 + 0.04) = 1 − 0.38 = 0.62.
Elimination
Eliminate the wrong options
A full decision tree has train acc = 1.00 and test acc = 0.76. After setting ccp_alpha = 0.05 (depth 3, 5 leaves), train = 0.84 and test = 0.84. Why does test accuracy increase even though train accuracy dropped?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The full tree (depth 8, 24 leaves) memorised training noise — high variance. Pruning to 5 leaves collapses low-information twigs, recovering 8+ percentage points on the test set by trading a small bias increase for a large variance decrease.
Check
Reason about the alpha parameter.
Check your understanding
A full decision tree has train acc = 1.00 and test acc = 0.76. After setting ccp_alpha = 0.05 (depth 3, 5 leaves), train = 0.84 and test = 0.84. Why does test accuracy increase even though train accuracy dropped?
Answer: A
Why: The full tree (depth 8, 24 leaves) memorised training noise — high variance. Pruning to 5 leaves collapses low-information twigs, recovering 8+ percentage points on the test set by trading a small bias increase for a large variance decrease.
Section
Project
Concept
Fit a full DecisionTreeClassifier on iris (entropy), visualize it, then prune it with ccp_alpha. Implement a minimal scratch version of the impurity functions.
| # | requirement | tool |
|---|---|---|
| 1 | entropy + IG functions from scratch | numpy, Counter |
| 2 | full + depth-3 tree on iris | DecisionTreeClassifier, export_text |
| 3 | pruning path + cross-validate alpha | cost_complexity_pruning_path |
Build rules: verify your entropy function against H(0.5,0.5)=1.0 and H(1.0)=0.0 before calling it inside any tree logic.
Counterexample
Discussion prompt
Fit a full DecisionTreeClassifier on iris (entropy), visualize it, then prune it with ccp_alpha. Implement a minimal scratch version of the impurity functions.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: verify your entropy function against H(0.5,0.5)=1.0 and H(1.0)=0.0 before calling it inside any tree logic.
Worked example
Your turn: write entropy(labels) and compute IG on the toy split [4+,6−] → {[3+,1−], [1+,5−]}. Predict IG.
Hint: use Counter(labels) to get class counts; guard against p=0 to avoid log2(0).
import numpy as np
from collections import Counter
def entropy(labels):
n = len(labels)
counts = Counter(labels)
return -sum((c/n)*np.log2(c/n) for c in counts.values() if c > 0)
parent = [1,1,1,1,0,0,0,0,0,0]
left = [1,1,1,0]
right = [1,0,0,0,0,0]
IG = entropy(parent) - (4/10)*entropy(left) - (6/10)*entropy(right)
print(f'IG = {IG:.4f}')| quantity | value |
|---|---|
| H(parent) | 0.9710 |
| H(left) | 0.8113 |
| H(right) | 0.6500 |
| IG | 0.2564 |
Pattern
Step through it
Step through Milestone 1 — entropy & IG from scratch one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: fit a full tree and a depth-3 tree on iris. Predict which split appears at the root and what test accuracy depth-3 achieves.
Hint: export_text(clf, feature_names=list(iris.feature_names)) prints the full decision path.
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
iris = load_iris()
X, y = iris.data, iris.target
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
clf = DecisionTreeClassifier(criterion='entropy', max_depth=3, random_state=42)
clf.fit(X_tr, y_tr)
print(export_text(clf, feature_names=list(iris.feature_names), max_depth=2))
print('depth-3 test acc:', accuracy_score(y_te, clf.predict(X_te)))| tree | depth | leaves | test acc |
|---|---|---|---|
| full (no limit) | 6 | 10 | 1.0000 |
| max_depth=3 | 3 | 5 | 1.0000 |
| root split | petal length ≤ 2.45 | — | — |
Trade off
Comparison matrix
From Milestone 2 — grow and inspect the iris tree: every row here is a choice with a cost. Fill the depth column, then say which row you would actually pick and what you give up for it.
| tree | depth | leaves | test acc |
|---|---|---|---|
| full (no limit) | 6 | 10 | 1.0000 |
| max_depth=3 | 3 | 5 | 1.0000 |
| root split | petal length ≤ 2.45 | — | — |
Worked example
Your turn: on make_classification (n=300), grow a full tree, extract the pruning path, and find the alpha where test accuracy peaks.
Hint: clf.cost_complexity_pruning_path(X_tr, y_tr) returns .ccp_alphas; iterate and score each.
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=300, n_features=10,
n_informative=5, n_redundant=2, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)
full = DecisionTreeClassifier(criterion='entropy', random_state=0).fit(X_tr, y_tr)
path = full.cost_complexity_pruning_path(X_tr, y_tr)
for a in [0.0, 0.028, 0.050, 0.071]:
c = DecisionTreeClassifier(criterion='entropy', ccp_alpha=a, random_state=0).fit(X_tr, y_tr)
print(f'a={a:.3f} depth={c.get_depth()} test={accuracy_score(y_te,c.predict(X_te)):.4f}')| ccp_alpha | depth | leaves | test acc |
|---|---|---|---|
| 0.000 | 8 | 24 | 0.7556 |
| 0.028 | 6 | 12 | 0.8111 |
| 0.050 | 3 | 5 | 0.8444 |
| 0.071 | 1 | 2 | 0.8444 |
Discrimination
Sort into buckets
Sort these by test acc, from memory, without looking back at Milestone 3 — pruning path and alpha selection. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
import numpy as np
from collections import Counter
from sklearn.datasets import load_iris, make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.metrics import accuracy_score
# --- Impurity functions ---
def entropy(labels):
n = len(labels); counts = Counter(labels)
return -sum((c/n)*np.log2(c/n) for c in counts.values() if c > 0)
parent=[1,1,1,1,0,0,0,0,0,0]; left=[1,1,1,0]; right=[1,0,0,0,0,0]
print('IG:', round(entropy(parent)-(4/10)*entropy(left)-(6/10)*entropy(right),4))
# --- Iris tree ---
iris = load_iris(); X,y = iris.data, iris.target
Xr,Xe,yr,ye = train_test_split(X,y,test_size=0.2,random_state=42)
clf = DecisionTreeClassifier(criterion='entropy',max_depth=3,random_state=42).fit(Xr,yr)
print('iris d3 acc:', accuracy_score(ye,clf.predict(Xe)))
# --- Pruning ---
X2,y2 = make_classification(n_samples=300,n_features=10,n_informative=5,random_state=0)
Xr2,Xe2,yr2,ye2 = train_test_split(X2,y2,test_size=0.3,random_state=0)
full = DecisionTreeClassifier(criterion='entropy',random_state=0).fit(Xr2,yr2)
for a in [0.0,0.050]:
c = DecisionTreeClassifier(criterion='entropy',ccp_alpha=a,random_state=0).fit(Xr2,yr2)
print(f'alpha={a}: depth={c.get_depth()} test={accuracy_score(ye2,c.predict(Xe2)):.4f}')| output line | value |
|---|---|
| IG: | 0.2564 |
| iris d3 acc: | 1.0 |
| alpha=0.0: depth=8 test= | 0.7556 |
| alpha=0.05: depth=3 test= | 0.8444 |
All four numbers match prior milestones. If your IG or test accuracy disagrees, check: (1) random_state=42/0, (2) test_size=0.2/0.3, (3) n_informative=5 in make_classification.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output line | value |
|---|---|
| IG: | 0.2564 |
| iris d3 acc: | 1.0 |
| alpha=0.0: depth=8 test= | 0.7556 |
| alpha=0.05: depth=3 test= | 0.8444 |
Concept
Out loud, slides closed: (1) derive IG from entropy for the toy split [4+,6−] → {[3+,1−],[1+,5−]}, (2) explain why CART's greedy split never guarantees a globally optimal tree, (3) describe what ccp_alpha penalises and how to select it without leaking the test set.
Stretch (USAAIO homework): (a) implement a DecisionTreeClassifier.fit from scratch — just binary Gini splits with max_depth — on iris; (b) prove that entropy is more sensitive to rare classes than Gini by constructing a two-class dataset where they pick different root splits; (c) implement cost-complexity pruning manually (weakest-link pass).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Impurity: entropy & information gain · Gini impurity · CART induction algorithm · Cost-complexity pruning · Your turn: build & prune. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
ccp_alpha) and select alpha via cross-validation, not the test setDecisionTreeClassifier with export_text and cost_complexity_pruning_path| idea | the one thing to remember |
|---|---|
| entropy | H = −Σp log₂p; pure node → 0, uniform → log₂K |
| information gain | IG = H(parent) − Σ(|child|/|parent|)·H(child) |
| Gini | G = 1 − Σp²; cheaper, less sensitive to rare classes |
| CART | exhaustive greedy search; binary; NP-hard globally |
| ccp_alpha | leaf-count penalty; larger α = shallower tree; pick by CV |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.