On 50 loan applicants it tries every split and keeps the one with the lowest Gini impurity. Two rules come out, income under 3.6k a month, then under 2.5 years on the job, for 96% accuracy you can read line by line.
python3 --version.pip install "numpy>=1.26"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "numpy>=1.26"
cd machine-learning/08-decision-trees python3 src/tree.py
from loans import X, y, NAMES, LABELS
def gini(y): # how mixed, weighted by size
return 2 * y.mean() * (1 - y.mean()) * len(y)
def best_split(X, y):
return min(
(gini(y[x < t]) + gini(y[x >= t]), f, t)
for f, x in enumerate(X.T)
for t in sorted(set(x))[1:])
def grow(X, y, say="", d=0): # d = depth
vote = round(y.mean())
if d == 2 or gini(y) == 0:
print(say + LABELS[vote])
return (y == vote).sum()
_, f, t = best_split(X, y)
q, m = f"{say}{NAMES[f]} < {t}? ", X[:, f] < t
return (grow(X[m], y[m], q + "yes: ", d + 1)
+ grow(X[~m], y[~m], q + "no: ", d + 1))
print(f"accuracy: {grow(X, y) / len(y):.0%}")This model plays twenty questions with your data, and you can read every answer. A decision tree. Ask a question, split the group, repeat. Say a bank approves small loans.
Fifty past applicants: monthly income, and years at their current job. Each dot is one person: approved, or declined. The tree's job: find the questions that sort them into clean groups. We load the fifty: X holds the two numbers, y holds the decision.
First, a score for how mixed a group is: Gini impurity. Two, times the share approved, times the share declined. All one kind? Zero.
A fifty-fifty mix? As high as it gets. Then times the group's size, so bigger groups count more. Now the search: every feature, every threshold.
Split the group in two, and add up the impurity of both sides. Min keeps the lowest. The best first question: income under 3.6.
Now we grow the tree. Each group votes: the majority decision wins. Two questions deep, or a pure group? Stop there.
Print the answer, and count how many it got right. Otherwise: find the best question, write it down, and mark who says yes. Then grow each side the same way, one level deeper. The last line prints the accuracy.
Let's run it. Income under 3.6? Yes: declined.
All sixteen of them. No? Next question: under 2.5 years at the job?
Declined. Otherwise, approved. Ninety-six percent right. Only two of the fifty don't fit the rules.
Now a new applicant: 5.1 thousand a month, one year at her job. Income under 3.6?
No. Under 2.5 years? Yes.
Declined. And you can tell her exactly why. Lenders often must explain a denial. A tree makes every decision readable.
Trees are also the building blocks of random forests and gradient boosting: the go-to models for tables, the kind of data most companies run on. Twenty questions: ask the one that best splits the group, then ask again.
Read the lesson on GitHub →