It tries every split and keeps the purest. On 50 loan applicants it learns two questions, income under 3.6k a month and under 2.5 years on the job, and gets 96% right.
python3 --version.pip install "numpy>=1.26"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "numpy>=1.26"
cd machine-learning/08-decision-trees python3 src/tree.py
from loans import X, y, NAMES, gini
ok = y == y # approve all 50, for now
for q in range(2): # ask two questions
_, f, t = min((gini(y[ok & (x < t)]) +
gini(y[ok & (x >= t)]), f, t)
for f, x in enumerate(X.T) for t in x)
print(f"{NAMES[f]} < {t}? declined")
ok &= X[:, f] >= t
print(f"accuracy: {(ok == y).mean():.0%}")This model plays twenty questions, and you can read every answer. Fifty loan applicants: income, and years on the job. Each one approved, or declined. Two questions.
Each tries every feature, every cut, Gini scores how mixed both sides are. Min keeps the lowest: income under 3.6. Below: declined.
The rest? It asks again. Last line: the accuracy. Run it.
Income under 3.6? Declined. Under 2.
5 years? Declined. Everyone else, approved. Ninety-six percent right.
Try every split. Keep the purest. Read every answer.
Read the lesson on GitHub →