With 3,000 customers and 20% churn, a model that says "nobody leaves" is 80% accurate and catches no one. Pick recall as the goal, compare candidates, and logistic regression wins at 0.76 while staying explainable.
python3 --version.pip install "pandas>=2.2"pip install "scikit-learn>=1.4"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pandas>=2.2" "scikit-learn>=1.4"
cd machine-learning/11-choosing-a-model python3 src/choose.py
from sklearn.model_selection import cross_validate
from customers import load_customers
from models import candidates
df = load_customers()
print("rows, cols:", df.shape,
"missing:", df.isna().sum().sum())
print(f"churn rate: {df.churn.mean():.0%}")
X, y = df.drop(columns="churn"), df.churn
M = ["accuracy", "recall", "precision"]
print("model acc rec prec")
recall = {}
for name, clf in candidates().items():
r = cross_validate(clf, X, y, cv=5, scoring=M)
s = [r["test_" + m].mean() for m in M]
recall[name] = s[1]
print(f"{name:<9}", *[f"{v:.2f}" for v in s])
top = max(recall.values()) - 0.02 # near the best
pick = next(n for n in recall if recall[n] >= top)
print(f"pick: {pick} (recall {recall[pick]:.2f})")There's no best algorithm. Here's how teams actually pick one. Data first. Goal second.
Model last. Say a subscription company wants to predict which customers will churn. The path: explore the data, set the goal, note the constraints, build a baseline, compare a few candidates fairly, and choose. First, know your data.
A helper loads the customer table. Print the shape and the missing values, so we know what needs cleaning. Then the churn rate. When few customers leave, accuracy can lie.
Features go in X, the churn label in y. Now the goal. Missing a churner loses a customer, so we score recall: of everyone who really left, how many did we catch? Precision counts false alarms.
Accuracy stays in as a warning. Then the constraints. The retention team must see why a customer is flagged, and scoring runs nightly on a few thousand rows. So no deep nets.
Candidates, simplest first. A baseline that always says: no churn. Then logistic regression, a decision tree, and a random forest. Each one gets the same five folds of cross-validation, so it's fair.
We average each metric, save the recall, and print a row. Last, the pick. Anything within 0.02 of the best recall makes the cut, and next takes the first one.
The list runs simplest first, so simple wins ties. Let's run it. 3,000 customers, 150 missing values, 20% churn. The baseline scores 80% accuracy, the best in the table, and catches zero churners.
That's the accuracy trap. Logistic regression catches 76%. The tree, 69. The forest, 65, even with the best accuracy of the three.
So we pick logistic regression. Here, the simplest model also scored best on our goal. And it explains itself: each coefficient shows which way a feature pushes the risk. In production, the right model is one you can trust, explain, afford to run, and maintain.
When the data shifts, run the comparison again. No best algorithm. Know your data, set the goal, start simple, compare fairly.
Read the lesson on GitHub →