If 20% of customers leave, always guessing "stays" scores 80%. Set the goal first, recall, and logistic regression catches 76% of the churners.
python3 --version.pip install "pandas>=2.2"pip install "scikit-learn>=1.4"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pandas>=2.2" "scikit-learn>=1.4"
cd machine-learning/11-choosing-a-model python3 src/choose.py
from sklearn.model_selection import cross_validate
from churn import load_churn, candidates
X, y = load_churn()
M, rec = ["accuracy", "recall"], {}
for n, clf in candidates().items():
r = cross_validate(clf, X, y, cv=5, scoring=M)
acc, rec[n] = (r["test_"+m].mean() for m in M)
print(f"{n:9} acc {acc:.2f} rec {rec[n]:.2f}")
print("pick:", max(rec, key=rec.get))This model is 80% accurate, and catches no one. Say 20% of customers churn. Guess nobody leaves, and you're 80% right. The fix: load 3,000 customers, pick the goal first: recall, the share of churners we catch.
Every candidate gets the same five folds. Average the scores, and pick the best recall. The baseline: 80% accuracy, zero recall. Logistic regression: 73% accurate, but it catches 76%.
So logistic wins. Pick the goal first, then the model. Accuracy alone can lie.
Read the lesson on GitHub →