With 8% positives in 200 rows, a plain split left 1 positive in 50 test rows, just 2%. Adding stratify=y to train_test_split keeps the real 8%.
python3 --version.pip install "scikit-learn>=1.4"python3 -m venv .venv source .venv/bin/activate
pip install "scikit-learn>=1.4"
python3 split.py
from sklearn.model_selection import (
train_test_split as split)
y = [1] * 16 + [0] * 184
_, plain = split(y, random_state=59)
_, strat = split(y, random_state=59, stratify=y)
for s in y, plain, strat: print(sum(s) / len(s))Your test set might miss the rare class. 16 positives, 200 rows. Plain split: one positive in 50 test rows. Add stratify equals y.
Ratio kept. Print each rate: 8, 2, 8 percent. Rare class? Always stratify.
Read the lesson on GitHub →