On ice cream sales, a degree-9 curve scores 0 error on training and misses new days by 69. A simple degree-3 curve scores 15 on both, because it learned the pattern, not the noise.
python3 --version.pip install "numpy>=1.26"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "numpy>=1.26"
cd machine-learning/10-overfitting python3 src/overfit.py
import numpy as np
rng = np.random.default_rng(149)
temp = np.arange(15, 34) # 19 days, °C
trend = 300 / (1 + np.exp((24 - temp) / 3))
cones = trend + rng.normal(0, 20, temp.size)
train = temp % 2 == 1 # 10 days
val = ~train # 9 unseen days
def rmse(p, days):
err = np.polyval(p, temp[days]) - cones[days]
return np.sqrt(np.mean(err ** 2))
scores = []
print("deg train val")
for deg in range(1, 10, 2):
p = np.polyfit(temp[train], cones[train], deg)
tr, va = rmse(p, train), rmse(p, val)
scores.append((va, deg))
print(f"{deg:3} {tr:6.0f} {va:4.0f}")
print("best degree:", min(scores)[1])This model got a perfect score on its training data. That's the problem. It's called overfitting. It memorized the data instead of learning the pattern.
Say you sell ice cream. Nineteen days: the temperature, and cones sold. Hot days sell more, then sales level off. Plus some random noise.
In code: one day per temperature, 15 to 33 degrees. Sales follow a smooth curve, plus random noise with a fixed seed. Odd temperatures are the training set: ten days to learn from. The other nine are held back.
It never sees them: the validation set. To score a curve, we take each miss, square it, average, then take the root. Now we try curves that bend more and more: degree 1, 3, 5, 7 and 9. polyfit finds the best curve of that degree, using only the training days.
Then we score it twice: on the days it learned, and on days it never saw. We keep each validation score, and print a little table. The winner is the degree with the lowest validation error. Let's run it.
Degree one is a straight line. It misses by about 27 cones, on both sets. Too simple. That's underfitting.
Degree three bends with the data: 15 on training, 15 on validation. Five and seven bend more. Training error keeps falling, but validation creeps up. Degree nine threads every single training point.
Zero error. But between those points, it swings wildly. At 32 degrees it predicts 110 cones. That day sold 281.
Across all nine unseen days, it's off by 69. The worst of the five. Plot both errors by degree, and you get the classic shape. Training error keeps falling.
Validation falls, then climbs. The bottom of that U is the sweet spot: degree three. In production, a model only ever sees new days. That's why we hold out validation data, cross-validate, and regularize.
Regularizing penalizes wild curves, so the model stays calm. A perfect training score means it memorized. Memorizing isn't learning.
Read the full article →