This model got a perfect score on its training data. That's the problem. It's called overfitting: it memorized the data instead of learning the pattern. The only way to see it is to score the model on data it never trained on, and this article does exactly that, in 22 lines of numpy.
What you’ll needPython 3.10+, NumPy
What you’ll need
- python.org/downloads →
Python 3.10 or newerRuns the code. Check yours with
python3 --version. - code.visualstudio.com →
VS Code or CursorEither works; so does any editor you like.
- git-scm.com/downloads →
gitGets the code from GitHub.
NumPy 1.26+Math on many numbers at once
pip install "numpy>=1.26"
- 1Get the code
git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
- 2Make a virtual environmentIt keeps this project's packages separate from the rest of your computer.
python3 -m venv .venv source .venv/bin/activate
- 3Install the packagesrequirements.txt installs every lesson's packages. This one needs only: NumPy.
pip install -r requirements.txt # or just this lesson: pip install "numpy>=1.26"
- 4Run it
cd machine-learning/10-overfitting python3 src/overfit.py
Project structure10-overfitting/
Project structure
- src/
- overfit.pymainthe 22 lines from the video
- README.mdthe idea, how to run it, a line-by-line walkthrough, things to try
The setup: ice cream and temperature
Say you sell ice cream. Nineteen days: the temperature, and cones sold. Hot days sell more, then sales level off. Plus some random noise, because real measurements are always a pattern you want to learn mixed with noise you don't.
The plan has four steps:
- Split: keep some data aside, the validation set. The model never sees it while fitting.
- Fit: train models of growing flexibility on the training set only.
- Score twice: measure the error on the training set and on the validation set.
- Pick: choose the model with the lowest validation error, not the lowest training error.
The code
import numpy as np
rng = np.random.default_rng(149)
temp = np.arange(15, 34) # 19 days, °C
trend = 300 / (1 + np.exp((24 - temp) / 3))
cones = trend + rng.normal(0, 20, temp.size)
train = temp % 2 == 1 # 10 days
val = ~train # 9 unseen days
def rmse(p, days):
err = np.polyval(p, temp[days]) - cones[days]
return np.sqrt(np.mean(err ** 2))
scores = []
print("deg train val")
for deg in range(1, 10, 2):
p = np.polyfit(temp[train], cones[train], deg)
tr, va = rmse(p, train), rmse(p, val)
scores.append((va, deg))
print(f"{deg:3} {tr:6.0f} {va:4.0f}")
print("best degree:", min(scores)[1])In order:
- Lines 3 to 6 make the data. One day per temperature, 15 to 33 °C. Sales follow a smooth S-curve that rises around 24 °C and levels off near 300 cones, plus random noise with a spread of 20 cones. The seed is fixed, so every run prints the same numbers.
- Lines 7 and 8 split it. Odd temperatures are the training set: ten days to learn from. The other nine are held back. The model never sees them.
- Lines 10 to 12 score a curve: take each miss, square it, average, then take the root. That's the root mean squared error, a typical miss in cones.
- Lines 16 to 20 try curves that bend more and more: degree 1, 3, 5, 7 and 9.
polyfitfinds the best curve of that degree using only the training days. Then we score it twice: on the days it learned, and on days it never saw. - Line 22 picks the winner: the degree with the lowest validation error. Putting the validation score first in
each pair makes
minsort by it.
The run
python3 src/overfit.pydeg train val
1 26 27
3 15 15
5 11 19
7 8 24
9 0 69
best degree: 3Degree one is a straight line. It misses by about 27 cones on both sets. Too simple. That's underfitting.
Degree three bends with the data: 15 on training, 15 on validation. Five and seven bend more. Training error keeps falling, but validation creeps up.
Degree nine threads every single training point. Zero error. But between those points, it swings wildly. At 32 degrees it predicts about 110 cones. That day sold 281. Across all nine unseen days, it's off by 69, the worst of the five.
Training error fell from 26 to 0. Validation error went from 27 down to 15, then back up to
69. The best model is the one in the middle.
Why degree 9 scores exactly zero
Ten training points and a degree-9 polynomial. A degree-9 polynomial has 10 coefficients, so it can pass exactly through any 10 points with different x values. The training error is zero by construction, not because the model learned anything. To hit every noisy point, the curve has to swing between them, and the swings are what the validation days catch.
This is why a perfect training score is suspicious in practice. A model with enough parameters can memorize its training set, noise included.
The U-shaped curve
Plot both errors by degree and you get the classic shape. Training error keeps falling as the model gets more flexible. Validation error falls, then climbs. Too simple is underfitting, too flexible is overfitting, and the bottom of the validation curve is the sweet spot: here, degree three.
What to do about it in production
In production, a model only ever sees new days, so the score that counts is the one on data it was not trained on. Three habits follow from that:
- Hold out validation data. Always score on data the model never fit.
- Cross-validate. A single odd/even split is tidy, but with so few days the numbers depend on which days land where. Cross-validation repeats the split several ways and averages the scores.
- Regularize. Regularization keeps the flexible model but adds a penalty for large coefficients (ridge regression is one example), so the curve stays calm. This file only changes the degree.
One more honest caveat. Choosing the degree by validation error means the winner's validation score is slightly optimistic. Real projects keep a third set, the test set, untouched until the very end.
Try this
- Change
range(1, 10, 2)torange(1, 10)to try every degree from 1 to 9. Is the U shape still there, and does the best degree change? - Change the noise in line 6 from
20to5. What happens to the gap between degree 3 and degree 9, and why? - Change the seed in line 3 to
7. Does degree 3 still win? Try a few seeds. - Swap the sets:
train = temp % 2 == 0(9 training days). Why does the validation error for degrees 7 and 9 get so much worse?
A perfect training score means it memorized. Memorizing isn't learning. The full lesson is in the repo under
machine-learning/10-overfitting.