Drop 3 pins into 300 dots, let every dot join its closest pin, move each pin to the middle of its crowd and repeat. You end with 3 clean groups of 100: unsupervised learning at its simplest.
python3 --version.pip install "numpy>=1.26"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "numpy>=1.26"
cd machine-learning/02-k-means python3 src/kmeans.py
import numpy as np
from shops import X # 300 unlabeled shops
C = X[[234, 263, 286]] # 3 random guesses
for step in range(8):
d = ((X[:, None] - C)**2).sum(2)
g = d.argmin(1) # nearest center
C = [X[g == j].mean(0) for j in (0, 1, 2)]
print("group sizes:", *np.bincount(g))Three hundred dots. Zero labels. Can a computer find the groups on its own? Yep.
It's called k-means, and it's weirdly simple. Drop three pins anywhere. Every dot runs to its closest pin. Then each pin slides to the middle of its crowd.
Repeat. Dots switch teams, pins slide again, until nothing moves. Three clean groups. A hundred each.
Nobody labeled a thing. That's k-means.
Read the lesson on GitHub →