In 1969 Minsky and Papert showed one neuron can't learn XOR. A 2-4-1 network with one hidden layer can: loss falls from 0.795 to 0.001 and the predictions come out 0, 1, 1, 0.
python3 --version.pip install "numpy>=1.26"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "numpy>=1.26"
cd machine-learning/09-neural-network-from-scratch python3 src/xor.py
import numpy as np
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
y = np.array([[0], [1], [1], [0]])
rng = np.random.default_rng(3)
W1, b1 = rng.normal(size=(2, 4)), np.zeros(4)
W2, b2 = rng.normal(size=(4, 1)), np.zeros(1)
lr = 0.5
for step in range(2000):
h = np.tanh(X @ W1 + b1)
p = 1 / (1 + np.exp(-(h @ W2 + b2)))
loss = np.sum((p - y) ** 2) / 2
if step in (0, 1999):
print(f"step {step:4}: loss {loss:.3f}")
dz2 = (p - y) * p * (1 - p)
dz1 = dz2 @ W2.T * (1 - h ** 2)
W2 -= lr * h.T @ dz2
b2 -= lr * dz2.sum(0)
W1 -= lr * X.T @ dz1
b1 -= lr * dz1.sum(0)
print("predictions:", *p.round().astype(int).flat)One straight line can't solve this four-point puzzle. Two layers can. A neural network, from scratch. The puzzle is XOR: output 1 when exactly one input is 1.
Zero-zero and one-one give 0. The mixed pairs give 1. Try one straight line. A point always lands on the wrong side.
In 1969, Minsky and Papert showed a single-layer perceptron can't learn XOR. Interest in neural networks cooled for years. The fix: one hidden layer, with a bend in it. NumPy only.
Four inputs, and the answers we want. Random weights: two inputs, four hidden neurons, one output. lr is the step size. Then, two thousand training steps.
Forward pass. Each hidden neuron mixes the inputs, and tanh bends the result. The output mixes those, and a sigmoid squeezes it into 0 to 1. The loss: how far off each guess is, squared.
We print it at the first and last step. Now backprop: the chain rule, run backwards. At the output: the error, times the sigmoid's slope. Send it back through W2, times the slope of tanh.
Now every hidden neuron knows its share of the blame. Then every weight takes a small step downhill. That's gradient descent, two thousand times over. Last, round each output to 0 or 1.
Let's run it. Step zero: loss 0.795. Watch the boundary: it stalls, then bends into a band.
Last step: loss 0.001. Predictions: 0, 1, 1, 0. XOR, solved.
Each hidden neuron drew its own straight line. The output combines them into a shape no single line can make. That's the recipe. Layers.
A non-linearity. A loss. Backprop. Gradient descent.
Every deep network, even a large language model, is this same recipe, scaled up. More layers. Billions of weights. The same chain rule.
One catch: change the seed to 0, and this exact code gets stuck. Where you start matters. One line can't. Two layers and backprop can.
Read the lesson on GitHub →