The repo is the textbook.
One folder per video, sorted by domain. Each lesson has the file you saw on screen, a Run it command and the exact output, plus a line-by-line table in its README.
git clone https://github.com/DayanEbrar0X/data-anatomy.ai.gitMachine learning
How models learn from data: loss, gradients, training loops, clustering. Small enough to read in one sitting.
00 · What is machine learning
Nobody told this line where to go.
src/learn.pyFull file →import numpy as np from houses import x, y # 40 examples w, b = 0.0, 0.0 # the model for epoch in range(2000): err = w * x + b - y # how wrong w -= 0.01 * (err * x).mean() # nudge w b -= 0.01 * err.mean() # nudge b print(f"price = {w:.2f} * size + {b:.2f}")Run itpython3 src/learn.py01 · Gradient descent
Every AI learns by walking downhill.
src/descent.pyFull file →import numpy as np loss = lambda w: .5 * (w[0]**2 + 10 * w[1]**2) grad = lambda w: np.array([w[0], 10 * w[1]]) w = np.array([-4.0, 1.6]) # start high lr, start = 0.17, loss(w) # step size for step in range(30): w = w - lr * grad(w) # step downhill print(f"{start:.1f} -> {loss(w):.4f}")Run itpython3 src/descent.py02 · K-means clustering
300 dots. Zero labels. Find the groups.
src/kmeans.pyFull file →import numpy as np from shops import X # 300 unlabeled shops C = X[[234, 263, 286]] # 3 random guesses for step in range(8): d = ((X[:, None] - C)**2).sum(2) g = d.argmin(1) # nearest center C = [X[g == j].mean(0) for j in (0, 1, 2)] print("group sizes:", *np.bincount(g))Run itpython3 src/kmeans.py07 · Linear regression, no libraries
15 lines of Python that learn. No imports.
src/delivery.pyFull file →km = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10] mins = [8, 9, 15, 15, 20, 20, 26, 27, 31, 35] w, b = 0.0, 0.0 lr = 0.02 n = len(km) for epoch in range(1000): dw, db = 0.0, 0.0 for x, y in zip(km, mins): err = w * x + b - y dw += err * x / n db += err / nFirst 12 of 17 linesRun itpython3 src/delivery.py08 · Decision trees
This model plays twenty questions with your data, and you can read every answer.
src/tree.pyFull file →from loans import X, y, NAMES, LABELS def gini(y): # how mixed, weighted by size return 2 * y.mean() * (1 - y.mean()) * len(y) def best_split(X, y): return min( (gini(y[x < t]) + gini(y[x >= t]), f, t) for f, x in enumerate(X.T) for t in sorted(set(x))[1:]) def grow(X, y, say="", d=0): # d = depthFirst 12 of 22 linesRun itpython3 src/tree.py09 · A neural network from scratch
One straight line can't solve this four-point puzzle. Two layers can.
src/xor.pyFull file →import numpy as np X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]]) y = np.array([[0], [1], [1], [0]]) rng = np.random.default_rng(3) W1, b1 = rng.normal(size=(2, 4)), np.zeros(4) W2, b2 = rng.normal(size=(4, 1)), np.zeros(1) lr = 0.5 for step in range(2000): h = np.tanh(X @ W1 + b1) p = 1 / (1 + np.exp(-(h @ W2 + b2))) loss = np.sum((p - y) ** 2) / 2First 12 of 22 linesRun itpython3 src/xor.py10 · Overfitting
A perfect score on the training data is a red flag.
src/overfit.pyFull file →import numpy as np rng = np.random.default_rng(149) temp = np.arange(15, 34) # 19 days, °C trend = 300 / (1 + np.exp((24 - temp) / 3)) cones = trend + rng.normal(0, 20, temp.size) train = temp % 2 == 1 # 10 days val = ~train # 9 unseen days def rmse(p, days): err = np.polyval(p, temp[days]) - cones[days] return np.sqrt(np.mean(err ** 2))First 12 of 22 linesRun itpython3 src/overfit.py11 · Choosing a model
There is no best algorithm. Here's how teams actually pick one.
src/choose.pyFull file →from sklearn.model_selection import cross_validate from customers import load_customers from models import candidates df = load_customers() print("rows, cols:", df.shape, "missing:", df.isna().sum().sum()) print(f"churn rate: {df.churn.mean():.0%}") X, y = df.drop(columns="churn"), df.churn M = ["accuracy", "recall", "precision"] print("model acc rec prec")First 12 of 22 linesRun itpython3 src/choose.py
AI engineering
Building with language models: agents that use tools, retrieval over your own documents, and the systems around the model.
03 · The AI agent loop
Every AI agent runs on one tiny loop: think, act, observe, repeat.
src/agent.pyFull file →from tools import search, calculator from model import think # the LLM call goal = "What is 15% of Q3 revenue?" memory = [goal] for step in range(5): action, arg = think(memory) if action == "answer": print("answer:", arg) break tool = {"search": search,First 12 of 16 linesRun itpython3 src/agent.py06 · RAG from scratch
Your AI never read your handbook.
src/rag.pyFull file →import math, re from docs import DOCS, VOCAB def embed(text): words = re.findall(r"[a-z]+", text.lower()) return [words.count(w) for w in VOCAB] def cosine(a, b): dot = sum(x * y for x, y in zip(a, b)) return dot / (math.hypot(*a) * math.hypot(*b)) ask = "How many vacation days do new hires get?"First 12 of 22 linesRun itpython3 src/rag.py
Methodologies
The ways of working that make AI systems reliable: modeling knowledge with ontologies, and improving systems in measured loops.
04 · What an ontology gives an AI agent
Why your AI agent can't connect the dots.
src/ontology.pyFull file →from facts import FACTS # the ontology: types, and how they connect ONTOLOGY = [("Supplier", "supplies", "Part"), ("Part", "used_in", "Product"), ("Product", "ordered_by", "Customer")] def follow(start, path): found = {start} for rel in path: found = {o for s, r, o in FACTS if s in found and r == rel}First 12 of 17 linesRun itpython3 src/ontology.py05 · Evals and loop engineering
From 60% to 100% without guessing.
src/evals.pyFull file →from cases import CASES def route(text, rules): for team, words in rules.items(): if any(w in text for w in words): return team return "other" def score(v, rules): fails = [t for t, want in CASES if route(t, rules) != want] rate = 1 - len(fails) / len(CASES)First 12 of 21 linesRun itpython3 src/evals.py
Data engineering
Moving and shaping data: pipelines, file formats, and the steps between a source system and an answer.
12 · ETL vs ELT
Same three letters, different order, and it changes your whole data platform.
src/pipeline.pyFull file →import duckdb import pandas as pd from report import report raw = pd.read_csv("orders.csv", dtype=str) # ETL: transform in Python, then load clean = raw.drop_duplicates() del clean["email"] # PII never lands usd = clean["amount"].str.strip("$") clean["amount"] = usd.astype(float) etl = duckdb.connect()First 12 of 22 linesRun itpip install pandas duckdb cd src python3 pipeline.py13 · PySpark: data too big for one machine
Your laptop chokes on a billion rows. Spark splits the job across many machines.
src/sales.pyFull file →from pyspark.sql import SparkSession from pyspark.sql import functions as F spark = (SparkSession.builder.master("local[4]") .config("spark.ui.enabled", "false") .getOrCreate()) spark.sparkContext.setLogLevel("ERROR") sales = (spark.range(100_000_000, numPartitions=4) .withColumn("store", F.col("id") % 50) .withColumn("amt", F.abs(F.hash("id")) % 500)) print("partitions:", sales.rdd.getNumPartitions())First 12 of 22 linesRun itjava -version # should print 17 or 21 pip install pyspark==3.5.4 python3 src/sales.py14 · Why Polars is fast
Same question, same data. One library reads the whole file, the other reads only what it needs.
scripts/make_trips.pyFull file →"""Write src/trips.parquet: 1.2 million seeded taxi trips, one row group per month.""" from pathlib import Path import numpy as np import polars as pl OUT = Path(__file__).resolve().parent.parent / "src" / "trips.parquet" ZONES = ["Airport", "Downtown", "Harbor", "Midtown", "Uptown", "Old Town"] TIP_PCT = np.array([17.0, 15.5, 13.0, 14.5, 19.0, 12.0]) # typical tip % per zone PER_MONTH = 100_000 rng = np.random.default_rng(14)First 12 of 30 linesRun itpython3 scripts/make_trips.py cd src python3 trips.pyBuild Lab 01 · API to Parquet
JSON in. Parquet out.
orders_pipeline/main.pyFull file →import duckdb from pathlib import Path from fetch import fetch_orders from transform import clean PQ = "data/orders.parquet" rows = fetch_orders() df = clean(rows) df.to_parquet(PQ, index=False) kb = lambda p: Path(p).stat().st_size / 1024 j = sum(kb(p) for p in Path("api").glob("*"))First 12 of 17 linesRun itpip install pandas pyarrow duckdb cd orders_pipeline python3 main.pyBuild Lab 02 · Documents to data
Scans in. SQL out. Then search them by meaning.
docs_pipeline/main.pyFull file →import json import duckdb import pandas as pd from ocr import ocr_folder from parse import parse texts = ocr_folder("invoices") records = [parse(t) for t in texts.values()] with open("data/docs.json", "w") as f: json.dump(records, f, indent=2) df = pd.DataFrame(records)First 12 of 22 linesRun itcd docs_pipeline python3 main.py # Part 1 python3 search.py # Part 2
Next on the bench
These folders are in the repo already; their lessons are on the way.