17 lessons · MIT licensed

The repo is the textbook.

One folder per video, sorted by domain. Each lesson has the file you saw on screen, a Run it command and the exact output, plus a line-by-line table in its README.

Open DayanEbrar0X/data-anatomy.aigit clone https://github.com/DayanEbrar0X/data-anatomy.ai.git

Machine learning

How models learn from data: loss, gradients, training loops, clustering. Small enough to read in one sitting.

Folder on GitHub →
  • 00 · What is machine learning

    Nobody told this line where to go.

    src/learn.pyFull file →
    import numpy as np
    from houses import x, y           # 40 examples
    w, b = 0.0, 0.0                   # the model
    for epoch in range(2000):
        err = w * x + b - y           # how wrong
        w -= 0.01 * (err * x).mean()  # nudge w
        b -= 0.01 * err.mean()        # nudge b
    print(f"price = {w:.2f} * size + {b:.2f}")
    Run itpython3 src/learn.py
  • 01 · Gradient descent

    Every AI learns by walking downhill.

    src/descent.pyFull file →
    import numpy as np
    loss = lambda w: .5 * (w[0]**2 + 10 * w[1]**2)
    grad = lambda w: np.array([w[0], 10 * w[1]])
    w = np.array([-4.0, 1.6])        # start high
    lr, start = 0.17, loss(w)        # step size
    for step in range(30):
        w = w - lr * grad(w)         # step downhill
    print(f"{start:.1f} -> {loss(w):.4f}")
    Run itpython3 src/descent.py
  • 02 · K-means clustering

    300 dots. Zero labels. Find the groups.

    src/kmeans.pyFull file →
    import numpy as np
    from shops import X       # 300 unlabeled shops
    C = X[[234, 263, 286]]        # 3 random guesses
    for step in range(8):
        d = ((X[:, None] - C)**2).sum(2)
        g = d.argmin(1)           # nearest center
        C = [X[g == j].mean(0) for j in (0, 1, 2)]
    print("group sizes:", *np.bincount(g))
    Run itpython3 src/kmeans.py
  • 07 · Linear regression, no libraries

    15 lines of Python that learn. No imports.

    src/delivery.pyFull file →
    km   = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
    mins = [8, 9, 15, 15, 20, 20, 26, 27, 31, 35]
    w, b = 0.0, 0.0
    lr = 0.02
    n = len(km)
    
    for epoch in range(1000):
        dw, db = 0.0, 0.0
        for x, y in zip(km, mins):
            err = w * x + b - y
            dw += err * x / n
            db += err / n
    First 12 of 17 lines
    Run itpython3 src/delivery.py
  • 08 · Decision trees

    This model plays twenty questions with your data, and you can read every answer.

    src/tree.pyFull file →
    from loans import X, y, NAMES, LABELS
    
    def gini(y):  # how mixed, weighted by size
        return 2 * y.mean() * (1 - y.mean()) * len(y)
    
    def best_split(X, y):
        return min(
            (gini(y[x < t]) + gini(y[x >= t]), f, t)
            for f, x in enumerate(X.T)
            for t in sorted(set(x))[1:])
    
    def grow(X, y, say="", d=0):  # d = depth
    First 12 of 22 lines
    Run itpython3 src/tree.py
  • 09 · A neural network from scratch

    One straight line can't solve this four-point puzzle. Two layers can.

    src/xor.pyFull file →
    import numpy as np
    X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
    y = np.array([[0], [1], [1], [0]])
    rng = np.random.default_rng(3)
    W1, b1 = rng.normal(size=(2, 4)), np.zeros(4)
    W2, b2 = rng.normal(size=(4, 1)), np.zeros(1)
    lr = 0.5
    
    for step in range(2000):
        h = np.tanh(X @ W1 + b1)
        p = 1 / (1 + np.exp(-(h @ W2 + b2)))
        loss = np.sum((p - y) ** 2) / 2
    First 12 of 22 lines
    Run itpython3 src/xor.py
  • 10 · Overfitting

    A perfect score on the training data is a red flag.

    src/overfit.pyFull file →
    import numpy as np
    
    rng = np.random.default_rng(149)
    temp = np.arange(15, 34)          # 19 days, °C
    trend = 300 / (1 + np.exp((24 - temp) / 3))
    cones = trend + rng.normal(0, 20, temp.size)
    train = temp % 2 == 1             # 10 days
    val = ~train                      # 9 unseen days
    
    def rmse(p, days):
        err = np.polyval(p, temp[days]) - cones[days]
        return np.sqrt(np.mean(err ** 2))
    First 12 of 22 lines
    Run itpython3 src/overfit.py
  • 11 · Choosing a model

    There is no best algorithm. Here's how teams actually pick one.

    src/choose.pyFull file →
    from sklearn.model_selection import cross_validate
    from customers import load_customers
    from models import candidates
    
    df = load_customers()
    print("rows, cols:", df.shape,
          "missing:", df.isna().sum().sum())
    print(f"churn rate: {df.churn.mean():.0%}")
    
    X, y = df.drop(columns="churn"), df.churn
    M = ["accuracy", "recall", "precision"]
    print("model     acc  rec  prec")
    First 12 of 22 lines
    Run itpython3 src/choose.py

AI engineering

Building with language models: agents that use tools, retrieval over your own documents, and the systems around the model.

Folder on GitHub →
  • 03 · The AI agent loop

    Every AI agent runs on one tiny loop: think, act, observe, repeat.

    src/agent.pyFull file →
    from tools import search, calculator
    from model import think          # the LLM call
    
    goal = "What is 15% of Q3 revenue?"
    memory = [goal]
    
    for step in range(5):
        action, arg = think(memory)
        if action == "answer":
            print("answer:", arg)
            break
        tool = {"search": search,
    First 12 of 16 lines
    Run itpython3 src/agent.py
  • 06 · RAG from scratch

    Your AI never read your handbook.

    src/rag.pyFull file →
    import math, re
    from docs import DOCS, VOCAB
    
    def embed(text):
        words = re.findall(r"[a-z]+", text.lower())
        return [words.count(w) for w in VOCAB]
    
    def cosine(a, b):
        dot = sum(x * y for x, y in zip(a, b))
        return dot / (math.hypot(*a) * math.hypot(*b))
    
    ask = "How many vacation days do new hires get?"
    First 12 of 22 lines
    Run itpython3 src/rag.py

Methodologies

The ways of working that make AI systems reliable: modeling knowledge with ontologies, and improving systems in measured loops.

Folder on GitHub →

Data engineering

Moving and shaping data: pipelines, file formats, and the steps between a source system and an answer.

Folder on GitHub →
  • 12 · ETL vs ELT

    Same three letters, different order, and it changes your whole data platform.

    src/pipeline.pyFull file →
    import duckdb
    import pandas as pd
    from report import report
    
    raw = pd.read_csv("orders.csv", dtype=str)
    
    # ETL: transform in Python, then load
    clean = raw.drop_duplicates()
    del clean["email"]  # PII never lands
    usd = clean["amount"].str.strip("$")
    clean["amount"] = usd.astype(float)
    etl = duckdb.connect()
    First 12 of 22 lines
    Run itpip install pandas duckdb cd src python3 pipeline.py
  • 13 · PySpark: data too big for one machine

    Your laptop chokes on a billion rows. Spark splits the job across many machines.

    src/sales.pyFull file →
    from pyspark.sql import SparkSession
    from pyspark.sql import functions as F
    
    spark = (SparkSession.builder.master("local[4]")
             .config("spark.ui.enabled", "false")
             .getOrCreate())
    spark.sparkContext.setLogLevel("ERROR")
    
    sales = (spark.range(100_000_000, numPartitions=4)
        .withColumn("store", F.col("id") % 50)
        .withColumn("amt", F.abs(F.hash("id")) % 500))
    print("partitions:", sales.rdd.getNumPartitions())
    First 12 of 22 lines
    Run itjava -version # should print 17 or 21 pip install pyspark==3.5.4 python3 src/sales.py
  • 14 · Why Polars is fast

    Same question, same data. One library reads the whole file, the other reads only what it needs.

    scripts/make_trips.pyFull file →
    """Write src/trips.parquet: 1.2 million seeded taxi trips, one row group per month."""
    from pathlib import Path
    
    import numpy as np
    import polars as pl
    
    OUT = Path(__file__).resolve().parent.parent / "src" / "trips.parquet"
    ZONES = ["Airport", "Downtown", "Harbor", "Midtown", "Uptown", "Old Town"]
    TIP_PCT = np.array([17.0, 15.5, 13.0, 14.5, 19.0, 12.0])  # typical tip % per zone
    PER_MONTH = 100_000
    
    rng = np.random.default_rng(14)
    First 12 of 30 lines
    Run itpython3 scripts/make_trips.py cd src python3 trips.py
  • Build Lab 01 · API to Parquet

    JSON in. Parquet out.

    orders_pipeline/main.pyFull file →
    import duckdb
    from pathlib import Path
    from fetch import fetch_orders
    from transform import clean
    
    PQ = "data/orders.parquet"
    rows = fetch_orders()
    df = clean(rows)
    df.to_parquet(PQ, index=False)
    
    kb = lambda p: Path(p).stat().st_size / 1024
    j = sum(kb(p) for p in Path("api").glob("*"))
    First 12 of 17 lines
    Run itpip install pandas pyarrow duckdb cd orders_pipeline python3 main.py
  • Build Lab 02 · Documents to data

    Scans in. SQL out. Then search them by meaning.

    docs_pipeline/main.pyFull file →
    import json
    import duckdb
    import pandas as pd
    from ocr import ocr_folder
    from parse import parse
    
    texts = ocr_folder("invoices")
    records = [parse(t) for t in texts.values()]
    with open("data/docs.json", "w") as f:
        json.dump(records, f, indent=2)
    
    df = pd.DataFrame(records)
    First 12 of 22 lines
    Run itcd docs_pipeline python3 main.py # Part 1 python3 search.py # Part 2

Next on the bench

These folders are in the repo already; their lessons are on the way.