Both pipelines end with 7 orders and $1,117.15. ETL cleans first; ELT loads raw and keeps that raw table, customer emails included.
python3 --version.pip install "pandas>=2.2"pip install "duckdb>=1.0"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pandas>=2.2" "duckdb>=1.0"
cd data-engineering/12-etl-vs-elt cd src python3 pipeline.py
import duckdb, pandas as pd, report
raw = pd.read_csv("orders.csv", dtype=str)
etl, elt = duckdb.connect(), duckdb.connect()
t = raw.drop_duplicates().drop(columns="email")
t.amount = t.amount.str.strip("$").astype(float)
etl.sql("CREATE TABLE orders AS FROM t")
elt.sql("CREATE TABLE raw_orders AS FROM raw")
elt.sql(open("orders.sql").read())
report.show(etl, elt)Same data, same total. But one pipeline keeps every email. ETL or ELT. Same letters, different order.
Eight raw orders: a duplicate, text amounts, and emails. ETL transforms first: drop the duplicate and the email, fix the amounts. Then loads only clean rows. ELT loads all eight raw rows first, then SQL cleans them, inside the warehouse.
Run it. Both: seven orders, about 1,117 dollars. But ELT also kept the raw table, emails included. Same total.
Where the T runs decides what lands.
Read the lesson on GitHub →