40 clean rows, 90% accuracy, a prediction served in under 1 ms: who builds which part.
python3 --version.pip install "pandas>=2.2"pip install "duckdb>=1.0"pip install "scikit-learn>=1.4"pip install "joblib>=1.3"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pandas>=2.2" "duckdb>=1.0" "scikit-learn>=1.4" "joblib>=1.3"
cd data-engineering/23-data-engineer-vs-data-scientist-vs-ml-engineer cd src python3 roles.py
import duckdb
from roles import engineer, scientist, ml_engineer
from report import report
db = duckdb.connect()
first = engineer(db) # data engineer
df = engineer(db) # scheduled rerun
m, acc, base = scientist(df) # data scientist
predict, ms = ml_engineer(m) # ML engineer
report(first, df, acc, base, predict, ms)Same data. Three jobs. Here's who builds what. One churn file.
42 raw rows. The data engineer builds a clean table, safe to rerun. The data scientist fits a model and measures it. The ML engineer saves it, and serves a fast, safe predict.
Run it. 40 clean rows, even on rerun. 90 percent accuracy, versus 50 guessing. Bad input rejected.
A churn score of 0.85, under a millisecond. Same data. One builds the table, one finds the answer, one ships it.
Read the lesson on GitHub →