One messy churn dataset, three roles in code: the engineer cleans 42 raw rows to 40 and reruns safely, the scientist reaches 90% accuracy against 50% for guessing, the ML engineer rejects bad input and serves 0.85 in under 1 ms.
python3 --version.pip install "pandas>=2.2"pip install "duckdb>=1.0"pip install "scikit-learn>=1.4"pip install "joblib>=1.3"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pandas>=2.2" "duckdb>=1.0" "scikit-learn>=1.4" "joblib>=1.3"
cd data-engineering/23-data-engineer-vs-data-scientist-vs-ml-engineer cd src python3 roles.py
import time, duckdb, joblib
from sklearn import linear_model as lm
from report import report
X = ["tenure", "calls"]
def engineer(db): # raw file -> typed table
db.sql("""CREATE OR REPLACE TABLE churn AS
SELECT DISTINCT id::INT id,
tenure::INT tenure, calls::INT calls,
churned::BOOL churned
FROM read_csv('churn.csv', all_varchar=1)
WHERE tenure <> ''""")
return db.sql("FROM churn ORDER BY id").df()
def scientist(df): # question -> model + metric
test = df.id % 4 == 0
tr, te = df[~test], df[test]
m = lm.LogisticRegression()
m.fit(tr[X].values, tr.churned)
acc = m.score(te[X].values, te.churned)
p = te.churned.mean() # always-guess baseline
return m, acc, max(p, 1 - p)
def ml_engineer(m): # model -> safe, fast serviceSame data. Three jobs. Here's who builds what. Data engineer.
Data scientist. ML engineer. Think of a restaurant. One person preps the ingredients, one invents the dish, one runs the kitchen.
Our ingredients: 42 raw customer rows. Who churned? One row came out twice. One has no tenure.
The data engineer's job: turn raw files into a table people trust. Create or replace makes it idempotent. Rerun it, same table. DISTINCT drops the duplicate, and every column gets a real type.
Read everything as text, then skip the row with no tenure. The data scientist starts from that table, with a question: who churns? Every fourth customer is held out as a test set. A logistic regression learns from tenure and support calls, then gets scored on unseen customers, against a baseline that always guesses one answer.
The ML engineer turns the model into something an app can call. Save it, and load it back, like a server would. The predict function refuses bad input, instead of guessing, and returns the chance this customer churns. Then a speed check: a thousand calls, timed.
Last, the pipeline runs twice, like a scheduler would, and each role hands off. Let's run it. Engineer: 42 raw rows, 40 clean. The rerun?
Still 40. Scientist: 90 percent accuracy, where guessing gets 50. ML engineer: bad input rejected, and a new customer scores 0.85, in under a millisecond.
In real teams, these lines blur. All three write Python and SQL. Scientists build pipelines too, and ML engineers train models. Titles vary by company, so look at what each one owns: trusted data, a tested answer, and a model running in production.
Same data, three jobs: build the table, find the answer, ship the model.
Read the lesson on GitHub →