Same 1.2 million taxi trips, but Polars reads 4 of 8 columns and skips 11 of 12 months. Same answer either way: Uptown tips best, at 19.0%.
python3 --version.pip install "polars>=1.0"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "polars>=1.0"
cd data-engineering/14-polars python3 scripts/make_trips.py cd src python3 trips.py
import polars as pl
tip = pl.col("tip") / pl.col("fare")
q = (pl.scan_parquet("trips.parquet")
.filter(pl.col("month") == 12)
.group_by("zone").agg(tip.mean()))
for line in q.explain().splitlines()[-3:-1]:
print(line.strip())
zone, share = q.collect().sort("tip").row(-1)
print(f"Top tippers: {zone}, {share:.1%}")Same question, same file. One library reads all of it. Polars reads only what it needs. The tip share is just an expression.
scan_parquet reads nothing yet. It starts a plan. Keep December, then average the tip by zone. explain() prints the optimized plan.
collect() runs it, and we print the winner. Run it. Projection pushdown: four of eight columns read. Predicate pushdown: 11 of 12 months never get read.
Top tippers: Uptown, at 19 percent. Same answer, a fraction of the work. Read only what you need.
Read the lesson on GitHub →