Load only the columns you need with usecols and store repeated text as category. The same data drops from 72 MB to 9 MB, measured with memory_usage(deep=True).
python3 --version.pip install "pandas>=2.2"python3 -m venv .venv source .venv/bin/activate
pip install "pandas>=2.2"
python3 slim.py
import pandas as pd
from sales import CSV # 1M orders, 6 columns
mb = lambda d: d.memory_usage(deep=True).sum()/1e6
a = mb(pd.read_csv(CSV))
b = mb(pd.read_csv(CSV, usecols=["city", "qty"],
dtype={"city": "category"}))
print(f"{a:.0f} MB -> {b:.0f} MB ({a/b:.0f}x)")Your pandas DataFrame is eating memory it doesn't need. Measure the full load, deep equals True. Now read only the columns you need, and store repeated text as category. 72 MB down to 9.
Eight times smaller. Load less. Store smarter.
Read the lesson on GitHub →