Compacting 5,500 tiny Parquet files into 4 made the same query run over 20x faster.
python3 --version.pip install "pyarrow>=18"pip install "duckdb>=1.0"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pyarrow>=18" "duckdb>=1.0"
cd data-engineering/22-small-files-problem cd src python3 smallfiles.py
import pyarrow.parquet as pq
from lake import day, query, faster # 4.4M rows
def write(d, n): # same rows, split into n files
size = len(day) // n
for i in range(n):
pq.write_table(day.slice(i * size, size),
f"{d}/{i}.parquet")
write("tiny", 5500); write("big", 4)
faster(query("tiny"), query("big"))Thousands of tiny files just made this query crawl. Same rows, same query. Only the file count changed. 4.
4 million events, and a helper that splits them into n Parquet files. Write them as 5,500 tiny files, then as four big ones. Then time the same query on both. Tiny: about half a second, and 3.
9 MB of it is footers. Big: under 50 milliseconds. Over 20 times faster. Fewer, bigger files.
Same data, faster queries.
Read the lesson on GitHub →