5,500 tiny Parquet files carried 3.9 MB of footers and took about 0.5 s to query. The same rows compacted into 4 files ran over 20x faster; exact speedups depend on the machine.
python3 --version.pip install "pyarrow>=18"pip install "duckdb>=1.0"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pyarrow>=18" "duckdb>=1.0"
cd data-engineering/22-small-files-problem cd src python3 smallfiles.py
import pyarrow.parquet as pq
from lake import clicks, reset, query, faster
day = clicks(4_400_000) # one day of events
reset("tiny", "big")
# streaming: one small file per micro-batch
for i in range(5500):
part = day.slice(i * 800, 800)
pq.write_table(part, f"tiny/{i:04}.parquet")
slow = query("tiny") # median of 7 runs
# compaction: read them all, write 4 big files
t = pq.read_table("tiny").combine_chunks()
for i in range(4):
part = t.slice(i * 1_100_000, 1_100_000)
pq.write_table(part, f"big/{i}.parquet")
fast = query("big")
faster(slow, fast)Thousands of tiny files just made this query crawl. Same rows, same query. Only the number of files changed. Every file costs something before any data is read: open it, read its footer, plan the scan.
It's like mailing a thousand envelopes instead of one box. The postage adds up. pyarrow writes the Parquet files. A small helper times DuckDB queries.
Our data: 4.4 million click events, and two empty folders. First, what streaming jobs often do: a small file per micro-batch. That's 5,500 slices of 800 rows, each written as its own Parquet file.
Then we time a group-by on that folder: the median of seven runs. Now compaction. Read all the tiny files back as one table, then write it out as four files of 1.1 million rows each.
Same query, on the big folder. Last, compare the two medians. Let's run it. Tiny: 5,500 files, 45.
8 MB. 3.9 MB of that is just footers. The query: about half a second.
Big: four files, 19.7 MB, 4 KB of footers, under 50 milliseconds. Same data. Over 20 times faster.
Tiny files compress worse too: big is less than half the size. Exact speedups depend on the machine. On S3, every file is also a network request, so the gap grows. Streaming jobs and frequent small writes create tiny files all day.
So table formats compact them: OPTIMIZE in Delta Lake, rewrite_data_files in Iceberg. Aim for files in the hundreds of megabytes, not kilobytes. Tiny files made it crawl. Compaction made it fast again.
Read the lesson on GitHub →