A real Iceberg table keeps every commit as a snapshot. A bad overwrite left 3 rows ($545); time travel showed the original 8 rows ($2,050), and one rollback brought them back.
python3 --version.pip install "pyarrow>=18"pip install "pyiceberg[pyarrow,sql-sqlite]>=0.12"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pyarrow>=18" "pyiceberg[pyarrow,sql-sqlite]>=0.12"
cd data-engineering/16-apache-iceberg python3 src/iceberg.py
from lake import catalog, orders, total
cat = catalog() # SQLite catalog + local folder
t = cat.create_table("shop.orders",
schema=orders.schema)
t.append(orders) # tonight's 8 orders
bad = orders.slice(0, 3) # a buggy job lost rows
t.overwrite(bad) # replaces the whole table
for n, s in enumerate(t.snapshots(), 1):
op = s.summary.operation.value
print(f"snap {n} {op}:",
s.summary["total-records"], "rows")
good = t.snapshots()[0].snapshot_id
print("now:", total(t.scan()))
print("snap 1:", total(t.scan(snapshot_id=good)))
t.manage_snapshots().rollback_to_snapshot(
good).commit()
print("rolled back:", total(t.scan()))A folder of Parquet files can't undo a mistake. This table can. Apache Iceberg is an open table format. It gives files a history.
Think of it like Git for a table. Every write is a commit, called a snapshot. The rows stay in plain Parquet files. The metadata lists which files each snapshot uses.
A catalog points to the current metadata. A commit just swaps that pointer, all at once. In code: a catalog backed by SQLite, a warehouse folder on disk, and an orders table. Append tonight's eight orders.
That's snapshot one. Then a buggy job keeps only three rows, and overwrites the whole table. Now list the snapshots. Each records what the write did, and the rows it left.
Grab snapshot one, and read the table as it was. That's time travel. Then roll back. The current pointer moves back to snapshot one.
Run it. Three snapshots: the append of eight rows, then the overwrite, as a delete and an append of three. Right now, the table holds three rows. $545.
Snapshot one still reads eight rows. $2,050. Roll back, and the table is whole again. Nothing was deleted on disk.
The old Parquet file never moved. Only the pointer did. And because a commit swaps one pointer, readers never see a half-written table. The same trick lets you add or rename a column without rewriting a single file.
Spark, Trino, Flink and Snowflake can all work with the same Iceberg tables. One gotcha: old snapshots keep old files around, so teams expire them on a schedule. Files can't undo. Snapshots can.
That's Apache Iceberg.
Read the lesson on GitHub →