Notes from an Iceberg migration

Lessons from moving 12 PB of data to Apache Iceberg without downtime — and what we'd do differently.

PNPriya Natarajan·Mar 15, 2026· 12 min read

Twelve petabytes is not a large migration by hyperscaler standards, but it is large enough that anything you get wrong will hurt for a long time.

What worked: dual-writing to Iceberg alongside the legacy Hive tables for six weeks, with automated diffing on every hourly batch. This let us build confidence without risking user-visible incidents.

What we underestimated: the operational surface of table maintenance. Compaction, snapshot expiration, and orphan file cleanup are all first-class concerns. Budget an engineer's time to own them, or you will discover the hard way that your query performance quietly regresses over months.

What we'd do differently: start the catalog decision earlier. Iceberg is a table format, not a catalog, and the catalog you pick — Nessie, Polaris, AWS Glue, Unity — will shape your governance story for years.

PN
Written by
Priya Natarajan
Head of Data Platform

Priya leads Dataamps' data platform practice, with 15 years shipping lakehouse and streaming systems inside Fortune 500 enterprises.

More from Priya

Get the next essay in your inbox.

One well-considered read from our practitioners, every other Friday.

Ready to amplify your data?

Book a 30-minute strategy call. We'll map your data landscape and identify the three highest-leverage moves.

See case studies