{"id":351,"date":"2024-11-26T09:25:22","date_gmt":"2024-11-26T09:25:22","guid":{"rendered":"https:\/\/joskodeboer.nl\/?p=351"},"modified":"2026-04-10T14:45:00","modified_gmt":"2026-04-10T14:45:00","slug":"apache-iceberg-the-next-generation-table-format-for-big-data","status":"publish","type":"post","link":"https:\/\/joskodeboer.nl\/index.php\/2024\/11\/26\/apache-iceberg-the-next-generation-table-format-for-big-data\/","title":{"rendered":"Is Apache Iceberg the future of data storage?"},"content":{"rendered":"\n<div class=\"et_pb_section_0 et_pb_section et_section_regular et_block_section preset--module--divi-section--default\"><div class=\"et_pb_row_0 et_pb_row et_block_row\"><div class=\"et_pb_column_0 et_pb_column et_pb_column_1_2 et_block_column et_pb_css_mix_blend_mode_passthrough\"><div class=\"et_pb_text_0 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p>Data platforms have been relying on columnar storage for a while, but is it sufficient to meet modern standards? The demands of scalability, flexibility, and governance call for more than just columnar storage (e.g. Parquet). Apache Iceberg has growing momentum and support. I was intrigued by the promise of redefining data management and decided to explore its architecture, use cases and adoption journey.<\/p>\n<p>Today, teams are managing datasets that rapidly expand and evolve. Meanwhile, they have to deal with complex queries, schema updates, compatibility, data sharing and cross-platform compatibility.\u00a0 Traditional formats are no longer enough \u2014 enter Apache Iceberg.<\/p>\n<\/div><\/div><\/div><div class=\"et_pb_column_1 et_pb_column et_pb_column_1_2 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\"><div class=\"et_pb_image_0 et_pb_image et_pb_module et_block_module preset--module--divi-image--default\"><span class=\"et_pb_image_wrap\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/joskodeboer.nl\/wp-content\/uploads\/2024\/10\/iceberg_blog.png\" width=\"1792\" height=\"1024\" srcset=\"https:\/\/joskodeboer.nl\/wp-content\/uploads\/2024\/10\/iceberg_blog.png 1792w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2024\/10\/iceberg_blog-1280x731.png 1280w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2024\/10\/iceberg_blog-980x560.png 980w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2024\/10\/iceberg_blog-480x274.png 480w\" sizes=\"(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) and (max-width: 1280px) 1280px, (min-width: 1281px) 1792px, 100vw\" class=\"wp-image-354\" title=\"iceberg_blog\" \/><\/span><\/div><\/div><\/div><div class=\"et_pb_row_1 et_pb_row et_block_row\"><div class=\"et_pb_column_2 et_pb_column et_pb_column_4_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\"><div class=\"et_pb_text_1 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p>Iceberg addresses the evolving demands of modern data environments, offering improved query performance, seamless schema evolution, and data lineage insights.<\/p>\n<p>Let's examine the core features of Apache Iceberg, explore its architecture, and talk a little about what I could find on migrations. This greater detail will allow us to decide whether we need to start using Apache Iceberg.<\/p>\n<\/div><\/div><\/div><\/div><\/div><div class=\"et_pb_section_1 et_pb_section et_section_regular et_block_section preset--module--divi-section--default\"><div class=\"et_pb_row_2 et_pb_row et_block_row\"><div class=\"et_pb_column_3 et_pb_column et_pb_column_4_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\"><div class=\"et_pb_text_2 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h2>Understanding the Apache Iceberg Architecture<\/h2>\n<\/div><\/div><div class=\"et_pb_text_3 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p>Apache Iceberg\u2019s architecture is built for the demands of large-scale data, offering a multi-layered design that optimizes metadata management, query performance, and scalability across platforms. Let\u2019s examine the catalog, metadata, and data layers, each purpose-built to tackle specific challenges in big data environments.<\/p>\n<h3><strong>Catalog Layer: Metadata Management<\/strong><\/h3>\n<p>At the core of the Iceberg, the catalog layer organizes and manages metadata, ensuring scalability and consistency across your environments. The layer is a single source of truth for all Iceberg tables, enabling (multiple) compute engines to access a unified view of the same table.<\/p>\n<p><strong>Key functions of Iceberg's catalog:<\/strong><\/p>\n<ul>\n<li><strong>Table Registration<\/strong>: Basic operations such as creating, listing, and managing tables across a distributed environment should be simple in a modern system.\u00a0 The table registration is a global namespace for all Iceberg tables, allowing seamless access regardless of the processing engine or underlying storage.<\/li>\n<li><strong>Snapshot Management:\u00a0<\/strong>Data versioning is handled with a snapshot-based approach. Each table contains a series of immutable snapshots, representing the state of the table at a specific datetime. The catalog enables capabilities by tracking these snapshots. For example, you can use time travel in your queries to access a particular snapshot or focus on differences by comparing the changes.<\/li>\n<li><strong>Variety of catalog backends:\u00a0<\/strong> Iceberg supports a variety of catalog backends, such as Hive Metastore (ideal for transitions from legacy), AWS Glue, or custom solutions tailored to your needs. The flexibility ensures that Iceberg integrates into your existing data ecosystem without drastic architectural changes.<\/li>\n<li><strong>Concurrent Transactions<\/strong>: The catalog layer supports concurrency control, allowing multiple users access for operations at the same time. By isolating the changes to the snapshot level, conflict is prevented and readers access a consistent version of the data.<\/li>\n<\/ul>\n<p><strong>Benefits<\/strong><\/p>\n<p>The catalog scales horizontally, making it suitable for environments with thousands of tables or billions of rows. By decoupling metadata from compute, the catalog ensures a consistent view of the table across engines such as Spark, Flink, Trino, etc. Administrations can manage all Iceberg tables from a single location making governance simpler and reducing operational overhead.<\/p>\n<p>Even though it is highly robust, users may encounter challenges. As the number of snapshots or tables grows, the amount of metadata can increase rapidly. Regular maintenance is essential, like cleaning up used snapshots.<\/p>\n<p>For teams transitioning from other table formats, integrating the catalog with existing workflows may be a bit of a learning curve. Using supported tools like AWS glue or Hive metastore can ease this transition.<\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div><div class=\"et_pb_text_4 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h3>Metadata Layer:<\/h3>\n<p>This layer is crucial for enabling data lineage and optimizing query performance, both of which are essential for managing large-scale datasets. The layer achieves this by organizing metadata in hierarchies, with key components like manifests and manifest lists providing detailed insights into the structure and evolution of data.<\/p>\n<p>The metadata files store critical information in lightweight JSON, such as schema definitions, partitioning details, and file locations. By pruning irrelevant data, Iceberg ensures systems read only the required segments, avoiding costly full-table scans. Manifest files contain row-level statistics, such as the min\/max values for columns, enabling precise filtering during execution. This saves time, reduces computational overhead, and lowers cost, especially for large datasets. Manifest lists connect high-level metadata and granular manifest files, summarizing partition and file details for query planning. This hierarchy allows Iceberg to scale efficiently even with thousands of partitions.<\/p>\n<h4>Data Lineage<\/h4>\n<p>The metadata layer supports data lineage, tracking how data evolves within a table over time. This is achieved through Icebergs snapshot-based architecture, where each snapshot represents the table's state at a specific point.<\/p>\n<p>By retaining versioned metadata, Iceberg enables tracking of changes to a table, such as additions, deletions, or schema modifications, providing detailed insights into the evolution of data. This presents a clear historical view of the table, which is useful for governance, auditing and compliance efforts.<\/p>\n<p>While Iceberg excels at tracking a table's evolution, tracing data lineage across multiple tables does require additional tools. For example, you can look at Apache Atlas to enable end-to-end lineage tracing, connecting your ETL processes and linking data transformations across tables.<\/p>\n<\/div><\/div><div class=\"et_pb_text_5 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p style=\"padding-left: 40px;\"><strong>ChatGPT: <\/strong>Integrating Apache Iceberg with Apache Atlas enables robust data lineage and governance across big data ecosystems. Apache Atlas provides metadata management and lineage capabilities, while Iceberg's metadata layer offers detailed snapshots and file-level information. By combining the two, organizations can create end-to-end lineage graphs that trace data origins, transformations, and destinations across tables and systems. This integration can be achieved by configuring Atlas to monitor ETL jobs and processes interacting with Iceberg tables, enabling a unified view of data flows for compliance, auditing, and impact analysis.<\/p>\n<\/div><\/div><div class=\"et_pb_text_6 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h3>Data Layer<\/h3>\n<p>The data layer handles files multiple engines can query while keeping data immutable for version control. This is why Iceberg excels at interoperability and scalability.<\/p>\n<p>Data is organized in immutable columnar files, often in Parquet or ORC format. When data is updated, new files are added while old ones are marked as obsolete but remain accessible for past snapshots. This setup provides a natural version control system.<\/p>\n<p>Hidden partitioning eliminates the need to manage physical partitions, letting users query by conditions like data or region without manual partitioning. Teams avoid reorganization costs, making it ideal for large dynamic datasets where partitions need frequent adjustment.\u00a0<\/p>\n<p>\u00a0The open format lets multiple engines interact with the same tables. This flexibility reduces the need for duplication and enables seamless integration into multi-engine environments, offering accessibility from multiple platforms. From my perspective, a future-proof implementation that does not require a full migration when you switch engines.<\/p>\n<\/div><\/div><div class=\"et_pb_text_7 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h2>ACID<\/h2>\n<p>In large-scale operations, ACID compliance is often required to guarantee consistency and support concurrent transactions. Teams have encountered challenges when managing concurrent writes, especially in high-traffic environments. Many data engineers built custom workarounds to have a \"soft\" guarantee of ACID,\u00a0 like appending data only and updating old data to be set to a different archive state.<\/p>\n<p>Each transaction in Iceberg \u2014 INSERT, DELETE, and even ALTER (schema change) \u2014 generates a new snapshot. This mechanism avoids partial updates, keeps historical versions accessible, and ensures independent operations.<\/p>\n<p>Iceberg uses optimistic concurrency control instead of traditional locks. Each writer checks the latest version before committing changes, and if there is a conflict, retries with the latest snapshot. This should reduce the performance cost of locking while supporting high-concurrency environments.<\/p>\n<\/div><\/div><div class=\"et_pb_text_8 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h2>Schema Evolution: Adapt to scaling<\/h2>\n<p>As your datasets evolve, they often require schema changes. Iceberg can keep pace with these changes without rewriting data files. Each column has a unique field ID, so Iceberg can apply changes like renaming or reordering without impacting stored data. This setup ensures schema changes won't break existing queries, saving resources in fast-evolving datasets.<\/p>\n<p>By storing versions in metadata files historical queries are accommodated while allowing new queries to use updated schemas. Schema evolution can be managed without disrupting downstream processes, a critical capability for efficient data operations.<\/p>\n<\/div><\/div><div class=\"et_pb_text_9 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h2>Migrating to Iceberg<\/h2>\n<p>This section explores the benefits and planning required to migrate to Iceberg for various current data platforms:<\/p>\n<ul>\n<li>Databricks offers native support, so the migration is straightforward. Spark SQL can be used to query Iceberg tables directly, and existing delta lake tables can be converted, enabling more flexible schema evolution and version control.<\/li>\n<li>Bigquery has <a href=\"https:\/\/cloud.google.com\/blog\/products\/data-analytics\/announcing-bigquery-tables-for-apache-iceberg\">announced support,<\/a> so migration should be easy enough (soon). The BigQuery API and SQL can be used to access your Iceberg tables.<\/li>\n<li>Redshift has a small extension called Spectrum, that allows it to query S3 and essentially turn Athena-like. Both spectrum and athena support Iceberg tables.<\/li>\n<li>Snowflake would require external tables, so a slightly harder migration is required. Essentially, you would have to export an internal table to cloud storage, and then turn it internal.<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<\/div><\/div><div class=\"et_pb_text_10 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h2>Why should I migrate?<\/h2>\n<p>There are several reasons to adopt Iceberg. The architecture reduces query time by scanning only relevant files, translating to significant performance gains. Iceberg supports multiple processing engines, reducing your reliance and lock-in to a specific platform. The schema evolution capabilities offer a streamlined solution to manage and query evolving data, especially useful when interacting with data sources out of your control or without a proper SLA. We discussed Iceberg's data lineage for each table, and coupled with Apache Atlas that turns into a full lineage and tracing solution.<\/p>\n<p>Of course, migration to Iceberg introduces more complexity into your system, which might not be warranted especially for smaller organizations. The catalog and metadata layers require resources that may not justify the gains.\u00a0<\/p>\n<\/div><\/div><div class=\"et_pb_text_11 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h2>Conclusion<\/h2>\n<p>Apache Iceberg redefines large-scale data management, offering a solution for complex datasets that's generic, robust, scalable, and flexible. Compelling reasons for this backend are ACID compliance, performance optimization, schema evolutions, and multi-platform operations. It's not a one-size-fits-all solution, but it addresses a critical need in modern data lake(houses) and supports moving platforms in the future.<\/p>\n<p>Iceberg offers a compelling opportunity to revolutionize and future-proof the data platform. Is your organization ready to take the leap into modern data management?<\/p>\n<\/div><\/div><div class=\"et_pb_text_12 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p>References and Further Reading:<\/p>\n<ul>\n<li><a rel=\"noopener\" target=\"_new\" href=\"https:\/\/iceberg.apache.org\/\"><span>Apache<\/span><span> Iceberg<\/span><span> Official<\/span><span> Documentation<\/span><\/a><\/li>\n<li><a rel=\"noopener\" target=\"_new\" href=\"https:\/\/atlas.apache.org\/\"><span>Apache<\/span><span> Atlas<\/span><span>: Data<\/span><span> Governance<\/span><span> and<\/span><span> Lineage<\/span><\/a><\/li>\n<li><a href=\"https:\/\/docs.aws.amazon.com\/glue\/latest\/dg\/aws-glue-programming-etl-format-iceberg.html\" title=\" AWS Documentation AWS Glue User Guide AWS Documentation AWS Glue User Guide Using the Iceberg framework in AWS Glue\"><span>Introduction<\/span><span> to<\/span><span> Apache<\/span><span> Iceberg<\/span><span> on<\/span><span> AWS<\/span><span> Glue<\/span><\/a><\/li>\n<li><a href=\"https:\/\/medium.com\/insiderengineering\/how-we-migrated-our-production-data-lake-to-apache-iceberg-4d6892eca6e6\">How We Migrated Our Data Lake to Apache Iceberg<\/a><\/li>\n<li><a href=\"Why Iceberg beats Delta Lake\">https:\/\/www.starburst.io\/blog\/why-apache-iceberg-databricks-delta-lake\/<\/a><\/li>\n<\/ul>\n<\/div><\/div><div class=\"et_pb_text_13 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<\/div><\/div><\/div><\/div><\/div>\n","protected":false},"excerpt":{"rendered":"","protected":false},"author":3,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-351","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts\/351","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/comments?post=351"}],"version-history":[{"count":19,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts\/351\/revisions"}],"predecessor-version":[{"id":556,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts\/351\/revisions\/556"}],"wp:attachment":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/media?parent=351"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/categories?post=351"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/tags?post=351"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}