{"id":603,"date":"2026-07-01T23:47:22","date_gmt":"2026-07-01T23:47:22","guid":{"rendered":"https:\/\/joskodeboer.nl\/?p=603"},"modified":"2026-09-06T21:18:46","modified_gmt":"2026-09-06T21:18:46","slug":"603","status":"publish","type":"post","link":"https:\/\/joskodeboer.nl\/index.php\/2026\/07\/01\/603\/","title":{"rendered":"Levelling up your data engineering: Metadata-driven systems"},"content":{"rendered":"<div class=\"et_pb_section_0 et_pb_section et_section_regular et_flex_section preset--module--divi-section--default\">\n<div class=\"et_pb_row_0 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_0 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_text_0 et_pb_text et_pb_bg_layout_light et_pb_module et_flex_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h1>Levelling up your data engineering: Metadata-driven systems<\/h1>\n<p>In a discussion around the value generative AI is bringing to teams, and why there's such a big difference in productivity gains, I was reminded of the situation at a large enterprise client I had a few years ago - one of the largest companies in the Netherlands. I was a part of the data platform team, managing infrastructure for around 5 <em>data pipeline<\/em> teams with a goal of scaling that to around 30 of these teams.<\/p>\n<p>In this blog post, I want to introduce you to metadata-driven abstractions.<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_1 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_1 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_image_0 et_pb_image et_pb_module et_flex_module preset--module--divi-image--default\"><span class=\"et_pb_image_wrap\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/metadatasystems.jpg\" title=\"metadatasystems\" width=\"1408\" height=\"768\" srcset=\"https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/metadatasystems.jpg 1408w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/metadatasystems-1280x698.jpg 1280w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/metadatasystems-980x535.jpg 980w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/metadatasystems-480x262.jpg 480w\" sizes=\"(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) and (max-width: 1280px) 1280px, (min-width: 1281px) 1408px, 100vw\" class=\"wp-image-614\" \/><\/span><\/div>\n\n<div class=\"et_pb_text_1 et_pb_text et_pb_bg_layout_light et_pb_module et_flex_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p><em>Caption: Metadata-driven abstractions will turn large amounts of repetitive constructs into a clean, generated pipeline based on metadata.<\/em><\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_2 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_2 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_text_2 et_pb_text et_pb_bg_layout_light et_pb_module et_flex_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><h2>Data pipelines<\/h2>\n<p>I think one of the earliest definitions of data engineer was 'someone who creates data pipelines'. In the history of BI, that started with real, physical servers. One would extract data from all matter of source systems, the next would transform it, the final one would load it to the database system (ETL).<\/p>\n<p>Of course, around a decade ago that all moved to the cloud. Physical systems disappeared, and the notion of ETL tooling was introduced. A rapid addition was scaling database engines that allowed us to move transformation into the engine itself, introducing ELT.<\/p>\n<p>Honestly, nowadays I usually just refer to them as 'ingestion' tools. These tools are a part of the data platform, and they take data from source systems and move it to the database engine. In a modern platform, transformation logic is usually done inside the database engine separate from the tool moving the data, so it doesn't make sense to call it an ELT tool. An EL tool sounds silly. And using the term ETL tool to refer to a tool that doesn't do transformations and has the order wrong is just confusing.<\/p>\n<h2>Data platforms<\/h2>\n<p>A data platform is many parts, usually displayed as something like this:<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_3 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_3 et_pb_column et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_12_24 et_flex_column_12_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_row_4 et_pb_row et_pb_row_nested et_flex_row\">\n<div class=\"et_pb_column_4 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_image_1 et_pb_image et_pb_module et_flex_module preset--module--divi-image--default\"><span class=\"et_pb_image_wrap\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/lakehouse_solution_architecture-scaled.png\" title=\"lakehouse_solution_architecture\" width=\"2560\" height=\"1920\" srcset=\"https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/lakehouse_solution_architecture-scaled.png 2560w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/lakehouse_solution_architecture-1280x960.png 1280w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/lakehouse_solution_architecture-980x735.png 980w, https:\/\/joskodeboer.nl\/wp-content\/uploads\/2026\/07\/lakehouse_solution_architecture-480x360.png 480w\" sizes=\"(min-width: 0px) and (max-width: 480px) 480px, (min-width: 481px) and (max-width: 980px) 980px, (min-width: 981px) and (max-width: 1280px) 1280px, (min-width: 1281px) 2560px, 100vw\" class=\"wp-image-607\" \/><\/span><\/div>\n<\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_column_5 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_12_24 et_flex_column_12_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_text_3 et_pb_text et_pb_bg_layout_light et_pb_module et_flex_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p><span style=\"color: #808080;\"><em>Caption: Coral marks where data enters the platform (ingestion), blue marks where it's stored (object storage and table format), amber marks the compute that acts on that storage (workloads, dbt), and purple marks where it's served out (APIs, dashboards). Gray bands aren't stages in the flow \u2014 they're control, security, and metadata concerns that apply across all of the above rather than at one point in the pipeline.<\/em><\/span><\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n\n<div class=\"et_pb_row_5 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_6 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone\">\n<div class=\"et_pb_text_4 et_pb_text et_pb_bg_layout_light et_pb_module et_flex_module preset--module--divi-text--default\"><div class=\"et_pb_text_inner\"><p>The part I wanted to talk about: Most organisations have a lot of data hosted in similar source systems. This can be each source team having a relational database or a nosql database.<\/p>\n<p>In the setup at the previous client, they wrote each pipeline by hand. For this client, they'd have a business activity, then a domain team within that containing tables.<br \/>We'll use the shorthand `activity.domain.table` to refer to those.<\/p>\n<p>They'd write a very similar pipeline for each. I'll display it in a yaml form, because this is metadata - it is data describing the data you're trying to move:<\/p>\n<pre class=\"prettyprint\">pipeline:\n- id: retrieve_config_activity_domain_table_A\n  type: retrieve_config\n- id: extract_activity_domain_table_A\n  type: extract\n- id: load_activity_domain_table_A\n  type: load\n- id:  transform_activity_domain_table_A\n  type: transform\n<\/pre>\n<p>We can already see this can be templated extremely easily:<\/p>\n<pre class=\"prettyprint\">pipeline:\n- id: retrieve_config_{{ source_table }}\n  type: retrieve_config\n- id: extract_{{ source_table }}\n  type: extract\n- id: load_{{ source_table }}\n  type: load\n- id: transform_{{ source_table }}\n  type: transform\n<\/pre>\n<p>But even if we do that, it's still <em>a lot <\/em>of copies of the same data. In software engineering, we have the principle: <strong>Do not repeat yourself<\/strong>. In this case, that should clearly lead to a generator:<\/p>\n<pre class=\"prettyprint\">tables = [\n  'activity_domain_table_A',\n  'activity_domain_table_B'\n]\n\nfor source_table in tables:\n  yield \"\"\"\npipeline:\n- id: retrieve_config_{{ source_table }}\n   type: retrieve_config\n- id: extract_{{ source_table }}\n   type: extract\n- id: load_{{ source_table }}\n   type: load\n- id: transform_{{ source_table }}\n   type: transform\n\"\"\"\n<\/pre>\n<p>The 20 000 lines of boilerplate code have been reduced to a simple table listing each of your source tables together with the required config.<\/p>\n<p>That's the point I was getting to. The right **abstraction** here, as an engineer, is to use metadata to generate the pipelines. Even if your transform is handled by something modern like dbt, the point is that we don't want to spend our days copy-pasting repeated pipelines.<\/p>\n<p>We want to get rid of that, and focus on value: Onboard the data fast, and start building _intelligence_ on that. Whether that's business intelligence or data science, it doesn't matter - we don't want to focus on repetitive work. Even in the age of generative AI.<\/p>\n<h2>Generative AI<\/h2>\n<p>Repetitive, structured work - the ideal use case of generative AI. Your code base grows - for 2000 tables, suddenly you're looking at 20 000 lines of yaml that is utterly repetitive, but it doesn't matter because you're having an AI handle that - right?<\/p>\n<p>We shouldn't just deploy AI-generated code to production. How will a human be able to check all that _slop_? How do we make sure we can move systems fast, able to adapt to innovation?<\/p>\n<p>Amusingly, this might be one of the reasons why generative AI is so powerful to many people. Repetitive work that had already **should have been abstracted away** can now be taken over by AI. You're replacing intelligence with sponsored tokens, and one day the bill will come and your slop's unmanageable.<\/p>\n<p>Or you lose access to your <a href=\"https:\/\/www.anthropic.com\/news\/fable-mythos-access\">premier model<\/a> and suddenly, your business is no longer functioning.<\/p>\n<h2>Engineer intelligent systems<\/h2>\n<p>What matters is the abstractions you engineer into your systems, not the tool it's in. The abstraction you end up with should be readable, understandable and maintainable for your team. A metadata-driven abstraction will let your team focus on things that matter. Engineer intelligence.<\/p>\n<p>Keep your humans in the loop, and use generative AI to help you build an intelligent system that you can manage and leverage.<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n<\/div>","protected":false},"excerpt":{"rendered":"","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-603","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts\/603","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/comments?post=603"}],"version-history":[{"count":12,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts\/603\/revisions"}],"predecessor-version":[{"id":622,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/posts\/603\/revisions\/622"}],"wp:attachment":[{"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/media?parent=603"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/categories?post=603"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/joskodeboer.nl\/index.php\/wp-json\/wp\/v2\/tags?post=603"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}