Levelling up your data engineering: Metadata-driven systems
In a discussion around the value generative AI is bringing to teams, and why there's such a big difference in productivity gains, I was reminded of the situation at a large enterprise client I had a few years ago - one of the largest companies in the Netherlands. I was a part of the data platform team, managing infrastructure for around 5 data pipeline teams with a goal of scaling that to around 30 of these teams.
In this blog post, I want to introduce you to metadata-driven abstractions.

Caption: Metadata-driven abstractions will turn large amounts of repetitive constructs into a clean, generated pipeline based on metadata.
Data pipelines
I think one of the earliest definitions of data engineer was 'someone who creates data pipelines'. In the history of BI, that started with real, physical servers. One would extract data from all matter of source systems, the next would transform it, the final one would load it to the database system (ETL).
Of course, around a decade ago that all moved to the cloud. Physical systems disappeared, and the notion of ETL tooling was introduced. A rapid addition was scaling database engines that allowed us to move transformation into the engine itself, introducing ELT.
Honestly, nowadays I usually just refer to them as 'ingestion' tools. These tools are a part of the data platform, and they take data from source systems and move it to the database engine. In a modern platform, transformation logic is usually done inside the database engine separate from the tool moving the data, so it doesn't make sense to call it an ELT tool. An EL tool sounds silly. And using the term ETL tool to refer to a tool that doesn't do transformations and has the order wrong is just confusing.
Data platforms
A data platform is many parts, usually displayed as something like this:

Caption: Coral marks where data enters the platform (ingestion), blue marks where it's stored (object storage and table format), amber marks the compute that acts on that storage (workloads, dbt), and purple marks where it's served out (APIs, dashboards). Gray bands aren't stages in the flow — they're control, security, and metadata concerns that apply across all of the above rather than at one point in the pipeline.
The part I wanted to talk about: Most organisations have a lot of data hosted in similar source systems. This can be each source team having a relational database or a nosql database.
In the setup at the previous client, they wrote each pipeline by hand. For this client, they'd have a business activity, then a domain team within that containing tables.
We'll use the shorthand `activity.domain.table` to refer to those.
They'd write a very similar pipeline for each. I'll display it in a yaml form, because this is metadata - it is data describing the data you're trying to move:
pipeline: - id: retrieve_config_activity_domain_table_A type: retrieve_config - id: extract_activity_domain_table_A type: extract - id: load_activity_domain_table_A type: load - id: transform_activity_domain_table_A type: transform
We can already see this can be templated extremely easily:
pipeline:
- id: retrieve_config_{{ source_table }}
type: retrieve_config
- id: extract_{{ source_table }}
type: extract
- id: load_{{ source_table }}
type: load
- id: transform_{{ source_table }}
type: transform
But even if we do that, it's still a lot of copies of the same data. In software engineering, we have the principle: Do not repeat yourself. In this case, that should clearly lead to a generator:
tables = [
'activity_domain_table_A',
'activity_domain_table_B'
]
for source_table in tables:
yield """
pipeline:
- id: retrieve_config_{{ source_table }}
type: retrieve_config
- id: extract_{{ source_table }}
type: extract
- id: load_{{ source_table }}
type: load
- id: transform_{{ source_table }}
type: transform
"""
The 20 000 lines of boilerplate code have been reduced to a simple table listing each of your source tables together with the required config.
That's the point I was getting to. The right **abstraction** here, as an engineer, is to use metadata to generate the pipelines. Even if your transform is handled by something modern like dbt, the point is that we don't want to spend our days copy-pasting repeated pipelines.
We want to get rid of that, and focus on value: Onboard the data fast, and start building _intelligence_ on that. Whether that's business intelligence or data science, it doesn't matter - we don't want to focus on repetitive work. Even in the age of generative AI.
Generative AI
Repetitive, structured work - the ideal use case of generative AI. Your code base grows - for 2000 tables, suddenly you're looking at 20 000 lines of yaml that is utterly repetitive, but it doesn't matter because you're having an AI handle that - right?
We shouldn't just deploy AI-generated code to production. How will a human be able to check all that _slop_? How do we make sure we can move systems fast, able to adapt to innovation?
Amusingly, this might be one of the reasons why generative AI is so powerful to many people. Repetitive work that had already **should have been abstracted away** can now be taken over by AI. You're replacing intelligence with sponsored tokens, and one day the bill will come and your slop's unmanageable.
Or you lose access to your premier model and suddenly, your business is no longer functioning.
Engineer intelligent systems
What matters is the abstractions you engineer into your systems, not the tool it's in. The abstraction you end up with should be readable, understandable and maintainable for your team. A metadata-driven abstraction will let your team focus on things that matter. Engineer intelligence.
Keep your humans in the loop, and use generative AI to help you build an intelligent system that you can manage and leverage.
