Skip to main content
AgentAddaAgentAdda
All Articles
Data EngineeringAI AgentsKnowledge ArchitectureEnterprise AIModernization

The Heritage of Data Engineering: Will We Keep Carrying the Burden?

Every generation of data engineers inherits the technical debt, the undocumented decisions, the design drift, and the lost knowledge of the generation before it. We add new tools, new platforms, new architectures — and we carry the same burden forward, wearing new clothes. This is the manifesto for the generation that finally breaks the pattern.

15 January 202614 min read·AgentAdda Collective

We didn't learn this from textbooks. We learned it from the wreckage of projects that were supposed to work, and from the hard-won triumphs that surprised everyone, including us.

Here is where this manifesto actually begins.

Two engineers on a new engagement, seventeen years ago. They inherited an Informatica PowerCenter estate with no documentation — not a mapping description, not a design document, not a comment in the code. Mappings running nightly, producing numbers the business had trusted for years, for reasons nobody could explain.

Two months. One optimization task. Embedded lookups and SQL overrides nested so deep the nightly run took sixteen hours for twenty million records. The acceptance criteria: make it faster without changing the source count or the target count. Not "document it." Not "understand it." Just: don't disturb the numbers.

They worked entirely from inference — 120-line SQL queries, tuning join sequences, rewriting lookup logic, rearchitecting session flow. No specification. No history. Only the evidence of what the code did.

It worked. The hours came down. The counts matched. The business never knew it happened. The mappings went on running, still undocumented, passing their knowledge forward in the only form it had ever existed: the code itself.

This happened to us. It has happened to everyone reading this. The names change — Informatica becomes PySpark, SQL becomes a Databricks notebook, sixteen hours becomes a Spark job that nobody wants to touch — but the situation is identical. Inherited code. No documentation. Acceptance criteria that measure output, not understanding. Engineers doing extraordinary work in the dark.

The Hereditary Problem

Data engineering has a hereditary problem. Every generation of engineers inherits the technical debt, the undocumented decisions, the design drift, and the lost knowledge of the generation before it. We add new tools, new platforms, new architectures — and we carry the same burden forward, wearing new clothes.

This is not a failure of discipline. We have watched disciplined teams fail at documentation. We have seen rigorous governance frameworks quietly abandoned under delivery pressure within two sprints. The problem is structural: every delivery methodology ever designed optimizes for producing the artifact — the pipeline, the mapping, the model — and treats the knowledge embedded in it as a byproduct. Byproducts evaporate.

The pattern plays out the same way every time.

A data architect with eleven years of experience joins a complex SAP Finance migration program as lead. Experienced, thorough, well-regarded. On her first week she asks for the design documentation. She receives an ERwin diagram from 2014, a Confluence page last updated in 2019, and a spreadsheet with 340 rows labeled "STTM — FINAL v7." Nobody can explain why the intercompany elimination logic in the BKPF pipeline excludes records with document type ZR — the engineer who wrote that filter left two years earlier. "It's always been there. We don't touch it."

Six months later, during dual-run reconciliation on the new Databricks platform, the ZR exclusion causes a $4.2M discrepancy in the trial balance. The fix takes four hours. The investigation takes three weeks.

Will the generation building on Databricks and Snowflake today hand the same undocumented burden to the next generation — faster, cheaper, and more confidently wrong — or will we be the generation that finally breaks the pattern?

Six Generations. One Unresolved Inheritance.

The technology arc of data engineering spans six distinct generations. Each one was convinced it was solving the problem. None of them were — because none of them were solving the right problem.

Generation 0 — Shell Scripts (late 1980s–mid 1990s): Before there was ETL, there was the shell script. Unix cron jobs calling SQL loaders, writing results to a reporting schema. The entire pipeline was text files on a shared drive — or not checked in at all. This established the pattern that would persist through every generation: the code was the specification, the implementation, and the documentation simultaneously — and was none of them properly.

Generation 1 — SQL and PL/SQL as ETL (early 1990s–early 2000s): As enterprise databases matured, integration logic moved into the database itself. PL/SQL packages and T-SQL stored procedures became the primary ETL mechanism. Microsoft's DTS served SQL Server shops for seven years before SSIS replaced it — itself forcing thousands of packages to be rebuilt from scratch. ETL for the ODS. ETL for the data warehouse. Decisions made in the physical schema. Reasoning not recorded anywhere.

Generation 2 — Dedicated ETL Tools (late 1990s–present, still running): Triggered by enterprises accumulating source systems too heterogeneous for any single database to integrate natively. Informatica PowerCenter dominated. DataStage dominated financial services. Ab Initio occupied the largest financial institutions — technically superior, aggressively proprietary, with NDA coverage standard in its contracts. Visual design, server execution, metadata repository. Binary formats. Vendor lock-in. Metadata perpetually underpopulated.

Generation 3 — Big Data (2010–2018): The Hadoop ecosystem arrived with a technically correct answer to a question most enterprises were not asking. The question it answered: how do you process petabytes of unstructured data? The question most enterprises were asking: why don't my numbers agree, and why does every report take a week to build? Schema-on-read without discipline produced the world's most expensive data swamps. The lasting contribution was Spark itself, and the distributed storage patterns that became the foundation for Delta Lake and Apache Iceberg.

Generation 4 — Cloud-Native (2015–present): AWS Redshift, Azure Synapse, GCP BigQuery solved the infrastructure problem. Elastic compute. Managed services. Storage-compute separation. Infrastructure as code. The cloud generation inherited the Big Data era's data swamps, the ETL tool era's undocumented mappings, and the warehouse era's design drift. It migrated them, at cloud speed, to new platforms. The knowledge problem was not migrated. It reset to zero and began accumulating again from scratch.

Generation 5 — The Modern Data Stack (2019–present): Databricks unified data engineering and ML on a Lakehouse architecture. Delta Lake gave the data lake ACID properties. The medallion architecture (Bronze → Silver → Gold) gave structure to what had been undifferentiated. Snowflake provided the SQL-native cloud warehouse. dbt applied software engineering discipline to SQL.

What remained unsolved in every generation: why the pipeline was designed the way it was. The engineering generation solved the container. The knowledge inside — the design rationale, the architecture decisions, the operational learning — remains as ephemeral as it was in the shell script era.

The Design Hierarchy That Drifts

In a well-governed delivery, five layers of design stay synchronized: conceptual (ERDs, domain models), logical (dimensional models, schemas), physical (DDL, partition strategy), transformation specification (column-level mappings, business rules), and implementation (PySpark, dbt, stored procedures).

In every program we have been part of, changes happen at the implementation layer, propagate upward imperfectly, and leave the upper layers progressively out of sync.

Year one: the hierarchy is complete and approximately accurate. Year two: implementation changes have not been reflected upward. Year three: a new team reads the design documentation and finds it describes a different system from the one that runs. Weeks of investigation establish which discrepancies are intentional and which are drift.

The gap is the undocumented history of every business requirement change, every performance optimization, and every edge-case patch that modified the implementation without updating the design. It is the direct, compounding cost of the hereditary problem.

What Every Project Loses

Consider what happens at the close of a successful data engineering program. The team disperses. The architect who made the SCD decisions goes to a new client. The lead engineer who knows which SAP fields are unreliable in practice moves to a new organization. The QA lead who understands why the reconciliation threshold is 99.7% takes a promotion.

The platform runs. The knowledge does not.

Four types of knowledge evaporate, every time:

Design rationale: Why star schema over 3NF? Why delta load? Why SCD Type 2 for customers and Type 1 for products? The decisions persist in the physical schema. The reasoning does not. When conditions change — and they always change — no one knows whether the original reasoning still applies.

Operational knowledge: Which fields are unreliable in practice. Which edge cases surface under specific business conditions. Which failure modes recur. What the workarounds are. This knowledge lives in incident tickets when captured at all, disconnected from the design artifacts it relates to.

Architecture patterns: The PySpark pattern for a watermark-based delta load from SAP. The Airflow DAG structure for a multi-dependency financial close pipeline. These exist in the heads of experienced engineers. Junior engineers reinvent them from scratch. The knowledge exists. The mechanism for transfer does not.

Cross-program context: The organization that has delivered ten SAP-to-Databricks migrations has encountered the same source system behaviors, the same architectural trade-offs, the same edge cases, ten times. That accumulated understanding is its most valuable delivery asset. It lives in the heads of engineers who worked on multiple programs.

After ten programs, the organization has delivered ten times but compounded zero times. The tenth program is not materially smarter than the first about the things that matter most.

What AI Changes — And What It Does Not

AI does not solve the hereditary problem. It creates, for the first time, the technical conditions under which it can be solved. These are not the same thing.

What genuinely changes:

An agentic system that reads the approved design specification and generates code from it is qualitatively different from a generic AI coding assistant. Not better pattern completion — code generated from the specific approved specification of this program, in this organization, for this source system.

A vector store containing all past design decisions, architecture patterns, and code implementations — searchable by meaning rather than keyword — surfaces the relevant precedent in seconds. The knowledge has always existed. The retrieval now works.

A knowledge graph connecting every pipeline to its design specification, recent changes, and past incident history surfaces a ranked root cause hypothesis before the engineer finishes reading the alert. The 45 minutes of context-gathering disappears.

Each program contributing its decisions, patterns, and incident resolutions to a shared knowledge base means the second engagement is faster than the first. The fifth is dramatically faster. This is compounding delivery — never achievable before.

What does not change:

AI does not govern a specification that was never created. An agent generating PySpark from an absent specification generates from training data patterns, not organizational intent. The speed of generation becomes the speed of error propagation.

AI does not resolve the gap between the architectural diagram and the implementation that has diverged from it. It inherits the divergence and amplifies it.

AI does not substitute for organizational decisions about what the data means and what the architecture should do. The agent can present options. The architect must decide.

The hereditary problem deployed at machine speed is worse than the hereditary problem deployed at human speed.

The sequence is not negotiable: canonical specification first. Design rationale captured. Architecture decisions recorded. Then agents — and only then.

The 70/30 Principle

Most data engineering programs solve the same core problem expressed differently. SAP BKPF carries approximately 40 standard core fields — identical across every SAP installation. The Bronze-to-Silver transformation for financial posting data follows the same deduplication, currency conversion, and SCD Type 2 pattern in every organization.

The 70% is the same everywhere. The 30% is different everywhere.

Pre-built target data models (Bronze, Silver, Gold DDL). Column-level transformation specifications with pre-assigned load strategies. Per-column data quality rules. Architecture decisions pre-resolved: the standard SCD strategy, key strategy, load sequence, and orchestration pattern — already decided, validated across prior programs, documented with rationale.

The program spends its time on the 30% that actually differs — the Z-tables, the naming conventions, the regulatory rules specific to this organization and this domain.

What currently requires months of discovery, design, and build for the standard portions can be reduced to days when those standard portions have been built and validated across dozens of prior programs.

The Knowledge Flywheel

The goal is to transform programs from consumers of knowledge into producers of it — and to build the flywheel that results.

Each program, when run with disciplined knowledge capture, contributes to a compounding base: design decisions and their rationale, architecture choices and the alternatives considered, production incidents and their resolutions, data quality patterns and edge cases encountered.

The second engagement retrieves patterns from the first. The fifth engagement operates with accumulated intelligence from the previous four. The organization gets materially smarter — not because it hired better engineers, but because it stopped letting knowledge evaporate when programs end.

The knowledge base compounds. The debt compounds. There is no neutral position — only the choice of which direction to compound toward.

The Six Moves to Start

Not "what organizations should do." What practitioners would actually do, starting this week.

One canonical specification, this week. Pick one active pipeline. Produce a machine-readable specification: what it ingests, what it transforms, what rules govern it, what the architecture decisions are and why they were made. Not in a Confluence page. In a structured format an agent can read as input. This is unglamorous. It is the foundation.

Run the audit before choosing anything else. Map where the knowledge gaps are largest — which programs have the worst drift, the most expensive onboarding, the most repeated incident investigations. This tells you where the first AI deployment will have maximum credibility.

Three bounded accelerators, not a platform rollout. First: the program where knowledge fragmentation is visibly costing time. Second: give engineers back the 60–70% of their time spent on specification creation. Third: a modernization discovery exercise demonstrating what no previous tooling could.

Staff two new roles. Knowledge Architect: fluent in the business domain and the data architecture; captures decisions at the moment they are made. Delivery Quality Steward: validates AI-generated artifacts against specification and domain knowledge; the trust-builder for AI-assisted delivery.

Capture at the moment of creation. Every architecture decision recorded with its reasoning when made, not retrospectively. Every production incident resolution written to the knowledge base before the ticket is closed. Every defect traced to its root cause and linked to a test case that prevents recurrence.

Govern the specification, not the tools. Small specification councils with genuine authority. Centralize the standards. Federate the content. Every domain owns its definitions; the standards for how definitions are expressed are universal.

The 12-24-36 Month Arc

This is a 36-month investment in turning delivery from episodic into compounding. Treat it as a 6-month tool deployment and you will be running the same zero-knowledge programs in 2030 that you ran in 2020.

Months 0–12 — Foundation: Audit complete. Canonical specification for three to five programs. Accelerators deployed. Knowledge Architect and Delivery Quality Steward staffed. Template library established. AI governance v1: specification compliance, human gates, immutable audit trail. Outcome: design drift not occurring in accelerator-scoped programs.

Months 12–24 — Expansion: Top 15 programs on canonical specification. Knowledge marketplace active with cross-program retrieval. Modernization accelerator live. Knowledge corpus accumulating: architecture decision records, root cause analyses, defect patterns — being retrieved and influencing new programs. Outcome: onboarding time measurably reduced; patterns from earlier programs surfacing in architecture context for later ones.

Months 24–36 — Compounding: All strategic programs on the specification layer. Knowledge roles as established career tracks. Knowledge corpus as competitive asset. AI governance mature and auditable. Tacit expert knowledge systematically captured before practitioners leave. Outcome: "Why does this pipeline work this way?" answered in seconds. The next modernization program starting from this estate moves at a fundamentally different speed.

The Heritage Ends Here

We have carried this burden long enough.

Shell scripts. PL/SQL packages. Informatica mappings. DataStage jobs. HiveQL scripts. PySpark notebooks. dbt models. Different containers. Same knowledge problem. Inherited forward unchanged.

The technology has caught up with the diagnosis. For the first time, the technical conditions exist to build a delivery system where knowledge is not a byproduct — where every design decision, every architecture choice, every production incident, every defect and its resolution becomes a first-class artifact in a compounding knowledge base that makes every subsequent program smarter.

The organizations building compounding delivery intelligence today will be unreachable by the organizations still starting from zero in three years. Not because their engineers are better. Because their knowledge base compounds and the other organization's does not.

The heritage of data engineering is the accumulated burden of knowledge that was never systematically captured, never transmitted forward, and reconstructed — at enormous and invisible cost — by every generation that inherited it.

We are the generation that builds the system to end that pattern.


© AgentAdda.in — Practitioner Series · Data Engineering Practices · 2026 Edition

AgentAdda is a collective of data practitioners sharing honest insights on AI, data engineering, and enterprise transformation.

Back to All Articles