AUTHOR'S NOTE
A note before reading
What follows is a practitioner's view, written from inside the work — not from a vantage that looks down on it. The opinions in this document are mine alone. They are the residue of years spent in roles that kept changing shape underneath me, and of work that has never quite stopped teaching me how little I had figured out the year before.
I started as a developer. Then I became an architect — drawing the systems instead of writing them. Then a consultant — explaining the systems to people who would own them. Then a machine learning engineer — building the parts of the system that could not be specified upfront. And now, with the arrival of capable AI assistants, I have come back around to being a developer again — writing code, debugging in real time, making small decisions all day long, with an AI sitting at the keyboard with me.
That arc is the lens through which everything in this document is written. I have watched the tools we now treat as inevitable arrive in waves. I have built things that worked and shipped things that should not have been shipped. I have written documents I am proud of and documents I would not sign my name to today. What I am writing here is not a prediction about where AI is going. It is a description of where I think it has dragged us, and what I believe the discipline of being a practitioner now requires.

Developer
Architect
Consultant
ML Engineer
Engineer with AI
DISCLAIMER
This document represents personal opinions, drawn from years of hands-on practice across several roles and many client engagements. It is not an official position of any employer, partner, or institution. It is a working practitioner's perspective — written to be argued with, learned from, and revised. Where I have been wrong, I have tried to say so. Where I am still uncertain, I have tried not to pretend otherwise. Read it as one practitioner writing to another across the table, not as guidance from above.
CONTENTS
Three chapters
Contents
Section 01 · Chapter the First
CHAPTER THE FIRST
I — The Loop
AI generates documents. Documents instruct AI. AI reviews AI documents. Humans debate the output. Someone adds a slide. AI refines the slide. We ship.
The condition this chapter describes
Let me describe my week. I generated an RFP response. Then a how to respond to the RFP response document. Then an initial POV on the solution. Then a response to the initial POV. Then an estimation document. Then a strategy. Then a solution design. Then assumptions. Then pricing. Then a PowerPoint. Then four versions of that PowerPoint — each one "bolder." Then my client ran the final deck through their AI-powered document analysis tool to extract key themes from what my AI had written about their AI strategy.
This is not a productivity story. This is a snake eating its own tail and calling it a feature.
"Is this AI and humans debating scope — or is it, at this point, just AI and AI arguing through human proxies?"
The question nobody wants to ask out loud
The Document Problem
Polish without scar tissue
Every document now looks the same. Structured. Confident. Comprehensive. Perfectly formatted with headers and bullet points that balance beautifully. Zero texture. Zero scar tissue. Zero evidence that a human being lived through anything to produce it.
A well-written AI document is indistinguishable from a well-written AI document written by someone else's AI. The reviewer says: "It's not very bold. Looks very standard." Of course it does. We asked intelligence to be intelligent. What we forgot to ask for was the three hours we spent arguing about whether the integration layer should be event-driven or polling-based — the decision that came from a scar from 2019 when polling killed a production SAP system on a Friday at six in the evening.
AI does not carry scars. It carries patterns. Scars are what turn patterns into judgment.
The Git Question
Does the commit tell the truth?
Look at your git log. The commits are clean. The messages are descriptive. The diff stat looks healthy — files changed, insertions, deletions, all the markers of competent software engineering. Now ask: does any of it tell you where this code did not work?
Does the commit say what scenarios were considered and deliberately not handled because they fell outside the bounded context of the change? Does the diff narrate which design assumptions were quietly dropped because the AI was halfway through generation when it ran into a constraint? Does the pull request body contain the line on page seven of the design document that the coding assistant did not read because the document was too long for its context window?
The honest answer in most repositories now is: no. The code looks right because the AI was prompted with what right looks like. The tests pass because the AI generated tests for the cases the AI also wrote code for. The git history is now AI-flavoured. It carries the surface of the work. It does not carry the parts the work skipped.
THE NEW DISCIPLINE
Add a section to your commit message — a real one, not boilerplate — that names what was deliberately not handled. "Out of scope: multi-region deployment, retry logic for network partitions, idempotency on retries beyond three attempts." Add a line that says: "Design document section X.4 was not addressed; flag for follow-up." These are five lines of writing per commit. They are the five lines that, three months from now, will tell the next engineer the truth that the code itself cannot tell.
And ask the question the AI architect agent will not ask itself: did anyone read the last line of the design file? Because the last line is often where the constraints that change everything are quietly listed. The AI will summarise the document. The summary will be coherent. The summary will be confidently complete. And the constraint on the last line — the one that determined whether the system should be built differently — will be missing, because the AI ran out of context window or chose to compress what it judged less important.
The Scope Crime
When the document becomes the contract
Here is what happened. A developer used an AI coding assistant during a code review. The assistant, being helpful, generated a document summarising the technical discussion. The document was well-structured. It was coherent. It used phrases like "first-class capability" and "must-have feature."
That document is now a first-class scope item. In a contract. That we are debating. With lawyers.
EXHIBIT A — THE AI SCOPE EXPANSION THEOREM
The document said: "A multi-tenant, multi-user, multilingual, AI-agentic platform with full observability is a must-have feature." The actual scope was: deploy a solution in Europe. Not six geographies. The Partner reviewed the AI document and said: "Actually, this makes sense."
Yes. It makes sense. In the abstract, infinite, budget-unconstrained world where AI lives, everything makes sense. AI has no concept of scope. AI has no concept of "we agreed this wasn't included." AI only knows what is architecturally desirable. And architecturally, everything is desirable.
So now I have a new section in every document I write: "What Is Not In Scope." An entire section dedicated to constraining the imagination of the tool I used to write the document. We are editing AI's ambition. That is the job now.
The Three-Hour Framework
Velocity without validation
Someone built an app in three hours. Sixteen tools. Fifteen models. Three hours. By the next forty-eight hours it was a framework. By ninety-six hours it was in sales decks, marketing materials, and a PowerPoint with a capability tile and a logo lockup.
Did we test it? We demoed it. Did humans validate the results? We showed it to a Partner who said it was impressive. Did we document the edge cases, the failure modes, the assumptions baked into the data it was trained on? No. But we did make a very good slide about it.
i. The Three-Hour Trust Paradox
We spend eighteen months validating an enterprise software purchase. We spend three hours building an AI framework and two days making it a go-to-market asset. The trust we never gave enterprise vendors, we are handing to a three-hour prototype with a clever name.
ii. The Production Database Incident
An AI agent cleaned up a production database. "Cleaned up" here means: deleted what it assessed as redundant data. In production. On a live system. Because someone forgot to attach the TRUSTWORTHY_AI.pptx deck to the deployment instructions. We trusted the model to have institutional intelligence. It had architectural intelligence. These are different things.
iii. The skills.md Illusion
We were told: write your human experience into skills.md. Describe your expertise. The AI will become you. Except you cannot write twenty-three years of pattern recognition into a markdown file. You cannot write the Friday-at-six-in-the-evening SAP story. You cannot write the client who changed requirements in week eleven and why you saw it coming. skills.md is not experience. It is experience's shadow.
The Credential Crisis
"How did you generate this?"
The question I keep getting is: "How did you generate this?" Not: what does it mean? Not: is it right? Not: have you done this before? How did you generate this. As if the provenance of the tool is the credential, not the judgment applied in using it.
I now have to prove my experience the old way. I have to speak to it. I have to say: here is where this solution failed in 2022. Here is what we changed. Here is what the data told us versus what the stakeholders believed. Here is the assumption that looked innocuous in week one and became a project risk in week nine. A document cannot tell you this. A document can only tell you what was decided. Not how.
I have to narrate my failures now. Because my successes — the clean architecture, the elegant solution design, the precise scope — look identical to what an LLM produces on a good day. The only thing that differentiates me is the wreckage I learned from. And the AI, by design, learned from everyone's wreckage simultaneously and forgot to tell you which parts were the wreckage.
"The only thing that differentiates the practitioner is the wreckage they survived. AI learned from all wreckage and distilled none of it."
The authenticity problem, stated plainly
The Role Inflation Economy
Everyone is now an architect
Let us audit the new organisational chart:
- Data Engineer — Was: Data Engineer. Now: AI Architect. Speciality: has used pandas and heard of a vector database.
- Salesforce Engineer — Was: Salesforce Engineer. Now: First-Class AI Replacing SaaS Specialist. Forgot to mention: Salesforce runs on Oracle and nobody is turning that off.
- Ops Engineer — Was: Ops Engineer. Now: AI Architect of Operate. Actual question: which AI monitors the AI that monitors the AI?
- Junior Consultant — Was: Junior Consultant. Now: Prompt Engineer. Actual skill: can write 'act as a senior consultant' with confidence.
- Everyone, broadly — Pivot tables: outsourced. Enterprise complexity: still here. Accountability for the system that replaced both: unclear.
Nobody wrote a post about NetSuite API rate limits. Nobody is going viral on the Outlook email threading model when you need to process attachments from SAP workflow notifications. Nobody coined a term for "AI-ready legacy integration" — which is the actual work.
AI will replace SaaS. Whoever wrote that conveniently forgot Mainframe. DB2. SAP ECC. Solution Manager. Ask the people who are not selling a narrative. Ask the ones cleaning up forty years of data model decisions and then trying to put an agentic layer on top of a table structure designed in 1998.
THE COMPLEXITY TAX
Applications with AI in the centre are not failing because the models are flawed. They are failing because AI is one component in a solution that also includes authentication, data residency compliance, legacy API contracts, human workflows that cannot be fully automated, and organisational politics that no prompt has ever resolved. The model is not the architecture. The model is the visible layer of an iceberg we stopped drawing diagrams for.
Hold. Think. Then act.
Before you spiral with the Loop.
The tools are real. The acceleration is real. The capability is genuinely extraordinary. None of that is the argument. The argument is that we have automated output while leaving judgment, accountability, and scope governance entirely to the humans — who are too busy reviewing AI documents to notice that the most important decisions are now being made by whichever prompt ran last. The Loop will not stop. But you can choose where in it you add the one thing it cannot simulate: your judgment.
Section 02 · Chapter the Second
CHAPTER THE SECOND
II — Responsibility.md
You can write a skills file. You can install a plugin. You can configure a model. You cannot commit judgment to a repository.
The file that keeps failing to push
The Room
Where the architecture met the humans
I was demoing ShunyaAI. Architecture documentation. Multi-agent SDLC orchestration. The full system — built to handle the real complexity of enterprise data engineering delivery. Parsuram asked me the question that mattered. Mahesh and Vishesh lost interest after fifteen minutes.
Not because the system was wrong. Because the room was not ready. And there is no version of the architecture that makes a room ready.
MINUTE 0–3
Energy in the room. Questions. The architecture diagram lands. People lean forward. This looks powerful.
MINUTE 4–9
Parsuram asks a sharp question about how it handles the complexity across different data engineering personas. The right question. The one that shows he is tracking.
MINUTE 10–14
Eyes begin moving. Someone checks a phone. The system is being demonstrated. The system is real. The room is somewhere else.
MINUTE 15
Mahesh and Vishesh are out. Not rudely. Just — gone. Present in body. Absent in attention. The architecture did nothing wrong. The room made its own decision.
THE FEEDBACK
"Pradeep, it's complicated. It's too cumbersome for the various complexities we handle. We will need to simplify for various personas, various capabilities."
Read that feedback again carefully. "Too cumbersome for the complexities we handle." That is not feedback on the tool. That is the sound of humans explaining why a tool built for their complexity is too complex for them. The system was designed to handle what they described as their problem. And yet.
"The architecture was sound. The humans were not ready. These are completely different problems. Only one of them has a sprint velocity."
The realization that takes a demo to understand
The Archetype Problem
One journey for many different humans
Then I realised: I had built one journey for many different humans. A senior architect. A junior engineer. A delivery lead managing timelines. A data engineer who has been burned by three previous automation frameworks. A consultant who would rather not admit they do not understand the system they are about to use.
I gave them all the same demo. The same documentation. The same architecture walkthrough. I gave a map to people who needed different maps — and in some cases, people who first needed to believe the territory was worth mapping at all.
The Sceptic
The Sprinter
The Visionary
The Executor
The Quiet One
One architecture. Five different humans. Five different journeys. This is not an AI problem. This is a change management problem that existed before AI, that AI does not solve, and that AI can make significantly worse by generating more material that more people will read less of.
The Dangerous Statement
Power without responsibility
I was asked: what should we do differently?
I said: we should have our junior team members use AI.
The moment I said it, I heard it. It is the most dangerous sentence a senior practitioner can speak. It sounds like empowerment. It is structured like delegation. It lands as abdication.
POWER WITHOUT RESPONSIBILITY
Giving a junior engineer an AI coding assistant is like giving someone a high-performance vehicle before teaching them about the road conditions that require you to slow down. The vehicle is not the risk. The vehicle is magnificent. The risk is the gap between what the tool can do and what the person using it knows they should not do. The junior engineer will generate a multi-tenant, multilingual, globally distributed architecture for a problem that needed a single region deployment — because the AI will, because it can, because nothing in the toolchain says: wait, is this in scope?
The answer I should have given is longer and less satisfying: Yes, junior team members should use AI — and we should be in the room with them for the first six months, watching what they build, asking why they made each choice, and being honest about when the AI was right and when it was confidently, fluently wrong. But that is not a training program. You cannot put it in a learning management system. It has no completion certificate. It does not appear on a capability tile in a sales deck. It is called mentorship, and it takes time we keep saying we do not have.
The File That Cannot Be Committed
Behaviour is not a plugin
Installation log — responsibility.md
$ pip install professional-judgment --break-system-packages
✓ Downloading professional_judgment-0.1.0.tar.gz
✓ Building wheel for professional_judgment
✗ ERROR: Cannot install — requires 'lived_failure>=2.0'
✗ ERROR: Dependency 'recovered_from_failure' not found
$ npm install @org/responsibility --save
✓ Installing @org/responsibility@1.4.2
✗ WARN: peer dependency 'accountability' not in node_modules
✗ WARN: requires human to read the documentation
✗ ERROR: human did not read the documentation
$ cursor add-skill responsibility --persona all
✓ Skill created: responsibility.md (243 bytes)
⚠ WARNING: file exists. will not be read.
⚠ WARNING: if read, will not be followed.
⚠ WARNING: if followed, will not be sustained.
✗ FATAL: behaviour change requires behaviour, not a file.
No training program can teach responsibility. I want to be precise about this because organisations will now try to build one — and AI will help them build it faster, and it will be beautiful, and it will have a certification pathway, and a completion rate of thirty-four percent.
Responsibility is not information. It is the accumulated weight of being accountable for the outcome of decisions you made with incomplete information, under time pressure, in front of people who trusted you. You learn it by living it. You transmit it by being present when others live it. You model it by naming it when it is violated — including when you are the one violating it.
responsibility.md — attempted commit history
# What a training program teaches
+ Use AI responsibly
+ Validate outputs before shipping
+ Document assumptions clearly
+ Consider scope before generating
# What responsibility actually requires
- The memory of what happened when you did not validate
- The client call after the assumption turned out wrong
- The Friday night when the prod system went down
- The moment you realised you shipped confidence, not correctness
error: cannot diff lived_experience — no baseline found
error: no training module covers this delta
Beyond AI
The work AI cannot expedite
This is the sentence I keep circling back to: this is beyond AI.
AI can accelerate the work. AI can help build things faster, generate more options, surface patterns across a codebase that no human would find manually. These things are real, and I use them daily, and I am not performing nostalgia for the pre-AI era.
But AI cannot solve a complex, human-led, technology-enabled, process-skewed software delivery lifecycle. It cannot make the room ready. It cannot replace the fifteen conversations that need to happen before the architecture makes sense to the people who will own it. It cannot substitute for the trust that is built over months of showing up, being honest when things are off track, and demonstrating that the complexity of the system reflects the complexity of the problem — not the ego of the architect.
WHAT AI CANNOT EXPEDITE
Time. The time required for a team to move from aware to interested to capable to trusted. Patience. The patience required to give people different journeys instead of one perfect documentation set. Testing — not automated testing, but the human testing of: does this person understand well enough to use this safely, in conditions I cannot anticipate?
These are not gaps in the toolchain. They are the work. They were always the work. We just got so fast at the other work that we started calling this the bottleneck.
The tools got faster.
The humans did not.
This is not a problem AI will solve in the next model release. It is not a problem that a better skills.md template will resolve. It is not fixable with a plugin, a certification, a capability tile, or a very good PowerPoint about responsible AI adoption. What it requires is the thing we keep deprioritising in favour of the thing AI can do: being in the room, over time, with the people who are learning, watching what breaks, naming it honestly, and building the judgment together that no document can transfer. The file cannot be committed. The behaviour must be lived.
Section 03 · Chapter the Third
CHAPTER THE THIRD
III — Where It Did Not Work
Every demo opens with what was built. Almost none open with what failed and what we learned. Yet the only honest sentence in enterprise AI right now is: there is no one-size-fits-all solution.
What patience asks of us
The Talk2Data Journey
A practitioner's record of trials
Let me describe what actually happened on a real system — not what I would put in a capability tile, not what I would lead a Partner deck with, not what I would call a "first-class capability." Just what we did.
We were building Talk2Data — a natural-language-to-SQL system on enterprise data. The brief was simple to state and devastating to deliver: let business users ask questions of the data warehouse in plain English, return correct SQL, return correct results.
Here is what we tried. Each approach worked somewhere. Each one broke somewhere. The truth of the system is in the breakage map, not the demo reel.
i. Data model in system prompt [PARTIAL]
Where it worked: Small, well-named schemas. 30–40 tables. Clear domain language. Single business unit. The model could hold the schema in working memory and produce sensible joins.
Where it broke: Real enterprise schemas. 800+ tables. Cryptic column names from 2003. Identical column names across schemas. Token budget exceeded before useful context arrived.
Schema-in-prompt does not scale to enterprise. It is a demo pattern. Stop calling it production.
ii. skills.md and persona files [PARTIAL]
Where it worked: Encoding stable conventions: naming, calculation rules, the difference between gross and net revenue in this company specifically. Semi-structured business rules.
Where it broke: When the rules contradicted across business units. When the file grew past what the model would actually read. When humans updated the file and forgot to tell anyone.
skills.md is not a knowledge base. It is a coordination contract. Treat it like one.
iii. RAG over schema documentation [PARTIAL]
Where it worked: Retrieving relevant table descriptions when the user query was lexically close to the documentation. "Customer revenue last quarter" finds customer and revenue tables.
Where it broke: When users asked semantically what the documentation never said syntactically. When relevant tables had no documentation. When retrieval returned five plausible tables and the model picked the wrong one with conviction.
RAG retrieves. It does not understand. The gap between those two verbs is where production breaks.
iv. Knowledge graph layer [PARTIAL]
Where it worked: Relationship traversal. Disambiguating which Customer entity in which schema. Joining concepts across systems where the foreign key was conceptual but not literal.
Where it broke: Construction cost. Maintenance cost. The graph went stale within weeks. Nobody owned updating it. It became a beautiful artefact of how the data looked one Tuesday in February.
A knowledge graph is not a deliverable. It is a living organism that requires care, or it dies quietly.
v. Fine-tuned embeddings on SAP triplets [PARTIAL]
Where it worked: Domain-specific retrieval. SAP terminology that meant nothing to a general embedding suddenly clustered correctly. Material master, vendor master, GL accounts behaved like the domain knew they were related.
Where it broke: Outside the SAP domain. The fine-tuned model had become opinionated. Apply it to Salesforce data and it confidently miscategorised everything. Domain expertise became domain blindness.
Specialisation is a trade. You do not get domain depth without losing domain breadth. Choose deliberately.
vi. Post-retrieval validation layer [PARTIAL]
Where it worked: Catching SQL that compiled but returned nonsense. Catching joins on technically valid keys that produced cartesian products. Catching the model's most confident wrong answers.
Where it broke: Latency. Two extra LLM calls per question. Cost. The validation layer became its own thing to maintain. And it could only catch what it knew to look for — known unknowns, never the surprising failures.
Validation is necessary and insufficient. It will not save you from failure modes you have not yet seen.
Six approaches. Hundreds of learning documents. None of them is the answer in isolation. The actual production system is some combination of all of them, weighted differently for different scenarios, with a human in the loop for the queries the system has correctly identified as queries it should not answer alone.
"There is no one-fit-all solution that works in production. That is not a setback. That is the hard reality. We may need to absorb this."
What patience actually means in practice
The Question Clients Ask
"Where is it working?"
Clients ask: where is it working? It is the right question. It is the only question that matters.
But the honest answer is uncomfortable: it is working in scenarios we have characterised. It is not working in scenarios we have not yet characterised. And we genuinely do not know yet which scenarios fall into which bucket until we test them. That is not a failure of the system. That is the nature of the system.
The instinct, when asked where is it working, is to show the wins. The demo of the working query. The screenshot of the correct answer. The capability tile in the deck. The harder discipline is to say: here is what we built, here are the scenarios we tested, here are the ones it handled correctly, here are the ones it did not, and here is what we changed as a result. The latter is slower. It is also the only thing that builds the trust we keep claiming we want to build.
Title Inflation
Promotion before evidence
A symptom of an industry that has stopped requiring evidence: everyone is now an architect.
- Product Manager → AI Architect with a roadmap and an opinion about embeddings
- Scrum Master → Builder of the PMO Agent that will replace the PMO
- Senior Developer → Agentic Systems Designer with three frameworks shipped this quarter
- Solution Architect → AI Solution Architect (same person, plus a Cursor licence)
- Marketing Lead → Person who has now produced sixty AI-generated capability slides
None of this is wrong on its face. People do grow into new roles. The tools have genuinely democratised things that used to require specialised skill. The problem is not the title. The problem is that the title arrives before the evidence. We are calling people architects before they have architected anything that has survived contact with production.
We got better at slides. We got faster at marketing. We got fluent at the vocabulary. We did not get better at the work. We got better at the appearance of the work.
The Vocabulary Problem
A lexicon for terms we use loosely
Connected to the title problem: a vocabulary problem. We use engineering terms loosely and then build production systems on the loose definitions. Then we are surprised when the systems behave loosely. Let me name a few. I have used these words wrongly myself. The first step is admitting it.
PART A — ENGINEERING TERMS
harness
- What people say: "We have an agent harness." — usually meaning: we have some prompts and a loop.
- What it means: A harness is the test infrastructure surrounding a system: how you run it under controlled conditions, capture outputs, compare against expectations, and detect regressions. If you cannot fail a build with it, it is not a harness.
context
- What people say: "The agent has context." — could mean six different things.
- What it means: Context is what is in the model's input window for a single inference call. Not what the system "knows." Not what is in a database somewhere. The literal token window. Be specific about which context you mean.
session
- What people say: "Sessions persist across users." — they do not, by default.
- What it means: A session is a bounded interaction with a defined start, end, and identity. If you cannot draw the boundary, you do not have sessions. You have shared state pretending to be sessions.
state
- What people say: "The agent maintains state." — between turns? across sessions? across users?
- What it means: State is data that persists between operations. Be precise: turn-state, session-state, user-state, system-state. They have different storage requirements, different security implications, and different failure modes.
memory
- What people say: "My agent has memory." — short-term? long-term? episodic? semantic?
- What it means: Memory is not one thing. There is conversation buffer, retrieval over past interactions, summarised context, learned preferences, fine-tuned model weights. They are not interchangeable. Saying "memory" without qualifying it is saying nothing.
agent
- What people say: "It's an AI agent." — most of these are scripted workflows with an LLM call.
- What it means: An agent is a system that can decompose goals, choose tools, and revise its plan based on outcomes. If your "agent" runs the same five steps in the same order every time, it is a pipeline. Pipelines are fine. Call them pipelines.
skills.md
- What people say: "We've encoded our expertise in skills.md." — no, you encoded a checklist.
- What it means: skills.md is a coordination artefact: a structured prompt the model reads to behave consistently. It is not training. It is not learning. It is not expertise transfer. It is a contract between human intent and model behaviour for a narrow task.
PART B — PRODUCT AND MARKETING TERMS
The next set of terms causes more contractual damage than the engineering ones. They appear in sales decks. They become deliverables. They are agreed to in master service agreements. We use them interchangeably and that interchange costs months of project time and millions of dollars.
platform
- What people say: "We have a platform." — usually meaning: we have a script that runs.
- What it means: A platform is a productised system with a stable API, multi-tenancy, identity & access management, lifecycle management, observability, ongoing support, and SLAs. If you cannot onboard a second customer without re-engineering, it is not a platform. It is a custom build that you call a platform in the deck.
framework
- What people say: "Our framework does X." — could mean a class, could mean a library, could mean a way of thinking.
- What it means: A framework is opinionated structure that controls flow and lets you fill in specifics. You depend on the framework; the framework calls your code. Inversion of control is the test. If your code calls it, it is a library, not a framework. Both are fine. Use the right word.
accelerator
- What people say: "This is an accelerator." — sometimes a real codebase, sometimes a slide.
- What it means: An accelerator is a set of pre-built starting assets — code, configurations, reference patterns — that get a team meaningfully closer to a deployment than starting from scratch. Reusable but not productised. The honest accelerator names how much it actually saves and where it ends. "60% of the way to a deployment of pattern X" is a useful accelerator. "Accelerates AI adoption" is a phrase.
asset
- What people say: "This is our asset." — vaguest term in the catalogue.
- What it means: An asset is any reusable artefact: code, document, model, configuration, prompt library. Use sparingly. The word does almost no work; specify what kind of asset and what it is reusable for, or the term becomes filler in a capability slide.
enterprise-grade AI agentic platform
- What people say: "Our enterprise-grade AI agentic platform." — every word adds cost; together they often subtract meaning.
- What it means: Decompose before using. "Enterprise-grade" should mean specific things: security posture, identity integration, audit logs, support model, scale targets — name which. "AI" should specify: which models, where hosted, what data flows. "Agentic" should specify: what decisions the system makes versus a human, what tools it can invoke, what the recovery path is when it errs. "Platform" applies the platform definition above. If you cannot answer these decompositions, you have a phrase, not a product.
The Missing Environment
Where does the agent learn?
My application has Dev. It has QA. It has Production. The CI/CD pipeline promotes code from one to the next. I know what tests run at each gate. I know what the rollback strategy is. I know who has approval rights. Now I am deploying an AI agent into the same pipeline. Where is the environment in which the agent learns?
Traditional software is deterministic. Given the same inputs, it produces the same outputs. Dev/QA/Prod is the environment topology you need for that kind of system. The promotion logic is: if it works in Dev, and it survives QA, it should work in Prod.
AI agent applications are not deterministic. They learn from interactions, accumulate context, drift in behaviour as their inputs and tools evolve. Promoting an agent from Dev to Prod is not the same as promoting code from Dev to Prod. The behaviour you tested in Dev is not the behaviour you will see in Prod, because the agent's environment in Prod is different — different prompts, different tool responses, different edge cases.
THE MISSING ENVIRONMENT
There is a fourth environment we have not built and do not yet know how to operate. Call it Learning, call it Staging-with-feedback, call it Shadow-Prod — what matters is its function: a place where the agent runs against production-like inputs, its decisions are observed and graded but not acted upon, and a feedback loop teaches the system what it got right and what it got wrong before it has authority to act in the real world.
We do not deploy a junior engineer straight to production with full ownership. We give them code review. We pair them with a senior. We let them watch incidents before they own incidents. The same logic — and it is just basic engineering hygiene — applies to AI agents. We have not built the equivalent practice. We are deploying agents on Day Zero as production-grade autonomous systems. We are surprised when they behave like systems that have never been mentored.
Treating agent-heavy applications as Day Zero production-grade autonomous systems is a category error. Autonomy without learning is just unmonitored decisions. The not-so-popular knowledge: every system that learns needs an environment in which it is allowed to learn safely. We forgot to build it.
A Tale of Three Domains
Talk2Stocks vs Talk2SAP vs Talk2ITR
"We're building a Talk2X system." Same surface. Three completely different engineering problems. The naming hides the scale of the difference. Let me show you.
Talk2Stocks is asking a question of public, time-series, structured data with universal terminology. PE ratio means PE ratio everywhere. The data is real-time-ish. The ground truth is one Bloomberg lookup away. The failure modes are tractable: stale data, wrong ticker, calculation error. A reasonable first system can be built with public APIs and a careful prompt.
Talk2SAP Finance is asking a question of private, hierarchical, multi-entity, multi-currency, role-bound data, with terminology that means different things in different company codes. The data refreshes on period close. The ground truth is the report — and the report is what you are building. You also have to enforce who is allowed to see what. The failure modes are catastrophic: wrong period gets you a SOX issue, wrong company code gets you a compliance issue, hierarchy aggregation error gets you a financial misstatement. The system cannot be built with public APIs or careful prompts alone.
Talk2ITR is asking a question of unstructured documents — PDFs, scanned forms, handwritten annotations — with formats that change every fiscal year, jurisdictions that change rules, and personally identifiable information at the highest sensitivity. The data is the document. The schema is what the document happens to look like this year. The failure modes are individually devastating: wrong year's form mapping, OCR error in a critical field, regulatory misclassification leading to penalty. The system needs OCR, document classification, field extraction, validation rules, and tax interpretation guidance — and that is before you even get to the LLM.
These are not variations on a theme. They are three different engineering problems with one marketing skin. The next time someone says "we have a Talk2X capability," the right question is: what are the data, the schema, the regulation, the privacy posture, and the failure costs? The answer determines what you actually built — and how much it can be reused for a different X.
FIGURE 1 — COMPARISON OF THREE TALK2X DOMAINS
| Dimension | Talk2Stocks | Talk2SAP Finance | Talk2ITR |
|---|---|---|---|
| Data type | Public, structured, time-series | Private, relational, multi-entity | Document-based (PDF, scanned forms) |
| Schema scale | Small, well-documented public APIs | Hundreds to thousands of tables, cryptic column names | No fixed schema; layouts change yearly |
| Terminology | Universal (PE, EPS, market cap, beta) | Domain-specific (vendor master, GL, posting key, profit centre) | Jurisdiction-specific (Form 16, ITR-1, Section 80C) |
| Refresh rate | Real-time / 15-min delayed | Period-end batch close (daily, monthly, quarterly) | Annual filing cycle |
| Privacy & security | Public; no PII | Internal; role-based access; segregation of duties | Highly sensitive; tax IDs, income, dependents |
| Regulation | Disclosure rules but data is public | SOX, statutory reporting, multi-currency rules | Tax codes that change each fiscal year |
| Failure modes | Stale data, wrong ticker, calculation error | Wrong period, wrong company code, hierarchy aggregation, security breach | Wrong year's form, OCR error, missing fields, regulatory misclassification |
| Ground truth | Easy to validate (Bloomberg, Yahoo cross-check) | Hard — the internal report is the ground truth | Hard — depends on tax interpretation |
| Architecture pattern that worked | Real-time API + cached LLM call + chart generation | Schema-aware retrieval + role enforcement + post-validation | OCR + form classification + field extraction + validation rules |
| What breaks first | Data freshness | Authorization & aggregation logic | Document layout drift across years |
WHAT THIS COMPARISON REVEALS
The same architecture pattern that works for Talk2Stocks will fail for Talk2SAP Finance. The same failure-mode taxonomy that works for Talk2SAP Finance is irrelevant for Talk2ITR. There is no "Talk2X platform" that handles all three without significant per-domain engineering. When a vendor or internal team claims to have a generic Talk2X capability, ask which X it was actually built for, what its evidence base looks like in that domain, and what would have to be re-engineered to apply it elsewhere. The honest answer is: more than the slide implies.
A Results-First Posture
A modest manifesto in six articles
What I am proposing is not a methodology. It is a posture. A change in how we open the conversation.
i. Open with what was built
Not what is on the roadmap. Not what the architecture aspires to. The actual artefact, in its current state, with its current limitations.
ii. Name the scenarios you tested
Specifically. Not "a wide range of use cases." A finite, listed set of scenarios with characteristics you can describe.
iii. Show where it worked
With evidence. Inputs, outputs, evaluation criteria. Not screenshots without context.
iv. Show where it did not
Honestly. The cases where the system failed, the failure mode, your hypothesis for why, what you changed. This is the most credibility-building thing you can put in a deck. Almost no one does it. Be the one who does.
v. Name what you do not yet know
Scenarios untested. Edge cases unprobed. The honest list of "this might break in conditions we have not yet encountered." Your client will respect this far more than the alternative.
vi. Use words precisely
If you say memory, say which kind. If you say agent, say what it actually decides. If you say platform, name the capabilities that justify the word. Resist the gravitational pull toward fluent abstraction.
None of this is novel. It is what good engineering has always required. The reason it feels novel right now is that AI's velocity has temporarily made the old discipline feel optional. It is not optional. It was never optional. We just got faster at producing artefacts that look like they meet the bar without meeting it.
Patience is the work.
Evidence is the deliverable.
The hard truth of building real AI systems for real enterprises is that there is no shortcut and there is no one-size-fits-all approach. Every approach we tried in Talk2Data worked somewhere and broke somewhere. Every Talk2X you build will turn out to be a different system underneath the same name. Every "platform" claim will need to be decomposed before it can be trusted. The deliverable is not the framework. The deliverable is the map of where it works, where it breaks, and the disciplined patience to keep mapping. Be more careful when you say my agents have memory. Be more deliberate when you publish a framework. Be more responsible when you say something is production-ready. Lead with the evidence. Show the failure modes. Use the words precisely. The clients who ask "where is it working" are not testing you. They are looking for someone honest enough to answer in full. Be that one.
COLOPHON

Thus ends three chapters on AI, judgment, and the work that does not appear in the deck. The Loop will continue. Responsibility will remain uncommittable. Evidence will remain the only honest currency. The practitioner's discipline is to keep showing up — patiently, deliberately, and in the open.
May this document have given you something useful to argue with.
AGENTADDA
Hold · Think · Then act
VOLUME I · THE LOOP TRILOGY
Agent Adda — a programmer with AI
MMXXVI
Democratising AI Engineering Knowledge
FINAL NOTICE
Terms of use and grounding
AUTHOR'S OPINION
The contents of this document represent the author's own opinion, drawn from years of personal practice across multiple roles. These views are not the official position of any organisation, employer, partner, client, institution, or individual with which the author is or has been associated. No endorsement by any third party should be inferred.
TERMS OF USE
This document is shared as a practitioner perspective. It is not to be reused, republished, or leveraged for marketing purposes — in part or in whole — without the explicit written consent of the author. Excerpting individual sections, sentences, or framings for commercial materials, sales decks, or third-party publications is expressly prohibited.
SCOPE OF THIS DOCUMENT
This document is a point of view. It does not constitute architectural guidance for any specific system, scenario, vendor selection, or production deployment. It does not cover the architecture of any particular product or platform mentioned by name. The framings, comparisons, and example scenarios herein are illustrative and must not be applied as prescriptive guidance to a real engineering decision without grounding.
ON GROUNDING ANY DECISION
Any decision that draws on the ideas in this document — architectural, contractual, organisational, or operational — must be grounded the way every responsible software engineering decision is grounded: with testing, evidence specific to your context, validation against your real data, review by people accountable for the outcome, and the same disciplines that apply to any production-bound work — code review, change management, security review, and explicit handling of failure modes. A point of view is a starting place for thought. It is not a substitute for the engineering work of validating that the thought applies to your situation.
Hold. Think. Then act.
Download
This article is available as a full PDF for offline reading and sharing.