Contents
Section 01 · Preface
PREFACE
A New Philosophy for AI in Software Delivery
This is not a story about how AI wrote our software. It is a story about what we had to unlearn, dismantle, and rebuild before AI could become genuinely useful in a real delivery context. That is a harder and more honest story. We tell it here because the easier version -- the one where an engineer types a prompt and production code appears -- is a myth, and that myth is costing teams real time, real quality, and real trust.

ShunyaAI began as an act of deliberate experimentation. Not a research project in a controlled environment, but a live, high-stakes exercise in governed, multi-agent software delivery. The platform we built is real. The data engineering challenges it addresses are real. The failures we encountered, the patterns we discovered, and the philosophy we arrived at were all earned through direct, unfiltered contact with the variability and unpredictability of AI-assisted work at the edge of the possible.
Why Governed AI SDLC -- and Why Now
The data engineering industry is at an inflection point that most practitioners have not yet named correctly. AI coding assistants are not productivity tools layered onto an existing process. They are a fundamental challenge to the process itself. When a tool can generate a complete SAP-to-Databricks Bronze extraction pipeline in seconds, the question is no longer 'how do we write faster code?' It is 'how do we know the code is right, traceable, defensible, and improvable?'
The answer is not more tooling. The answer is governance -- not as bureaucracy, but as the structural foundation that makes AI-generated work trustworthy. Without it, speed becomes liability. With it, speed becomes a compounding advantage.
A Governed AI SDLC means that every phase of delivery -- requirements, architecture, design, build, test, deploy, operate -- has an owner, a contract, an evidence trail, and a gate. The AI participates at every phase, but it does not own any phase. The human does not disappear from the process. The human becomes the architect of a more capable, more disciplined process than was possible before AI existed.
Governance is not what slows AI down. The absence of governance is what slows delivery down -- it just does so three months later, when you cannot trace why a reconciliation is failing, whose specification assumption caused it, or how to prevent it recurring in the next wave.
Why Agent Adda
Agent Adda is a name with a philosophy inside it. An adda is a gathering place -- a corner of a city where people come to argue, think aloud, challenge each other's assumptions, and leave with clearer ideas than they arrived with. We chose that name deliberately.
The AI SDLC conversation happening across the industry is too often a monologue. One tool. One model. One way of working. Agent Adda was founded on the belief that the real learning comes from the argument -- from putting multiple assistants in the same room, giving them different roles, watching them conflict, resolving those conflicts through governance, and capturing what was learned so the next project does not start from zero.
We are practitioners, not researchers. We built ShunyaAI the way a real team builds a real product -- under resource constraints, against real delivery pressures, with real source systems that behave badly, and with AI assistants that are genuinely powerful but genuinely unpredictable. Agent Adda is the practice group that holds those experiments, extracts the patterns, and makes them available to the next team that does not have time to make every mistake from scratch.
Breaking the Myth
The dominant narrative about AI coding assistants rests on three myths that our experiments broke systematically.
Myth 1: The AI knows what you want. It does not. It knows what you said. When what you said is ambiguous, incomplete, or contradicted by a prior decision in the same codebase, the AI generates something plausible that is wrong in ways that are difficult to detect. The fix is not a better prompt. The fix is a frozen specification that removes ambiguity before generation begins.
Myth 2: Tests prove the AI's output is correct. Tests prove the AI's output is consistent. If the specification is wrong, the generated code implements the wrong behaviour, and the generated tests validate the wrong behaviour. The chain is only as strong as the specification at its head. This is why ShunyaAI treats the STTM, the ADR, and the API contract as the source of truth -- and why changing a frozen specification requires a Change Request, not a chat message.
Myth 3: One AI assistant is enough. Different assistants have genuinely different strengths and failure modes. Claude Code bootstrapped fearlessly but sometimes over-reached. Copilot executed reliably but needed strict ownership boundaries. Cursor reviewed with detail but produced findings faster than they could be actioned. Codex reasoned deeply but conflicted with other agents' file ownership. The winning pattern was not one universal assistant. It was a governed ensemble of roles -- each bounded, each accountable, each contributing to a shared evidence trail.
Delearning, Deconstructing, Building New
The hardest part of building ShunyaAI was not the technology. It was the delearning. Every experienced engineer carries a mental model of how software delivery works that was formed without AI. That model is not wrong -- it is incomplete. The question is which parts of it to keep, which parts to discard, and which parts to replace with something genuinely new.
We had to delearn the instinct to treat AI output as a starting point and human review as the quality gate. In a governed AI SDLC, the specification is the starting point, and AI generation against a frozen specification is not a draft -- it is a certified implementation or a deviation that must be explained.
We had to delearn the separation between delivery artefacts and knowledge. In a traditional SDLC, the code is the deliverable and the documentation is the afterthought. In a governed AI SDLC, the specification is the deliverable. The code is generated from it. The knowledge accumulated from delivery -- every resolved defect, every approved ADR, every incident resolution -- is not a byproduct. It is the asset that makes every future delivery better than the last.
We had to deconstruct the assumption that governance and agility are opposites. They are not. Governance, properly designed, is the thing that makes agility sustainable past the first sprint. Without governance, velocity compounds defects. With governance, velocity compounds knowledge.
What we built in its place is a philosophy, not a framework. A philosophy that begins with the premise that AI is a powerful collaborator in a structured system -- and that the structure is what gives the power a direction. Specification before code. Decision before design. Context before generation. Evidence before certification. Knowledge before the next project.
The Variability Problem
Anyone who has worked with LLMs in production knows the variability problem. The same prompt, given to the same model, produces different outputs on different runs. The same model, given the same context in a different order, reasons differently. The same assistant that produces brilliant code in one session produces subtly wrong code in the next.
This is not a bug to be fixed. It is a fundamental property of probabilistic systems that must be designed around. The governance model in ShunyaAI is largely a response to this variability. Every generated artefact is validated against a deterministic specification. Every ADR makes a decision explicit so the next generation does not relitigate it. Every defect is classified to a tier so the fix is applied at the right level and not just the most convenient one.
The agentic harness was built specifically to test variability at the UX level. Can a delivery lead -- simulated as an LLM persona -- navigate a workspace it has never seen before and understand whether the system is ready? Can a QA engineer persona find the defect evidence it needs? Can a data architect persona complete a STTM review without getting stuck? These are questions that deterministic test scripts cannot answer. They require the same kind of intelligent, variable, exploratory behaviour that real users bring. The harness provides it systematically, repeatably, and cheaply.
Why These Experiments Are Key to the Future
The experiments documented in this thesis are not academic. They are the early evidence base for a set of practices that will, in our view, define how serious software delivery teams work with AI over the next decade.
The current wave of AI tooling is providing speed. The next wave must provide trust. Trust requires traceability -- the ability to answer, for any piece of generated code, what specification it implements, what decision governed that specification, what requirement drove that decision, and what test validates the result. ShunyaAI implements that chain end to end. Every function references its specification. Every specification references its ADR. Every ADR references its requirement. Every requirement has acceptance criteria that become test cases.
The compounding intelligence model is the other piece of evidence that matters. Traditional software delivery is a largely flat learning curve -- each project draws on the expertise of the individuals involved, and when those individuals leave, the expertise leaves with them. A governed AI SDLC with systematic knowledge capture is a genuinely different model. The platform learns from every project. The defect patterns, the architectural decisions, the validated code templates, the operational runbooks -- all of it is retained, indexed, and made available to the next team before they have even written their first backlog item.
The proof of concept in this thesis is modest -- one platform, one team, seven weeks of intense iteration. But the numbers tell a meaningful story. 516 commits. 869 backlog items. 96 harness reports. 5,807 harness screenshots. 13 million tokens of exploratory test evidence. A wiki that documents every module, every route, every decision. A collaboration model that four different AI assistants could operate within without creating chaos.
The question we set out to answer was not "can AI build software?" It clearly can. The question was "can AI build software that a real team can trust, trace, extend, and hand to the next team without starting from zero?" The answer, we now believe, is yes -- but only inside a governed delivery system, and only when the team is willing to do the harder work of building that system before they start generating.
How to Read This Document
This thesis is structured as a working record, not a polished retrospective. It documents what we built, how we built it, what broke, what we learned, and what we would do differently. The evidence is real -- the screenshots are from the live application, the volumetrics are from the actual repository, the harness reports are from real test runs.
The Executive Summary that follows gives the platform overview. The subsequent sections move from origin and philosophy through the collaboration model, the backlog evolution, the knowledge pipeline, the testing architecture, and the product journey screenshots. The Lessons Learned section is where the philosophy becomes practical. The Conclusion points forward to the next phase -- knowledge-first enhancement -- which is where ShunyaAI's approach moves from building new systems with AI to governing change in existing ones.
A note on authorship: This document was written by Pradeep Gorai, based on the work of Agent Adda. Some sections were drafted with AI assistance -- and that is appropriate, because this entire thesis is about learning to work with AI honestly. Every claim in this document is grounded in evidence from the repository.
Section 02 · Executive Summary
EXECUTIVE SUMMARY
ShunyaAI was built as both a product and a research exercise in governed, multi-agent software delivery. The product is a phase-gated SDLC platform for data engineering and agent-assisted application build-out — combining project planning, requirements, architecture, STTM, source analysis, logical and physical modelling, code generation, test management, deployment, operations, knowledge management, and LLM-assisted workspaces.
The build process is as important as the product itself. We did not begin with a perfect system, a complete architecture, or a polished UX vision. We began with a clear problem: data engineering delivery is fragmented, context is lost across phases, knowledge is scattered across people and documents, and AI coding assistants need stronger governance to contribute safely to real systems.
The Practical Answer: A Governed AI SDLC
Backlog as Memory
Use a growing backlog as executable product memory — every item carries a contract with acceptance criteria, dependencies, and evidence.
GitHub as Platform
Use GitHub as the shared knowledge and audit platform — commits, PRs, Actions, wiki, and reports all become the durable record.
Multi-Agent Roles
Use multiple coding assistants with distinct roles and personas — each bounded, each accountable, each contributing to a shared evidence trail.
Engineering Infrastructure
Treat collaboration rules, skills, prompts, commands, runbooks, logs, tests, and reports as engineering infrastructure — not afterthoughts.
Human Ownership
Keep the human product owner in control of priority, conflict resolution, and product judgement — AI participates, human decides.
Agentic Testing
Build agentic testing — because normal tests cannot fully validate behaviour, UX, workflow coherence, and LLM-assisted outcomes.
ShunyaAI evolved from a problem-first idea beginning in mid-March 2026 into a broad platform with hundreds of backlog items, thousands of files, a formal collaboration agreement, an Obsidian-style wiki graph, and an LLM-driven agentic harness that tests the experience from persona-based viewpoints.
Section 03 · The Origin
THE ORIGIN — MID-MARCH 2026
ShunyaAI began not with a product brief but with a frustration. In mid-March 2026, Pradeep Gorai was revisiting a pattern he had encountered too many times across data engineering engagements: the context problem. Every project started from scratch. Senior engineers re-discovered the same source system quirks. Business rules lived in people's heads, not in documentation. AI coding assistants could write code but had no mechanism to write the right code — because 'right' depended on context that nobody had encoded anywhere.
The question that launched ShunyaAI was deceptively simple: what if the delivery system itself had memory? Not personal memory stored in a single engineer's head, but institutional memory that survived handoffs, compounded across projects, and could be given to an AI assistant as structured context before it generated a single line.
The Name: Zero as a Beginning
ShunyaAI takes its name from shunyaa — the Sanskrit word for zero. For me, the name is personal. We began our careers in data engineering from zero — no inherited systems, no domain expertise handed down. Just complex source systems, tight deadlines, and the expectation that the data would flow.
Zero is not an absence. In Sanskrit and in mathematics, zero is the origin point — the concept that made all subsequent numbers possible. ShunyaAI builds from zero in the same spirit. Not because the field was empty, but because the existing approaches were not working. The platform begins where conventional delivery methods run out.
The Agent Adda Concept
'Agent Adda' is the frame for how ShunyaAI was built. An adda is a place where people gather, debate, challenge, and refine direction. In ShunyaAI, the adda was not only human conversation. It was a multi-agent engineering room built around GitHub, collaboration files, backlog rows, wiki pages, review logs, and test evidence.
From the start: The build process was designed to mirror the product philosophy — agents, contracts, gates, logs, and evidence. We were not just building a collaboration platform; we were using one to build ShunyaAI.
Section 04 · What We Wanted to Build
WHAT WE WANTED TO BUILD
We wanted to build a platform for agentic delivery of data engineering and application work. The product vision was not simply 'an AI chatbot for software delivery.' It was a governed workspace where project intent, business requirements, data structures, source systems, mappings, generated artefacts, tests, reviews, deployment evidence, and operational knowledge remain connected throughout the entire delivery lifecycle.
The Core Problem Areas
Context Rediscovery
Data engineering teams spend too much time rediscovering context that someone else already captured — every project starts from scratch.
Fragmented Tooling
Requirements, STTM, source profiling, logical models, physical models, generated code, test cases, deployment records, and defects are stored in different tools with no traceable connection.
AI Without Contracts
AI assistants can generate useful code, but they need contracts, context, constraints, and verification to generate the right code.
Enhancement Blindness
Enhancement work against existing systems is hard because current system knowledge is incomplete, stale, or undocumented.
UX Judgement Gap
UI and UX judgement is hard for teams that are primarily data, AI, and platform engineers — it requires a different kind of evaluation.
Agentic Test Gap
Testing agentic applications requires more than unit tests and static Playwright scripts — it needs intelligent, exploratory, persona-driven validation.
ShunyaAI therefore became a product with two linked identities:
- A governed SDLC platform for projects, artefacts, phase gates, knowledge, and agents.
- A research platform for learning how multiple AI coding assistants can collaborate under human governance.
The Delivery Cycle
We knew the problems before we knew the product shape. That distinction mattered. Rather than designing everything upfront, we allowed the product to emerge from repeated evidence-driven cycles:
| Step | Input | Output |
|---|---|---|
| 1. Identify pain point | Delivery experience, engineering friction | Candidate backlog item |
| 2. Capture as backlog item | Pain point understanding | Contract with acceptance criteria |
| 3. Build thin vertical slice | Backlog contract | Working code against the contract |
| 4. Test with API, UI, workflow | Working slice | Evidence of correctness |
| 5. Review UX | Working screens, harness reports | UX findings, accessibility gaps |
| 6. Convert gaps to backlog | Review findings | Next sprint's contracts |
| 7. Preserve learning | All of the above | Compounding system knowledge |
Core principle: That cycle is the real engine of ShunyaAI. The product is not the output of a design phase. It is the accumulated output of evidence-driven iteration.
Section 05 · Governed AI SDLC
GOVERNED AI SDLC
The phrase 'Governed AI SDLC' means that AI can participate across the delivery lifecycle — but every important step has an owner, contract, evidence trail, and gate. Without governance, AI assistants amplify both productivity and inconsistency. With governance, they become useful contributors inside a bounded operating model.
Setup → Requirements → Architecture → Build → Test → Deploy → Operate
Governance Principles
- Backlog before implementation: Meaningful work gets a B-* item or is tied to an explicit reviewer issue.
- Role ownership: Implementers write implementer logs. Reviewers own reviewer queues. Test agents own test gaps.
- GitHub as audit: Commits, branches, PRs, Actions, and generated reports become the durable record.
- Human priority: The human product owner decides priority and resolves disputes.
- Testing evidence: Work is not complete because an assistant says it is — only when evidence exists.
- Knowledge refresh: Major work updates docs, wiki, or reports so future agents inherit the learning.
The Multi-Assistant Collaboration Model
Each assistant developed a working persona. The operating model made those roles explicit — not as a constraint but as the mechanism that made parallel AI work coherent.
| Assistant | Persona | Primary Role | Main Risk |
|---|---|---|---|
| Claude Code | Early implementer | Bootstrapping features, large edits, initial momentum | May over-change without tight contracts. |
| GitHub Copilot (Optimus) | Primary implementer | Backlog execution, feature delivery, tests | Needs strict ownership boundaries. |
| Cursor | Reviewer and test lead | Review queue, test strategy, UX review | Findings must convert to backlog rows. |
| Codex (Jetfire) | Reasoning partner | Debugging, planning, harness and thesis work | Must not conflict with other file owners. |
| GitHub Actions / CI | Non-human gatekeeper | Repeatable checks, security scans, e2e gates | Can fail late if contracts drift. |
| Human product owner | Final authority | Priority, product taste, conflict resolution | Bandwidth is scarce — keep decisions crisp. |
The lesson was not that one assistant is best. The lesson was that different assistants behave like different engineering roles. The operating model must make those roles explicit.
Section 06 · The Backlog as Product Memory
THE BACKLOG AS PRODUCT MEMORY
The backlog began as a list of planned work and grew into the primary memory of ShunyaAI. By May 2026, it contained 869 unique B-* identifiers across features, epics, defects, design gaps, harness findings, architecture follow-ups, and knowledge-first enhancement work.
| Metric | Value | Notes |
|---|---|---|
| B-* IDs | 869 | Unique backlog items |
| Lines | 2,556 | Current backlog |
| Total lines | 3,919 | Including archives |
| Optimus sessions | 380 | Implementer log entries |
Growth Channel Examples
| Channel | Examples |
|---|---|
| Initial product vision | SDLC platform, projects, phases, gates, artefacts, agents. |
| Implementation discoveries | Missing APIs, stale UI states, schema drift, inadequate navigation. |
| Reviewer findings (Cursor) | Implementation gaps converted into open items and backlog rows. |
| Architect reconciliation | Cross-layer action items from Kaleen-style deep reviews. |
| Harness findings | UX and workflow issues with concrete reproduction evidence. |
| Research outcomes | Knowledge-first enhancements, planning depth, STTM governance, estimation. |
Backlog as contract: An assistant should not guess the product. It should execute a contract with evidence. Each backlog row contains: ID, title, priority, effort, area, acceptance criteria, status, implementation notes, evidence, and dependencies.
Section 07 · Progress Over Time
PROGRESS OVER TIME — MARCH TO MAY 2026
The product progressed through waves rather than a single top-down release. What began as a problem statement in mid-March 2026 became, by early May, a platform with 770 API routes, 16 distinct product workspaces, a full agentic test harness, and 5,807 harness screenshots.
| Period | Theme | Key Evidence |
|---|---|---|
| Mid-March 2026 | Problem definition and initial thesis | Core problem areas identified. Governed AI SDLC concept formed. ShunyaAI name and zero origin concept established. |
| Late March 2026 | Architecture and foundational decisions | First collaboration agreement. Agent role definitions. Core tech stack selected: FastAPI, Next.js, PostgreSQL, Chroma. |
| Early April 2026 | Platform structure and collaboration model | Collaboration agreement ratified. Role definitions firm. Early backlog (first B-* items) and logs established. |
| Mid April 2026 | Wiki, agents, data plane, testing structure | Wiki index created. Module pages, agent framework, platform KB, database and route catalogues. |
| Late April 2026 | Planning, STTM, source analysis, knowledge, harness | B-PLAN, B-STTM, B-KNOW, B-UX-AUD, B-TEST backlog sections active. |
| 27–30 April 2026 | Artefact spine, lineage, STTM governance, harness evidence | Agentic harness reports, screenshot capture, lineage backlog, STTM lifecycle work. |
| 1 May 2026 | Planning merge, estimation workbook, CI stabilisation | Planning PR merge, estimation XLSX, OpenAPI drift gates, modernisation CI fixes. |
| 2 May 2026 | Estimation UX hardening, deep harness scenario, lineage workspace | Estimation readiness tests, deep harness report, Data Lineage workspace, toast system. |
Iteration pace: On 1 and 2 May alone there were 50 commits — planning CI stabilisation, estimation, STTM, lineage, UX, and harness improvements all shipped in 48 hours.
Section 08 · Wiki Graph and Knowledge Pipeline
WIKI GRAPH AND KNOWLEDGE PIPELINE
The ShunyaAI wiki is an Obsidian-style knowledge graph under ShunyaAI/wiki/. It was designed to make the system know itself — so that future agents could answer: What modules exist? Which routes own a capability? What decisions have already been made? What flows define expected behaviour?
| Dimension | Count | Examples |
|---|---|---|
| Modules | 27 | Knowledge base, Agent framework, Data Architect Chat, Source Analysis, Code Blueprint, Analytics Hub, Sprint Management |
| Concepts | 12 | Prompt safety, Audit events, Phase gates, Lineage, Knowledge retrieval |
| Decisions | 3 | Collaboration model, Testing strategy, Knowledge pipeline design |
| Entities | 4 | Project, Story, Agent, Knowledge item |
| Flows | 4 | Build flow, Knowledge ingestion, Agent execution, Harness run |
| Source pages | 8 | FastAPI (770 routes), Next.js (62 pages), PostgreSQL (153 tables, 3 schemas), 152 services |
For enhancement projects, this pattern becomes essential. An AI assistant cannot safely change an existing codebase unless it can first build a trustworthy baseline of what exists.
Section 09 · Volumetrics
VOLUMETRICS
Measured from the local repository on 2 May 2026:
| Metric | Value | Notes |
|---|---|---|
| Git commits (full history) | 516 | Concentrated in April–May 2026 |
| Commits since 1 May 2026 | 50 | Shows pace of recent iteration |
| Tracked files | 5,344 | git ls-files |
| Unique B-* backlog IDs | 869 | From docs/backlog.md |
| Backlog lines (current) | 2,556 | docs/backlog.md |
| Backlog including archives | 3,919 | Including done archives and pending |
| Agentic harness reports | 96 | Markdown report files |
| Harness screenshots | 5,807 | PNG evidence files |
| Summed harness tokens | 13,042,701 | Harness runs only — not coding sessions |
| Summed harness cost | $25.79 | Based on per-report estimates |
| Optimus implementer sessions | 380 | copilot-actions.md entries |
| Jetfire sessions | 45 | Codex-style implementer records |
Section 10 · The Testing Problem
THE TESTING PROBLEM
The biggest challenge was testing. ShunyaAI is not a simple CRUD application. It has APIs, database constraints, generated schemas, source-analysis artefacts, STTM rows, lineage relationships, LLM-generated drafts, planning baselines, estimation runs, UX workflows, cross-phase traceability, role-based navigation, knowledge retrieval, agent behaviour, and CI gates. Normal tests were necessary but not sufficient.
What Normal Tests Could Not Answer
- Can a delivery lead understand whether a workspace is ready?
- Does the UI expose enough evidence for a user to trust the state?
- Can a QA persona trace requirements to stories, STTM rows, tests, and gates?
- Does an LLM actor get stuck because the page lacks accessible actions?
- Are empty states truthful or misleading?
- Does the product experience match the mental model of data engineering delivery?
Test Layer Architecture
| Layer | What It Validates | Tools |
|---|---|---|
| Unit and component | Individual functions return expected values | Vitest, pytest |
| API and service | Endpoints accept and return correct shapes | pytest, httpx |
| Database and schema | Constraints, seeds, drift detection | Custom migration tests |
| OpenAPI contract | Client matches server definition | OpenAPI drift gates |
| Scripted Playwright | Deterministic UI flows pass | Playwright e2e |
| LLM agentic harness | Persona-based exploratory behaviour, UX coherence, workflow integrity | Custom harness + gpt-4o |
| Human and reviewer acceptance | Product taste, architectural coherence | Cursor review queue |
Agentic Harness — Estimation Deep Run Evidence
| Attribute | Value |
|---|---|
| Scenario | estimation-deep |
| Persona | delivery_lead |
| Model | gpt-4o |
| Phases completed / aborted | 5 completed, 0 aborted |
| Steps | 14 OK, 0 failed |
| Artifact checks | 3 checks, 0 failed |
| Estimate runs found | 2 (minimum 2 required — PASS) |
| Screenshots captured | 5 |
| Total tokens | 47,319 |
| Estimated cost | $0.1267 |
Harness value: It did not merely check that the page rendered. It checked whether a delivery lead could see Estimation readiness signals, run evidence, sign-off affordances, and export/handoff cues.
Section 11 · Application Screenshots
APPLICATION SCREENSHOTS — PRODUCT JOURNEY
The following screenshots document the ShunyaAI product journey from a deterministic Playwright capture at 1440x900 viewport. These are stable, named, and repeatable — the product as it stands at 2 May 2026.
Figure 1: Platform Dashboard

Cross-project landing surface. The dashboard shows 50 active projects across all archetypes (Greenfield, Enhancement, Modernization), the full capability grid including Requirements, STTM Engine, 10 AI Agents, Knowledge Assistant, Review Agent, Full Lineage, Phase Gates, and Accelerators. Navigation spans Dashboard, Knowledge, Assistant, Ops, Agents, Templates, Marketplace, and Settings.
Figure 2: Project Overview — SAP Finance Datalake

The project journey model showing all 8 steps (Setup, Planning, Requirements, Design, Build, Test, Deploy, Operate) with status indicators. The 'My Work' panel shows role-filtered action items for Developer, BA, Architect, QA, and Delivery Lead. Current step: Design at 50% complete. The overview is the primary orientation surface for every persona.
Figure 3: Planning Workbench

Plan Workbench showing a Baselined v1 plan. Tasks, stories, story points, and sprint assignments visible. Generate, Refresh, and Materialize bulk actions at the bottom. Lock Baseline control top-right. The planning assistant and sprint planning tabs provide AI-assisted plan generation and refinement.
Figure 4: Estimation Workbench

Effort model workspace with overview, inputs, task drivers, phases, runs, comparison, and sign-off. Monte Carlo estimation runs with XLSX export and scenario support. The estimation workbench provides the quantitative baseline for delivery lead sign-off before Build begins.
Figure 5: Requirements Workspace

Requirements and story capture in the SDLC spine. AI extraction from documents, KPI catalogue, and business rules register. Stories linked to requirements with traceability maintained throughout the delivery lifecycle.
Figure 6: Architecture Workspace
Design decisions, architecture evidence, and ADR management. Agent-assisted architecture review with structured output. The architecture workspace is where design decisions become traceable artefacts before any code is written.
Figure 7: STTM Workbench

Source-to-Target Mapping as a governed data engineering artefact. STTM authoring, lifecycle governance, comments, bulk actions, and DQ rules management. Every mapping row has a status, owner, and audit trail. The STTM is the specification that code generation operates against.
Figure 8: Source Analysis Workspace

Existing and source system understanding before design and build. Source profiling, schema discovery, and complexity scoring. The source analysis workspace provides the factual baseline that prevents the common failure mode of designing against incomplete source knowledge.
Figure 9: Logical Data Model
Conceptual and logical data modelling workspace. Domain entities, relationships, and business semantics captured before physical implementation. The logical model bridges requirements and physical design.
Figure 10: Physical Model

Target implementation shape and DDL-oriented design workspace. Physical tables, columns, constraints, and index definitions. The physical model is the implementation contract that code generation and testing validate against.
Figure 11: Data Lineage

Cross-artefact lineage, impact, and traceability workspace. Lineage graph from requirement to code to test, with refresh, quality, diff, and provenance APIs. Every change's downstream impact is visible before it is made.
Figure 12: Engineer Build Workspace

Agent-assisted build workspace and generated artefact flow. Code Blueprint and metamodel-driven generation. The Build workspace generates implementation artefacts from the frozen STTM specification, with the Review Agent providing pre-certification quality checks.
Figure 13: Testing Workspace

Test strategy, cases, defects, and verification context. API tests, database checks, and Playwright scripts alongside the agentic harness. Test cases are linked to STTM rows and requirements to maintain full traceability.
Figure 14: Operate Analytics

Operate phase, delivery health, and analytics workspace. Operational KPIs, incident tracking, and delivery analytics. The Operate workspace is where the platform closes the loop — incidents feed back into the knowledge base for future delivery intelligence.
Figure 15: Knowledge Manager

Knowledge ingestion, indexing, and retrieval operations. RAG-powered chat with project context, impact analysis, and citations. Seven knowledge base types: requirements, STTM, stories, documents, code, ADRs, and incidents — all indexed and semantically searchable.
Figure 16: Marketplace

Reusable accelerators, agents, and platform packaging. Pre-built STTM, DQ rules, and test cases for SAP, Salesforce, Oracle, and Workday. Domain accelerator packs reduce cold-start time on new engagements from days to hours.
Section 12 · Agentic Harness Evidence
AGENTIC HARNESS EVIDENCE
The following screenshots are from the estimation-deep harness run on 2 May 2026. They show the LLM actor navigating the platform as a delivery_lead persona, verifying that readiness signals, evidence, and sign-off affordances are surfaced correctly.
Figure 17: Harness Phase 1 — Estimation Entry

LLM actor (delivery_lead persona) entering the Estimation workspace. The harness verifies that the workspace is discoverable, that the navigation to Estimation is accessible, and that readiness signals are visible on entry.
Figure 18: Harness Phase 2 — Estimation Runs
LLM actor inspecting estimation run results. The harness verifies that run evidence is surfaced (2 runs found, minimum 2 required — PASS), that comparison data is visible, and that sign-off affordances are present.
Figure 19: Harness Phase 4 — Sign-Off and Export

LLM actor reaching sign-off and export stage. The harness verifies that the delivery lead can confirm estimation readiness and initiate XLSX export. 14 steps OK, 0 failed across the full run.
Section 13 · Lessons Learned
LESSONS LEARNED
1. Multi-Agent Work Needs Contracts
Multiple coding assistants can move quickly, but without contracts they can also create confusion. The collaboration agreement, runbook, backlog, locks, logs, and review queues made the work tractable. The O-* / B-* queue separation — reviewer issues versus backlog items — prevented the most common failure mode: an implementer declaring its own work accepted.
2. The Backlog Became Product Memory
The backlog started as a delivery tool and became the memory of the product. It captured not only tasks but also history, rationale, acceptance criteria, dependencies, and evidence. By May 2026, 869 B-* items represented the full intellectual history of every decision made since mid-March.
3. Testing Had to Become Agentic
API tests, database checks, unit tests, and Playwright scripts were necessary but not sufficient for a system that used LLMs. The agentic harness filled the gap between deterministic automation and human exploratory testing — 96 harness reports, 5,807 screenshots, and 13 million tokens of exploration evidence.
4. UX Emerged Through Evidence
Because the builders were primarily data and AI engineers, UX improved through screenshots, accessible actions, harness notes, reviewer findings, and repeated passes. Every harness failure classification — product defect, UX clarity gap, accessibility issue, missing stable selector — became a specific, actionable backlog item.
5. GitHub Was the Knowledge Sharing Platform
GitHub was not only source control. It became the shared memory layer for commits, branches, PRs, Actions, reports, docs, wiki, backlog, and collaboration records. The Obsidian-style wiki graph transformed that activity into queryable system knowledge.
6. Knowledge-First Enhancement Is the Next Frontier
Enhancement projects against existing systems require a trustworthy baseline: what code exists, what docs exist, what tests exist, what STTM and lineage exist, what is stale, what is missing. The B-KNOW-REV backlog section turns that need into an implementable programme — knowledge baseline contract, deterministic parsers, confidence model, cross-artefact linker, assessment engine, and promotion workflow.
7. AI Assistants Need Different Personas
The winning pattern was not one universal assistant. It was a governed set of roles. Claude Code bootstrapped. Copilot delivered. Cursor reviewed. Codex reasoned. CI gated. The human product owner held final authority. Each role had explicit scope, explicit outputs, and explicit evidence requirements.
Section 14 · Conclusion
CONCLUSION
ShunyaAI was built by accepting uncertainty and governing it. Starting in mid-March 2026 with a clear problem but no fixed product shape, it became — by May 2026 — a platform with 770 API routes, 16 product workspaces, 869 backlog items, 5,807 harness screenshots, and a collaboration model that made multi-agent software delivery tractable.
The answer was not to ask an assistant to build everything in one pass. The answer was to create a governed AI SDLC:
- Backlog as contract.
- Collaboration agreement as operating law.
- GitHub as knowledge and audit platform.
- Wiki as system memory.
- Skills and commands as repeatable agent behaviour.
- Tests and harness reports as evidence.
- Human product ownership as the final authority.
ShunyaAI is therefore both the application and the proof of method. It shows that agentic application build-out can work — when AI assistants are not treated as magic, but as collaborators inside a disciplined delivery system.
The next major step is knowledge-first enhancement: using the existing codebase, documents, tests, STTM, lineage, reports, and wiki to build a trustworthy baseline before any change request is planned. That is where ShunyaAI moves from building new features with agents to governing change in real existing systems.
Pradeep Gorai Agent Adda · Data Engineering Practice · 2026
Download
This article is available as a full PDF for offline reading and sharing.