
Contents
Section 01 · Foreword
FOREWORD
"We don't write about tools. We write about what happens when tools meet reality — and what you need to know before that collision."
Agent Adda was built on a single conviction: the most valuable AI engineering knowledge lives in the field, not in model cards or conference keynotes. It lives in the audit logs of failed training runs. In the postmortems of agents that hallucinated at 3am. In the meeting where a client asked "can your AI just read the SAP docs?" and the practitioner had to explain why that's not how any of this works.
This Point of View is the third in our SDLC Disruption Series. The first two established why enterprise AI delivery requires knowledge-first architecture. This one asks the harder question: if AI agents can now pair-program across an entire software lifecycle — what does that actually look like, which barriers still matter, and what does the future hold?
We don't answer these questions from a whiteboard. We answer them from an NL→SQL fine-tuning project on SAP financial data, from a multi-regional requirements reconciliation, from an AI operations transformation engagement — and from the accumulated OQ tables and AUDIT logs of real delivery cycles.
The relay race metaphor is not decorative. It is the precise diagnosis of why enterprise software delivery is structurally inefficient — and why the agent pair is structurally its solution. What follows is the honest map.
Section 02 · The Core Argument
THE CORE ARGUMENT
The SDLC is a workaround, not a law of nature
Sequential software delivery was invented to manage the cognitive and communicative limitations of human teams. Those limitations no longer universally apply. The agent pair changes the physics.
The traditional SDLC — Requirements → Design → Build → Test → Deploy → Operate — is architecturally a relay race: sequential, handoff-dependent, and lossy by design. AI agent pairs operating with shared working memory, complementary cognitive modes, and goal-backward reasoning can collapse this relay into a closed loop. Every barrier that justifies the sequence — human knowledge limits, coordination overhead, siloed tooling, client communication friction — is now a solvable engineering problem, not a fixed constraint.
The claim is not that humans become irrelevant. The claim is that the structure of how software gets built must change — and that the change is already visible in production systems today, if you know where to look.
We will use a real case: fine-tuning an embedding model for natural-language-to-SQL on SAP financial data, where Cursor and Claude operated as a genuine pair programmer team across multiple sessions, with a shared audit log as working memory. The patterns generalize across every phase of the SDLC.
Section 03 · The Method
THE METHOD
Tree of Thoughts as a coordination protocol
Complex software is a tree of interdependent decisions — not a list of steps. Tree of Thoughts is the cognitive framework that lets agent pairs navigate this tree without losing coherence.
session-log.md — nl→sql embedding fine-tune · live audit trail
[GOAL: Recall@1 > 0.70 on weakest intents · NL→SQL Embedding Fine-Tune · SAP FI/SCM]
│
├── [CURSOR · Data Quality Branch] // fast, local, pattern-completing
│ ├── expand ANCHOR_PROMPT with intent examples
│ ├── wire phrasing_variant negatives in generators.py
│ ├── add --skip-llm-negatives CLI flag to orchestrator
│ └── BLIND SPOT: JARGON filter may destroy signal ← Cursor cannot see this
│
├── [CLAUDE · Data Quality Audit] // slow, goal-backward, domain-aware
│ ├── flags JARGON quality filter as semantic risk
│ ├── enforces: paraphrase → negative → quality (NOT quality first → silent data loss)
│ └── writes cursor-actions.md §Review rules ← shared working memory artifact
│
├── [CURSOR · Training Stability Branch]
│ ├── adds --triplet-margin to _train_triplet_loss
│ ├── adds --max-grad-norm ← only in triplet, NOT in MNRL sibling
│ └── bumps cross-encoder pairs: 300 → 600 ← ungrounded extrapolation, no data plan
│
├── [CLAUDE · Training Audit — Rule #2: sibling parity]
│ ├── catches missing --max-grad-norm in _train_mnrl
│ ├── proposes _fit_common_kwargs() refactor to share safety params
│ ├── reverts 600 → ~300 ← no data plan · skeptic role enforced
│ └── verifies all call sites reach training_meta.json
│
├── [CURSOR · Two-Pass Driver] // pure pattern-completion
│ └── writes run_v2_two_pass.sh (bash, env vars, MPS detection, bash -n validation)
│
└── [CLAUDE · Contract Definition]
├── Pass 2 must not start if modules.json absent ← safety gate
├── --smoke must reduce epochs AND skip eval independently
└── MPS defaults injectable via env vars ← CI/CD contract, never hardcoded

"Cursor's confidence is indistinguishable from its correctness. It writes gpt-40-mini and gpt-4o-mini with equal fluency. It bumps metrics from 300 to 600 with no hesitation. The pair programming contract must be explicit: Cursor implements, Claude audits before merge."
Section 04 · The Disruption Map
THE DISRUPTION MAP
Six phases. Six disruptions.
Every phase of the traditional SDLC has a specific set of assumptions that AI agent pairs invalidate. Here is the honest map — with the real failure modes when you don't design the pair correctly.
01 — REQUIREMENTS: The Assumption That's Already Dead
Requirements exist as a phase because translating business intent into engineering spec requires human-to-human negotiation — messy, slow, lossy. BAs interview stakeholders. Analysts write 60-page FRDs nobody reads.
Cursor failure mode: Produces 40 coherent-sounding user stories in 4 minutes. Internally consistent. Disconnected from actual constraints — exactly like the 300→600 extrapolation. Plausible. Ungrounded.
Claude's role: Goal-backward validation. Trace every requirement to a stated business outcome. Flag scope with no owner. Revert numbers with no data plan.
- CURSOR: enumerate
- CLAUDE: audit
- HUMAN: own the goal
02 — DESIGN: The Handoff That Eats Months
Architecture design is where the tree first becomes visible — and where human teams most consistently fail to maintain coherence. A solution architect makes a vector index decision without knowing the embedding model isn't finalized.
OQ-2 in practice: Upgrading from 384-dim bge-small to 768-dim bge-base invalidates the entire vector index. This cascade is tracked in the audit log. Cursor has no memory of it. Claude does.
The pair design session: Claude authors the decision tree. Cursor explores branches rapidly. Claude validates convergence points.
- CURSOR: prototype
- CLAUDE: convergence
- HUMAN: own the goal
03 — BUILD: Where the Pair is Most Powerful — and Most Dangerous
Build is Cursor's home territory and its most dangerous phase. The risk is not syntax errors. The risk is silent failures: a typo (gpt-40-mini) that passes py_compile but fails at runtime. Scope creep disguised as features.
The pair loop: Cursor implements → appends AUDIT entry. Claude reads AUDIT → reviews against plan → flags. Cursor reads review rules → self-validates next pass. No human required in every iteration — only at checkpoints.
The structural insight: Sibling parity is a design constraint no linter can express. It requires semantic understanding. That is Claude's contribution.
- CURSOR: implement
- CLAUDE: sibling audit
- HUMAN: define contract
04 — TEST: The Phase That Gets Skipped — Until It Shouldn't
Traditional testing comes after build because humans write code in one pass and test in another. The agent pair dissolves this boundary entirely. Tests are not a phase — they are a constraint encoded in the plan before code is written.
The smoke contract: --smoke must reduce epochs AND skip eval independently of device type. Written by Claude. Implemented by Cursor. Verified by Claude. Design-first testing — the opposite of how most teams operate.
Every numeric claim in the plan requires a falsifiable exit criterion before any code runs.
- CLAUDE: define exit criteria
- CURSOR: implement tests
- HUMAN: define contract
05 — DEPLOY: The Ceremony That Becomes a Script
Deployment is the phase most completely automatable by the current generation. Bash scripting, CI/CD configuration, environment variable management, device detection — pure pattern-completion work Cursor handles reliably.
The catch: Claude's deployment contract must be encoded before Cursor writes the script. MPS defaults must be injectable via env vars, not hardcoded. Pass 2 must not start if modules.json is absent. Design constraints, not implementation choices.
What 2027 looks like: Claude specifies. Cursor generates infra-as-code. Observation agent validates. Human approves checkpoints.
- CLAUDE: write contract
- CURSOR: generate infra
- HUMAN: approve
06 — OPERATE: Where the Closed Loop Closes
Operations has always fed back into requirements — but with a 6-month lag. A production incident triggers a postmortem, a ticket, the backlog, Q3 planning. The feedback loop exists in theory. In practice, it's broken by handoffs.
The agent pair's closed loop: Claude reads training_meta.json from each run. Detects R@1 decline on vendor-payment intents. Traces root cause to data quality drift. Generates a new OQ entry — a new branch. Cursor implements. Human reviews the audit trail.
The tree of thoughts becomes a living document. There is no hard boundary between "running the system" and "improving it."
- CLAUDE: monitor + diagnose
- CURSOR: patch + deploy
- HUMAN: approve changes
Section 05 · The Real Blockers
THE REAL BLOCKERS
The barriers are real — and solvable
Every enterprise engagement surfaces the same four objections. Here is the honest answer to each — not the vendor pitch, the field answer.
Human Barriers
The objection: "Our developers will not trust an AI agent to write production code." The real issue is not trust — it is legibility. Nobody can read a 40-file PR authored by Cursor in 8 minutes and know what's in it. The audit log pattern (AUDIT entry per session, OQ table per open question) solves this. It makes the agent's reasoning legible to the human reviewer. The barrier is the absence of a reviewable decision trail, not a fundamental capability gap.
Client Barriers
The objection: "Our client governance requires human sign-off at every phase gate." This is legitimate — and entirely compatible with agent-pair delivery if you treat human approval as a checkpoint, not a workflow. The client signs off on the ToT plan doc (Claude-authored). Cursor implements within the approved plan. The client reviews the AUDIT trail at the next checkpoint. The ceremony changes; the governance does not.
Process Barriers
The objection: "Our SDLC is ISO 27001 certified." Process certification is about evidence — can you demonstrate that the right checks were performed in the right order? An agent-pair with a structured audit log produces more evidence, more consistently, than a human-driven process where decisions live in meeting notes and Slack threads. The audit log is the process artifact. The question is whether your certification body can map it to their checklist — and that is a documentation problem, not a capability problem.
TOOL BARRIERS
The objection: "Cursor doesn't have access to our SAP system / Jira / internal wiki." This is the knowledge-first architecture problem. The agent pair is only as good as the knowledge it can access. The solution is externalizing your enterprise ecosystem knowledge into a form the agents can consume: a domain knowledge graph, a structured retrieval layer, a context document that travels with the session. This is a retrieval architecture problem, and it is solvable today.
Section 06 · The Agent Map
THE AGENT MAP
Claude and Cursor: two hemispheres
The mental model that makes this work: a single cognitive unit with two hemispheres — fast/associative and slow/deliberate. Neither can do the other's job.
Cursor — Fast Hemisphere
Autocompletes at thought speed — boilerplate, CLI wiring, env plumbing. Pattern-matches correctly across familiar code structures. Excellent bash scripting, device detection, file generation. Misses sibling function parity — sees local structure, not semantic contract. Inflates numbers — extrapolates plausible figures without data. Stateless across sessions — no memory of OQ-2, OQ-6 without the audit log. Confidence indistinguishable from correctness.
Claude — Slow Hemisphere
Goal-backward reasoning — traces every change to the stated plan. Catches semantic contracts Cursor is blind to (JARGON exemption, sibling parity). Skeptical on numerics — demands data or reverts the claim. Cross-session coherence via the audit log — knows the full decision tree. Authors plan docs, ToT trees, review rules, deployment contracts. Slow code generation — not optimized for boilerplate output. Overhead on purely mechanical tasks (bash, env wiring).

"The pair's most important artifact is not a code file. It is the audit log — the shared working memory that lets Claude audit what Cursor has done, and lets Cursor self-validate on the next pass. Without it, neither agent is trustworthy. With it, the pair operates like a team with a shared decision tree."
Section 07 · The Future State
THE FUTURE STATE
The SDLC as a closed loop
The future is not faster waterfall. It is a different topology — parallel, converging, continuously validated. Here is what the agent-pair SDLC looks like when it's running correctly.

Human checkpoints are injected at intent, build, and deploy — not at every phase transition.
What changes in Build — the pair loop in detail
Step 1 — Claude writes the review rules. Before Cursor touches the codebase, Claude authors explicit checks that Cursor must satisfy: sibling function parity, quantitative grounding, domain constraint compliance. These live in cursor-actions.md — the shared working memory.
Step 2 — Cursor implements. Works fast. Produces good boilerplate. Also makes its systematic mistakes: typos in model names, missing arguments in sibling functions, ungrounded scope additions.
Step 3 — Claude audits. Reading the AUDIT entries Cursor appended, Claude checks each change against the review rules and the plan doc. Flags violations. Proposes fixes. Updates the OQ table. Produces a structured review document — not a Slack message.
Step 4 — Cursor self-validates. On the next session, Cursor reads Claude's review output from the previous session. It self-validates against the rules before writing. The feedback loop runs without human intervention in the middle.
The human role: define the contract, approve the plan, review the audit trail at checkpoints — not in every loop.
REAL EXAMPLE · NL→SQL FINE-TUNE · OQ-2: THE 768-DIM CASCADE
Upgrading the embedding model from bge-small-en-v1.5 (384 dimensions) to bge-base-en-v1.5 (768 dimensions) was a design decision made mid-project based on R@1 ceiling evidence. The decision was correct. But it invalidated the entire vector index that Cursor had already built.
Without the audit log, this would have been discovered at integration time — a multi-day debugging exercise with no trail. With the audit log, OQ-2 was created the moment the dimension change was proposed: "bge-base upgrade blocked: vector index rebuild required before training can proceed." The blocked decision was tracked. The unblocking work was sequenced correctly. Total cost of the design change: 4 hours instead of 4 days.
This is not an AI story. This is a knowledge management story — the agent pair happened to be the most reliable knowledge management system in the project.
Section 08 · The Practitioner Action
THE PRACTITIONER ACTION
What to build right now
Theory without implementation is just a conference talk. Five specific artifacts — buildable on your next engagement, starting Monday.
1. The Audit Log
A single markdown file — cursor-actions.md — at the root of the repository. Every Cursor session appends a structured AUDIT entry: what was changed, what open questions were created or resolved, what review rules were applied. Claude reads this at the start of every review session. This is the shared working memory. Without it, the pair is two stateless agents. With it, the pair has institutional memory.
2. The Open Questions Table
A structured table embedded in the audit log: OQ-ID, description, blocked decisions, owner, resolution. Every unresolved design decision is a row. Every row has an owner — either Claude (will resolve on next review), Cursor (will implement once unblocked), or Human (requires domain knowledge or approval). The table is never empty while the project is running. Its length is a health indicator.
3. The Review Rules
A dedicated section of the audit log listing the semantic contracts Cursor must satisfy. Not linting rules — reasoning rules. "All sibling training functions must share safety parameters." "All numeric targets in plan docs must be derivable from data." "All model name strings must be verified against the provider's current API docs." These rules encode the patterns of Claude's catches — making the audit scalable and consistent across sessions.
4. The Plan Doc as Decision Tree
The project's plan document is not a gantt chart. It is a tree: goal to branch to sub-decision to constraint to exit criterion. Claude authors it. The human approves it. Cursor reads it. Every code change can be traced back to a branch in the tree. If it cannot, it is scope creep — and it gets reverted.
5. The Knowledge Layer
The agent pair is only as intelligent as the knowledge it can access. For enterprise projects: externalize client ecosystem knowledge — SAP transaction codes, business process maps, data dictionaries, integration contracts. This is the knowledge-first architecture problem. The knowledge layer is what separates a demonstration from a production system.
The teams that win the next five years of enterprise software delivery will not be the ones with the most AI licenses. They will be the ones who figured out how to make AI agent pairs legible — to clients, to governance bodies, and to each other. The audit log is not a developer tool. It is a competitive advantage. Start building it on your next engagement.
Section 09 · Horizon Scan
HORIZON SCAN
The possibilities are not incremental
We are not describing a 20% efficiency gain on the existing SDLC. We are describing a phase transition — from linear delivery to recursive, self-improving systems. Here is an honest map of what the next three horizons look like, grounded in current trajectory, not science fiction.
HORIZON 0 · IN PRODUCTION TODAY
The Audit-Log Pair
Cursor and Claude operating with a shared cursor-actions.md audit log as working memory. Cursor implements across sessions. Claude audits against the plan. Review rules encode the semantic contracts. The human owns the goal, approves the plan, and reviews at checkpoints. This is not a future state. This pattern is running on live enterprise engagements today — an NL→SQL embedding fine-tune, a multi-regional requirements reconciliation, an AI operations transformation. The infrastructure is simple. The results are not.
HORIZON 1 · NEAR-TERM · 12–24 MONTHS
The Self-Auditing Loop
The audit log and review rules become a machine-readable contract — not a markdown file but a structured schema that Cursor can query and Claude can enforce programmatically. The pair begins to self-correct without human-initiated review cycles. Claude monitors CI/CD outputs and flags spec deviations automatically. Cursor receives structured diff instructions and applies them in the same session. A third "observation agent" watches both agents and generates governance reports for client sign-off. The human's role shifts from reviewer to approver — still essential, but no longer in the loop on every decision.
Client governance adapts: audit trails become machine-readable evidence artifacts that satisfy ISO 27001, SOX change management, and FCA model risk guidelines — not because lawyers agreed, but because the format is richer than what human processes produced.

HORIZON 2 · MEDIUM-TERM · 24–48 MONTHS
The Knowledge-First Delivery System
The agent pair is no longer constrained by what is in the session context. A persistent, enterprise-wide knowledge graph — encoding client business processes, SAP configurations, integration contracts, regulatory constraints — is the foundation. Cursor and Claude query this graph before every decision. The knowledge-first architecture that practitioners are building manually today becomes a managed platform.
The SDLC no longer has phases. It has a continuous specification layer (Claude-maintained, human-approved), an execution layer (Cursor-driven, test-gated), and an observation layer (monitoring for spec drift, performance degradation, and regulatory exposure). Requirements, design, build, test, deploy, and operate are concurrent — different threads of the same loop, not sequential stages. Delivery velocity for a well-specified feature drops from weeks to days.
HORIZON 3 · LONG-TERM · 48+ MONTHS
The System That Improves the System
The observation layer closes the final loop: production telemetry, user behavior signals, and model performance metrics feed back into the specification layer automatically. The agent pair not only builds the system — it iteratively improves it without a human-initiated sprint cycle. Failure patterns in production trigger new branches in the plan tree. Claude proposes the fix. Cursor implements it. The observation agent validates it against the existing spec. The human approves changes above a defined risk threshold.
This is not autonomous AI replacing engineering teams. It is a fundamentally different model of human oversight: engineers set the goal, define the risk tolerance, and steer the evolution — but they are no longer the ones translating business intent into code, line by line. That translation has been automated. The remaining human work is judgment — which is, arguably, what it always should have been.
Six Bold Scenarios
Specific futures worth designing for
These are not thought experiments. Each is an extrapolation of a current, observable trend — what happens when the agent pair pattern scales to the enterprise, to regulated industries, and to the full software estate.
Enterprise Intelligence
Every enterprise system — SAP, Salesforce, Workday — ships with a Claude-mode agent that understands its own data model, business rules, and integration constraints. Cursor implements customizations against a formal spec. The agent pair replaces the SI consultant for tier-2 configuration work. Not entirely — judgment on business process design remains human. But the code that implements the judgment is generated, tested, and deployed by the pair. Horizon 1-2 — SI consulting market reshaped.
Continuous Compliance
Regulatory compliance — GDPR, FCA model risk, SOX change management — requires evidence that specific controls were applied in a specific sequence. The agent pair's audit log is that evidence, natively. The observation agent generates compliance reports as a side-effect of normal delivery, not as a separate effort. The compliance audit becomes a data query, not a 3-month exercise. Regulated industries — banking, pharma, healthcare — adopt the agent pair not for speed but for auditability. Horizon 2 — Compliance overhead near-zero.
Multi-Agent Orchestration
Cursor and Claude are the first two agents. The pattern scales. A security agent reviews every PR for vulnerability patterns. A performance agent benchmarks every build against SLA contracts. A domain agent encodes client-specific business rules. A UX agent validates interfaces against design system contracts. Each agent runs asynchronously, contributes to the shared audit log, and gates on the shared review rules. The human engineers coordinate the agent team — a fundamentally new skill set. Horizon 2-3 — New engineering org design.
Requirements Revolution
The agent pair's plan document becomes the primary interface between client and delivery team. The client describes intent in natural language. Claude structures it as a ToT decision tree with acceptance criteria. Cursor generates a working prototype in 48 hours. The client reviews working software, not wireframes. Feedback is captured in the audit log as new branches. The "requirements phase" collapses into a continuous conversation between client intent and running code — mediated by the pair. Horizon 1-2 — BA role redefined.
Legacy Modernisation
The single largest unsolved problem in enterprise IT is legacy modernisation — millions of lines of COBOL, PL/1, and ABAP that banks, insurers, and governments cannot safely replace. The agent pair changes the economics: Claude reads the legacy codebase and builds a formal spec of what it does. Cursor generates equivalent modern code. The observation agent validates behavioral equivalence. The human approves the migration segments. What currently takes 7 years takes 18 months. Horizon 2 — $1.5T legacy problem addressable.
Self-Healing Systems
Production incidents trigger Claude's diagnosis layer: read the logs, trace the failure to a branch in the plan tree, identify the divergence between spec and behavior. Cursor generates the patch. The observation agent validates it against the test suite. For incidents below a defined blast radius, the fix is deployed automatically with a human notification. For higher-risk changes, the patch is staged for approval. Mean time to recovery drops from hours into minutes. On-call engineering shifts from firefighting to oversight. Horizon 2-3 — MTTR in minutes not hours.
Honest Risks · Not Hype Suppression
What could go wrong — and why it matters more than the upside
Cursor's confidence problem scales. If the agent pair pattern is adopted without the audit log and review rules, Cursor's systematic overconfidence — inflated numbers, ungrounded scope, silent failures — scales to production systems. The pattern only works with the skeptic in the loop. Removing Claude from the pair to reduce cost is how you get a production incident at 3x the scale.
The knowledge layer becomes a single point of failure. In the knowledge-first architecture, if the enterprise knowledge graph is wrong, every agent that queries it is wrong — consistently and at machine speed. Garbage in, garbage at scale. The knowledge layer requires more governance, not less, as agent reliance on it grows.
Client governance lags capability. The agent pair can deliver faster than most enterprise procurement and change management processes can approve. The bottleneck shifts from delivery to governance. Organizations that don't invest in modernizing their approval processes will not capture the gains — not because the technology failed, but because the org model didn't adapt.
The human skill set erodes in the wrong places. If practitioners stop reading code because Cursor generates it, the capacity to audit Cursor degrades. If analysts stop writing specs because Claude drafts them, the capacity to evaluate Claude's specs degrades. The human role is not eliminated — it is elevated. But elevation requires investment in new skills, not the passive atrophy of old ones.
Model dependency risk. The audit log pattern is model-agnostic in theory. In practice, teams will optimize their review rules and plan docs for the specific behavior of today's Claude and Cursor. Model updates, pricing changes, or deprecations can invalidate hard-won institutional knowledge encoded in those documents. The knowledge layer must be designed to survive model churn.
Section 10 · Closing
CLOSING
The relay race is over. The loop begins.
The SDLC was never a methodology. It was an engineering solution to a cognitive bandwidth problem — how do humans coordinate the construction of systems too complex for any one person to hold in their head? The answer was sequence: fragment the problem, hand it off, reassemble. It worked. For forty years, it worked.
The agent pair does not fragment the problem. It holds the entire tree simultaneously. Claude maintains plan coherence across sessions. Cursor explores branches at machine speed. The human sets the goal, approves the plan, and steers at checkpoints. The loop runs continuously. There are no more relay handoffs — only convergence.
The practitioners who will define the next era of software delivery are building this infrastructure today — the audit logs, the review rules, the ToT plan docs, the knowledge layers. Not because the tools are mature. Because the pattern is proven. An embedding fine-tuning project, an AI operations architecture, a multi-region requirements reconciliation — these are not experiments. They are the new baseline.
The last relay has been run.
Build the loop.
Download
This article is available as a full PDF for offline reading and sharing.