APIs, integration & security — in depth

Spec-Driven Development With AI Coding Agents

Specifications must execute as enforcement gates, not static documents agents can ignore.

Correspondent · · 10 min read
Cover illustration for “Spec-Driven Development With AI Coding Agents”
Agentic Coding Workflows · October 2, 2026 · 10 min read · 2,247 words

Spec-Driven Development works because it turns the specification into something that blocks a merge, not because it explains intent more clearly to an AI agent. That distinction, between a document an agent reads once and a contract an agent cannot get past, is what separates teams shipping reliable code from teams drowning in architectural drift they can't locate, let alone fix.

Why vibe coding breaks at production scale

A production team running multiple AI agents against a shared codebase hits a wall that a solo developer prototyping a side project never sees. Devoteam's 2026 analysis describes vibe coding in complex corporate environments as leaving "a trail of implicit decisions that no one remembers making and serious inconsistencies between components". The trouble isn't velocity. By 2026, generating a complete feature in minutes has become the baseline expectation, not a competitive edge, and the real damage comes from speed without structure, which produces drift that compounds every time another agent builds on top of the last one's work. Running several agents in parallel without explicit coordination piles up inconsistencies, redundant code, and architectural deviations fast enough that no one can audit them after the fact. The arXiv process taxonomy published by Macedo in June 2026 names the mechanism directly: without specification, traceability, defined roles, and validation, an autonomous agent session just reproduces, at a much larger scale, the same problems that isolated prompting always had: lost context, decisions nobody wrote down, and a review process that can't keep up. The goal was always structural governance of what agents are allowed to produce.

SDD reverses code into a derived artifact

SDD doesn't layer specifications on top of an unchanged development process. It inverts which artifact holds authority. Devoteam's analysis states the principle without hedging: "code becomes a strictly derived artifact… the specification becomes the true source of truth; if the required behaviour changes, the specification is modified, and automated systems take care of regenerating, adapting, and auditing the resulting code". Once that inversion holds, the engineer's job changes shape. Instead of translating a set of requirements into code by hand, the work becomes defining what the system must do, specifying the conditions under which a given behavior counts as correct, and building the automated means to check that condition holds.

Picture the flow as a straight line: business requirements turn into an executable SDD specification, and that specification drives the AI agents that produce derived code. The specification sits at the center of that flow as the interface between human intent and agent execution. That's a different claim than "write better documentation." Final authority over what ships determines everything downstream about how enforcement can even be built.

The two structural failure modes that make enforcement non-optional

Spec-driven workflows built in good faith still fail, and they fail in two specific, recurring ways that only enforcement, not better intentions, can close. The Spec Growth Engine, published by Hartwig Grabowski at Hochschule Offenburg in June 2026, names both. The first is context explosion: an agent reasoning over an entire repository at once fills its context window and output quality degrades, a state the paper calls the "Dumb Zone," in which the agent starts conflating concerns from unrelated modules and produces fixes that reach across boundaries nobody asked it to touch. The second is silent spec-code drift: code keeps evolving while the specification stays frozen, the gap between them never gets flagged, tests keep passing, and the system ships with that gap built in. Each undisciplined round of changes adds to the divergence until the specification stops functioning as a living contract and becomes, in the paper's words, a historical artifact, with future agent runs guided by a document that no longer describes the system it claims to describe.

The drift failure is the more dangerous of the two at agent speed. An agent producing hundreds of lines a minute against a stale spec accumulates damage far faster than a human team ever could: a linter won't catch it, CI won't catch it, and the code ships anyway. The concrete patterns are specific and recognizable. An agent adds a repository implementation that imports directly from the domain layer into an infrastructure adapter, violating hexagonal architecture boundaries it was never told about. Another flattens a module hierarchy because its context window never had visibility into the project's structural conventions. Circular dependencies appear because agents generate cross-module utility functions for each other without any shared map of who owns what. Both failure modes trace back to the same root: nothing sets a principled boundary on what an agent is allowed to know and do, and without that boundary, a developer is left with two bad options, hand the agent everything (expensive and noisy) or guess at what's relevant (fragile and inconsistent). That's the structural case for treating SDD as more than a documentation habit. It has to function as a control system with hard boundaries built in.

Limits of the traditional spec-as-document approach

A spec that governs agent behavior and a spec that doesn't differ not in how well they are written but in whether the spec runs as a validation gate. A January 2026 preprint, Spec-Driven Development: From Code to Contract in the Age of AI Coding Assistants, draws the formal line: a traditional spec is read by a human, while an SDD spec executes, whether as a BDD scenario, an API contract test, or a model simulation. That changes what the artifact can do.

The Macedo process taxonomy, published on arXiv, surveyed frameworks across the field and found the same risk recurring in nearly all of them: "drift between specification and code, excessive trust in generated artifacts". These are tools producing documents an agent reads once, rather than living assets an agent executes against on every run. A separate 2026 tool evaluation found the failure showing up in practice rather than theory: static-spec tools "produce documents that drift from implementation within hours," and the evaluation treats the spec's lifecycle, whether it stays current or goes stale the moment code changes, as the first question any team has to answer before adopting a given tool. A specification that sits outside the enforcement layer gives an agent nothing to push against. An agent that can ignore a spec will ignore it because it has no mechanism for detecting what it just violated.

Much of what experienced developers used to assume silently, the architectural conventions nobody wrote down because everyone on the team just knew them, now has to be made explicit enough to implement and test directly. The spec has to answer "under what conditions is this behavior correct?" and not stop at "what should the system do?" Devoteam's analysis states the shift without qualification: the specification "has ceased to be a boring static PDF document or a forgotten page on Confluence"; it functions as "an executable operational contract that does not merely describe the system passively" but "actively governs it".

The review bottleneck shows what happens when enforcement is missing upstream

When nothing enforces architectural constraints before a merge, every bit of that cost lands on code review, and review was never built to absorb it at the speed agents now produce code. Faros AI's 2026 telemetry, drawn from more than 20,000 developers, found median code review time climbing sharply even as task throughput rose at the same time: AI-assisted pull requests run considerably larger than human-authored ones and sit waiting for a reviewer far longer. The same dataset shows pull request size growing in step with review time and with bug rates per developer. More code is shipping alongside more defects that still need to be caught somewhere downstream. A separate lab-versus-reality study from Faros AI found that AI adoption lets developers juggle more concurrent workstreams at once, more tasks and more pull requests moving through the pipeline each day, and that compounds the bottleneck directly: the volume of code needing verification grows faster than the humans available to verify it.

Agent failure modes aren't random. They cluster around CI configuration, duplicated logic, and edge-case handling, not around fresh, well-tested algorithmic code. A spec gate enforcing contracts can catch most of these before a reviewer ever opens the diff. The direction this points is clear: the next real advantage for AI-assisted teams won't come from bigger prompts or faster agents. It will come from review systems that force reproduction, small diffs, tests, and receipts before any code reaches a human reviewer at all. That's the case for enforcement sitting upstream of review rather than inside it. By the time a human opens a sprawling diff spread across dozens of files, the architectural violation already happened, and review at that point is cleanup.

What executable enforcement looks like in practice

Executable enforcement means spec-code divergence becomes a condition that blocks a merge outright, a technical gate rather than a discipline problem left to a team's conscience. The Spec Growth Engine proposes the hardest version of this invariant: spec and code may never diverge silently, full stop, so the drift gate turns any divergence into a blocking merge error, and the design couples spec and code into the same commit so every code change carries its spec update along with it. In CI practice, the enforcement toolkit looks like static analysis run before merge, tools like ArchUnit or ESLint alongside custom scripts, catching architectural violations before the code in question ever reaches the main branch; code that breaks a constraint simply doesn't merge.

Standard CI/CD pipelines were never built to catch what agents introduce. Traditional CI catches syntax errors and type mismatches, but it misses behavioral contract violations, infrastructure RBAC policy mismatches, SLO degradation, and dependencies an agent hallucinated outright. Closing that gap takes spec validation running as a first-class CI stage, plus a verifier gate that blocks a merge the moment agent output drifts from the plan it was given. Researchers at the Universidad Politécnica de Madrid describe this enforcement harness as operating in two layers at once: a technical harness around the agent itself (the spec, context scoping, validation gates) and a methodological harness around the team (roles, review checkpoints, human-in-the-loop patterns). Both layers are necessary, and neither one does the job alone.

Context scoping belongs inside that control system as a structural requirement, not an optional performance tweak. The Spec Growth Engine's Spine assembler restricts what an agent can see to the ownership path relevant to its actual task: an agent working on a payment module gets only the path leading to payment plus the contracts of its declared dependencies, not the rest of the repository. That turns context explosion into a solvable engineering problem rather than something left to prompt-engineering guesswork. GitHub Spec Kit's command structure shows what this looks like when it's running day to day: /speckit.specify captures business context and success criteria, /speckit.plan turns specs into architectural decisions, /speckit.tasks breaks plans into testable units, /speckit.implement runs agents inside those constraints, /speckit.constitution encodes the project principles every generation has to respect, and /speckit.clarify resolves ambiguity before planning even starts, among a broader command set that also includes /speckit.converge, /speckit.analyze, /speckit.checklist, and /speckit.taskstoissues. None of that is a document an agent reads once. It's a runtime the agent operates inside, continuously. An architecture graph mapping every module, file, class, function, and call relationship, kept current as agents write code, gives engineers the structural visibility to steer that process deliberately and run constraint checks as part of CI; a pattern like gr init to map the repository and gr check to enforce constraints at merge time applies the same graph-as-source-of-truth idea directly to SDD enforcement.

The enforcement spectrum across six current frameworks

The six frameworks the 2026 research evaluates sit at different points on a spectrum running from document-as-spec at one end to enforcement-as-merge-gate at the other, and the gap between them is not marginal.

One platform in that evaluation, built around living, auto-updating specs at organization-wide scope, uses a Context Engine that reasons across a large codebase and an Organization Knowledge layer that compounds across sessions. When an agent changed an API response shape midway through the four-microservice refactor, the living spec reflected that change immediately, so every downstream agent referenced the updated contract rather than a stale one. It's priced on a flat-rate business plan with no per-seat charge, entered public preview on May 3, 2026, and reached general availability on June 3, 2026. It fits multi-service teams that need their specs to stay current across sessions best, and a solo developer working a single repository will likely find its organization-wide scope more than the job calls for.

Amazon's Kiro takes a different approach: static specs written in EARS notation ("WHEN [condition] THE SYSTEM SHALL [behavior]"), organized across a three-document system of requirements.md, design.md, and tasks.md. A 2026 Requirements Analysis feature applies formal logic and SMT solvers to catch contradictions before code generation even starts. Kiro launched as a new agentic IDE on July 14, 2025, and Amazon's earlier Q Developer was deprecated in 2026 with Kiro named as its official successor. It offers a free tier with a monthly credit allowance that doesn't roll over, paid tiers above that at increasing price points, and a custom enterprise option beyond those. Its specs stay static and don't update as implementation evolves, which makes it the strongest fit for AWS-native greenfield projects rather than long-lived systems under constant architectural change.

GitHub Spec Kit is a third point on the spectrum: static markdown specs, agent-agnostic by design, and fully open source.

Sources

  1. From Prompt to Process: a Process Taxonomy and Comparative Assessment of Frameworks Supporting AI Software Development Agents
  2. Spec-Driven Development in 2026: The end of code as the center of development?
  3. The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development

More in Agentic Coding Workflows