Scaling AI Coding Agents From Prototype to Production Fleet
Architectural enforcement in CI, not prompts, is how teams keep agent fleets coherent.

Most organizations running AI coding agents in production have not reached fleet scale, and the reason has little to do with appetite for the technology. The blockers that keep those deployments from scaling into coordinated fleets are architectural, and that distinction sets the agenda for everything that follows.
The gap is not model quality. Leading agents can now research a repository, form a plan, edit files across that repository, run tests, and submit a pull request with minimal human direction, a capability GitHub documents in its Copilot agent tooling and that independent 2026 surveys of enterprise agent deployments confirm. Capability stopped being the ceiling some time ago. Multiple agents doing that simultaneously, against the same repository, is a different category of problem, and it is the one this piece addresses.
From one agent to many
When you run several agents concurrently against a shared repository, productivity does not scale in a straight line. What scales instead is architectural surface area, and it grows faster than any team can review it. The 2026 Agentic Coding Trends Report describes this directly: single agents are evolving into coordinated teams of agents working through orchestrator and sub-agent hierarchies, a structural shift already underway in production systems.
At fleet scale, agent sessions run in parallel, often on adjacent or overlapping parts of the same codebase, each one operating from its own context window with no visibility into what the others are changing at the same moment. That isolation is the core hazard. Agents do not share architectural memory across sessions. Each one reasons from a partial view of the repository, so when a decision needs coordination, like where a layer boundary sits or which module owns a given responsibility, nobody makes it deliberately. They get assumed, and different agents frequently assume them differently.
The output of that isolation is not code that reads as wrong line by line. A diff can look clean, pass its unit tests, and still violate the structural rules the system depends on: a controller that reaches past the service layer straight into the database, a bounded context that quietly starts importing from a neighboring domain. They pass CI as it's conventionally configured. They rot the architecture from the inside, and nothing about a syntax check or a green test suite will tell you that it happened.
Why agent-generated diffs are the wrong unit of review
The diff was never built to answer the question that fleet-scale agent work actually raises. It shows what changed in a set of files. It says nothing about whether the structure the team designed still holds after that change lands, and at fleet scale that gap turns the diff from an imperfect tool into an actively misleading one.
When a human developer opens a pull request, they usually explain the approach they took, the alternatives they considered and discarded, and the constraints that shaped the decision. Fleet scale then multiplies how many pull requests arrive per day, on top of a review burden that is already high for each one.
AI-based reviewers do not close that gap; research examining the shift from code review to code critique found that existing AI code review tools over-index on low-value suggestions such as style and best-practice nitpicks, while under-indexing on the concerns human reviewers actually prioritize: correctness, security, and performance. Adding an AI reviewer on top of agent-generated output often buries the structural violations that matter under a pile of cosmetic comments about naming and formatting.
Meta's internal RADAR system offers the clearest evidence of what automated review at scale actually requires. RADAR has reviewed more than 535,000 diffs and landed more than 331,000, reaching a peak throughput above 25,000 diffs per day. That volume could not be handled by a standard review queue; it required purpose-built infrastructure designed specifically for the problem. Even a system built at that scale optimizes for diff throughput, for getting changes reviewed and merged quickly, and verifying that the architecture underneath those diffs remains coherent is a separate problem it does not solve. Throughput and coherence are different problems, and solving one does not touch the other.
The obvious objection is that better prompts and more detailed agent instructions will close the gap. Instructions guide an agent on the task directly in front of it. They do not enforce structural invariants when multiple concurrent agents, each reasoning from its own isolated context window, work against a single shared codebase. The limitation is the absence of a shared architectural model that every agent in the fleet is required to satisfy, regardless of what its own instructions say.
What architectural drift looks like in a fleet-managed codebase
Architectural drift rarely arrives as one catastrophic merge. It accumulates as a long sequence of individually plausible changes, each one reasonable in isolation, that together erase the design a team started with.
Agent-assisted codebases show three failure patterns most consistently, with silent layer violations the third: code that skips the abstraction layer the architecture depends on, without tripping a test failure or a lint warning, because nothing in the toolchain is checking for that specific violation.
Context engineering at the file level, things like per-directory instruction files, read-deny rules, and hierarchical guidance passed down to an agent before it starts work, can steer an individual agent toward better local decisions. Each agent still reasons from its own window onto the repository, so however well you write an instruction file, it can't encode the full dependency graph of a real production codebase.
The organizational signal that drift has already set in is unmistakable once it appears: teams start spending more time untangling dependencies that agents introduced than they spend shipping new features. Each violation that lands in main becomes part of what the next agent reasons from, so the cost compounds. An architecture problem left unresolved for one more merge cycle is not static; it is the foundation the next round of agent-generated code gets built on top of.
Why enforcement belongs in CI, not agent instructions
Instructions can guide an agent's behavior. Only enforcement protects the architecture, and enforcement has exactly one place it can live and still mean anything: the CI pipeline, where a rule is unconditional, version-controlled alongside the code it governs, and cannot be talked around by a cleverly worded prompt.
The governing principle, drawn from how infrastructure teams have long handled policy and compliance, is that constraints belong upstream of generation. Catching a structural violation before it reaches the main branch is categorically cheaper than untangling it after three more agents have already built new work on top of it. Static analysis running in CI can encode architectural principles directly and mechanically: layered separation, so that controllers cannot reach past the service layer into the database; dependency inversion; bounded context boundaries drawn from domain-driven design. These are structural invariants that a test suite, by its nature, cannot verify, because a test checks behavior, not shape.
A CI architectural gate is only useful if it has one non-negotiable property: it exits non-zero and blocks the merge when it finds a violation. A gate that reports a problem but still lets the code through is a dashboard. It is not a gate, and calling it one invites the exact drift it was supposed to prevent.
The natural objection is that this slows the fleet down. It runs without slowing the agent. The agent commits its work and moves to the next task in its queue; the gate runs asynchronously against that commit and blocks the merge, not the generation. The fleet keeps producing work at its own pace. Only the code that fails to conform gets held back, and it gets held back before it can become the context the next agent inherits.
What those constraints must encode
An architectural gate only enforces what someone has already made explicit, and most teams beginning a fleet deployment have never written those constraints down in a form a machine could check. A gate with nothing encoded has nothing to enforce.
You need to make three categories of constraint explicit before enforcement means anything. Operational constraints cover pagination limits, query parameter requirements, and retry policies; a model reasoning about a single task cannot see these, but they decide how the system performs once real traffic hits it.
The right place for these constraints is the repository itself: version-controlled alongside the code, readable by the CI system that enforces them and by the engineers who maintain them. A diagram sitting in a wiki page or a document sitting in a shared drive cannot be checked by a machine and will drift out of sync with the code it was meant to describe the first time someone forgets to update it. Research formalizing this approach, published as arXiv:2602.02584 under the name Constitutional Spec-Driven Development, describes exactly this: organizations encoding architectural principles such as layered separation and bounded context boundaries from domain-driven design, specifically to prevent AI-generated code from violating structural invariants that testing alone struggles to detect.
A graph turns a rule like "no controller may call a repository directly" into a query that runs in seconds, replacing a convention written in a document that every engineer is simply trusted to remember and every agent has no way of seeing.
Building the architectural gate into the agent workflow end to end
So you need architectural enforcement built into every stage of the workflow, from how an agent receives its task to how its output gets verified before it merges. Bolting a check on at the end, after the pattern of drift is already visible, is a much harder problem than building the check in from the start.
At the planning stage, you should have agents operate against a shared architectural model, a graph of the repository's current structure, so task planning is informed by what already exists rather than by what an agent happens to remember from its own limited context. That shared view is what makes "find the existing helper function" the default behavior instead of "write a new one," so agents stop duplicating code that already exists. At the generation stage, per-directory instruction files, read-deny rules, and hierarchical context guidance reduce the odds of a local violation, but they remain probabilistic nudges. The gate is what makes the guarantee.
At the verification stage, you run a CI constraint check such as gr check on every agent-generated pull request, and it evaluates that output against the encoded architectural graph and exits non-zero the moment it finds a violation. The merge stays blocked until the violation is resolved, whether that resolution comes from the agent running a follow-up pass or from an engineer stepping in directly. Doctolib's integration of Claude Code shows this pattern already running in production: the tool's headless mode runs directly inside the CI pipeline, automatically opening pull requests for routine maintenance tasks, with every code change triggering a CI job that keeps technical documentation current. The infrastructure treats agent output as a first-class input to the verification pipeline, so it needs no separate handling as a special case.
At fleet scale, this gate does not slow anything down. Agents keep generating work in parallel, but only the pull requests that fail to conform get held at the merge gate, so everything that passes lands continuously. The developer's role in this system changes accordingly. Less time goes into reconstructing an agent's intent from a two-thousand-line diff, and more time goes into maintaining the architectural model that every agent in the fleet is measured against. The developer who has built this system is not reading through diffs line by line. That developer is navigating the graph to see what the fleet built and confirming that the structure held, which is a fundamentally different form of oversight than code review has ever asked for.
Steps teams should take before they scale to a fleet
If a team waits until after fleet scaling has already begun to impose architectural constraints, they will find the codebase has drifted past the point where the graph is clean enough to serve as a trustworthy source of truth. You need the constraints in place before the second agent starts running, not after the first few hundred pull requests have already landed.
The first step is to map the repository as it actually exists today, not as a design document claims it should be, but as the code on disk is genuinely structured. That map is the baseline every constraint will be measured against, and it is the baseline that will drift the moment it's left unguarded.
The second step is to make the architectural invariants explicit and machine-readable: layer boundaries, module ownership, operational constraints, written down before any agent touches the codebase at fleet scale. If a team can't state its own rules as checkable propositions, then no tool can enforce them, however capable it is.
So the third step is this: wire a constraint check such as gr check into CI as a required status check on every agent pull request, before you switch on parallel agent runs. The gate needs to be in place before the fleet starts running, because you can't add it once drift has become visible to everyone.
The fourth step is to treat the architectural graph as a living document: updated deliberately whenever the design changes on purpose, and checked continuously in CI to confirm that agent output matches the current intended graph.
Teams that have all of those controls in place and still lack architectural enforcement are running a fleet with no structural guardrails, and that is the far more common failure mode among organizations that believed scaling an agent fleet was primarily a governance problem.


