Monorepo Architecture Constraints Under AI Agent Workflows
Monorepos amplify both agent capabilities and architectural risks simultaneously.

Monorepos solve three specific structural failures that make AI agents unreliable in polyrepo environments, and that alone makes them the better substrate for agentic development work. Call them the read wall, the write wall, and the memory wall. Each one describes a different way an agent loses the information it needs to make a correct change, and each one disappears, or at least weakens considerably, when the codebase lives in a single repository.
The read wall is the most immediate. In a polyrepo, an agent operates on one project at a time, with no visibility into how a change ripples outward. Nx's analysis of this problem is direct: an agent asked to modify a shared UI component library has no way of knowing which downstream applications consume that library or how they'll break when it changes. The agent isn't being careless. A monorepo removes that blindness by putting the whole dependency graph in view, so the agent can trace a change from its origin to every place it lands.
The write wall follows naturally from the read wall. Every handoff is a point where the agent's autonomy ends and a person's calendar begins. A monorepo collapses that coordination problem into a single commit. The agent can complete work that would otherwise require a chain of synchronized human approvals.
The memory wall is the subtlest of the three and the one Nx frames as the "Memento Problem," named for the character who loses his memory every few minutes. Every repository boundary wipes an agent's working context clean. In a monorepo, the agent works inside one continuous context from start to finish. Mistakes still happen, but they happen inside a single coherent frame of reference, which makes them easier to trace and easier to catch than errors scattered across disconnected, context-free repos.
Nx's own illustration of this is concrete. Ask an agent to add a field to an API response inside a monorepo: it finds the schema, the corresponding TypeScript type, and the component that renders the data, all in the same pass, because all three live inside one visible graph. Asking it to do the same thing across a multi-repo setup makes it frequently hallucinate the frontend type, since it has no way to see what the backend actually returns. The tool caught it specifically because the backend change and the UI change touched the same pull request, in the same repository, where both halves of the mistake were visible together. In separate repos, that error had a real chance of slipping through unnoticed on either side.
A second-order effect follows from this, too. Monorepos remove enough of that friction that work previously written off as too costly becomes achievable, which is a genuine, structural gain for teams adopting agentic development.
What the same openness makes possible, and dangerous
The properties that let an agent trace a dependency graph across an entire codebase are the same properties that let an agent's mistakes travel across that same codebase. Full visibility into a project graph is what allows an agent to understand that changing a shared library will affect a dozen downstream applications. An agent that misjudges scope can touch all twelve of them in a single commit for the same reason. The blast radius of an agent's error in a monorepo scales with exactly the same mechanism that makes its correct changes so powerful.
Context overload compounds the problem in a related but distinct way. A monorepo doesn't scope an agent's context automatically; it just makes everything available, and availability is not the same as relevance.
Volume adds a third layer to the same dynamic. AI-assisted pull requests tend to run substantially larger than pull requests written without AI assistance, and agentic systems produce that volume at a pace no human-only team could sustain. Airbnb's experience with an LLM-driven automation pipeline shows the upside of that speed concretely. The result is genuinely impressive, and it carries an implication that matters just as much as the achievement itself: governance has to scale at the same rate as velocity, or the organization is simply accumulating risk faster than it can inspect it.
That shift changes what the job of a developer actually is. Writing code stops being the bottleneck. Reviewing and verifying code someone, or something, else wrote becomes the bottleneck instead. Agent-generated diffs break that assumption, because they are often large, heterogeneous, and delivered with the same uniform confidence whether the change is trivial or architecturally significant.
Architectural drift as the failure mode monorepos surface
When agents make broad, cross-cutting changes with full write access and no enforced boundaries, the design a monorepo was built to unify can disappear without a single bad commit ever being identifiable as the cause. It erodes instead through an accumulation of small, individually plausible violations, each one buried inside a diff too large for any one reviewer to fully trace. That erosion is architectural drift, and it is the specific failure mode that monorepos make visible at scale.
Drift is silent by nature. Nothing in a passing CI run tells a reviewer that a UI layer just started talking directly to a database layer it was never supposed to touch.
The monorepo structure makes this worse precisely because it makes so much more possible. The downside is that structural violations now travel at exactly the same scale as legitimate improvements, carried by the same mechanism, with no built-in way to tell the two apart.
The persistence of agent instructions compounds the risk further. An empirical study of GitHub Agentic Workflows found that among files with at least 120 days of observed activity, a large majority still receive updates in month four, which confirms that these are not one-off prompts but ongoing specifications executed repeatedly over time. Any architectural misconception embedded in one of those instructions doesn't produce a single bad outcome. It compounds, run after run, for as long as the instruction remains in place and unexamined.
Context engineering research gives this dynamic a useful frame: the entire informational environment in which an agent makes decisions functions as the lever controlling its behavior. Poorly specified or unenforced architectural context doesn't cost a team one flawed pull request. It shapes every subsequent run that draws on that same context. An unexamined architectural gap compounds just like an unexamined instruction does.
An objection surfaces here: teams already use code review to catch exactly this kind of structural violation. The objection misunderstands what code review was built for. The constraint has to be enforced before the pull request ever reaches a reviewer's queue. Reviewability isn't the problem; no mechanism enforces the design before the code shows up for review at all.
AGENTS.md files and architectural intent at the package level
AGENTS.md is a plain Markdown file that tells an AI coding agent how to build, test, and change a project, and it has become the primary surface for encoding architectural constraints that linters and formatters were never built to express.
In a monorepo, a single global AGENTS.md file at the root of the repository is the wrong design. A rule about how the billing service handles currency rounding has no business sitting in the context window of an agent editing a frontend component, and a hierarchical AGENTS.md structure keeps it out.
OpenAI's main repository is the clearest real-world demonstration of this pattern at scale: 88 separate AGENTS.md files spread across its directory tree, each one scoped to the part of the codebase it governs. That number shows what a mature, agent-first monorepo actually looks like in practice: not a single document trying to hold an entire architecture's worth of rules, but dozens of focused, package-level specifications that stay relevant to whatever task an agent happens to be working on.
The empirical study of GitHub Agentic Workflow files backs up the claim that practitioners are already taking this seriously. The same study also found that only a small fraction of files explicitly address prompt-injection defense, which suggests that even the more mature constraint files in production today have real blind spots on the adversarial side of the problem.
What AGENTS.md files give a team, in short, is a working, scoped, machine-readable statement of architectural intent. What they don't give a team is any guarantee that an agent will actually follow it.
Why declarative constraints alone do not enforce the architecture
An agent reading an AGENTS.md file is not the same as an agent obeying it. Declarative constraint files are necessary, but they stop short of enforcement, because nothing about a Markdown instruction forces an agent's output to comply with what the file says.
Programmatic enforcement is missing: dependency rules applied through automated linting. A team can specify, for instance, that a UI layer must never query a database layer directly, encode that rule as a verifiable dependency constraint, and let automated linting catch any violation the moment an agent attempts it, flagging the change before it goes any further. That is a fundamentally different kind of safeguard than a sentence in a Markdown file asking an agent to respect a layer boundary, because it does not depend on the agent choosing to comply.
Context engineering research states the underlying principle with precision: whoever controls the agent's context controls its behavior, whoever controls its intent controls its strategy, and whoever controls its specifications controls its scale. A constraint that exists only inside a Markdown file controls context. It does not control execution. The agent can read the rule, understand the rule, and still produce code that violates the rule, because nothing enforces the boundary at the moment the code is written.
That gap separates two distinct problems that are easy to conflate: specifying a constraint and enforcing it. A team that writes a thorough AGENTS.md file, scoped correctly across every package, has done real and valuable work. It has documented its architecture clearly enough for a human or an agent to understand the design's intent. It has not yet protected that design from violation, and those are not the same accomplishment.
CI as the non-negotiable enforcement gate for architectural constraints
Continuous integration is the one point in an agentic workflow where an architectural constraint can be made truly unconditional. A check that exits non-zero and blocks a merge cannot be reasoned with, deprioritized under time pressure, or quietly dropped from an agent's context window the way a Markdown instruction can be. CI either passes or it doesn't, and that binary quality is what a constraint needs to function as a constraint rather than a suggestion.
The stakes of getting this wrong rise sharply as agent autonomy increases. The critical risk at that level of autonomy concerns deployment safety specifically. If an agent can trigger a deployment on its own, the rate at which changes fail in production can rise unless constitution-level constraints are enforced directly in the pipeline, rules like never deploying without passing integration tests, or requiring human approval before anything reaches production.
The principle of least privilege applies here with the same force it applies anywhere else in a system that grants autonomous action. Agents deploying through CI need strict role-based access control, scoped to the minimum permissions required for the task in front of them. This governs not just what an agent is permitted to read across the codebase, but what it is permitted to merge, which is a different and in some ways higher-stakes boundary.
The infrastructure around this problem is maturing quickly. Sonar's acquisition of Gitar in May 2026 is a useful marker of that shift: Gitar monitors CI pipelines, performs root cause analysis on build failures, and iterates automatically until the build passes, and Sonar has continued selling it as a standalone product alongside SonarQube. That's a signal that agentic CI integration has moved past the startup-experiment stage and into enterprise infrastructure that organizations are willing to build their pipelines around.
What CI enforces, in the end, is the gap between documenting an architecture and defending it. An AGENTS.md file documents the design. A CI gate defends it, because it is the one place in the pipeline where a violation gets stopped rather than merely described.
Reviewing agent-generated diffs without losing architectural signal
Even with CI enforcing hard constraints, some agent-generated changes still reach a human reviewer, and the review process itself has to change shape to handle them. The review has to be restructured around hypotheses: a reviewer should approach a large diff by asking what specifically could have gone wrong given the scope of the change, then drilling selectively into the files most likely to carry that risk, rather than attempting to verify every line with equal attention.
This matters because equal attention across an agent-generated diff is itself a kind of failure. The review has to be weighted toward the places where architectural rules are most likely to be tested, the layer boundaries, the dependency edges, the parts of the system an AGENTS.md file and a CI rule were written specifically to protect.
Review process design is, in that sense, itself an architectural constraint, not a separate workflow concern layered on top of one. A team that writes precise AGENTS.md files, enforces dependency rules through CI, and still reviews every agent diff with the undifferentiated attention built for human-scale pull requests has only solved part of the problem. The architecture survives when every layer of the workflow, the context file, the linter, the pipeline gate, and the reviewer's own judgment, is built around the same understanding of what the system is supposed to be, and what it cannot be allowed to become.


