APIs, integration & security — in depth

AGENTS.md and Agent Context Files for Large Repos

Editorial team · · 10 min read
Cover illustration for “AGENTS.md and Agent Context Files for Large Repos”
Agentic Coding Workflows · October 5, 2026 · 10 min read · 2,208 words

AGENTS.md exists because coding agents had no shared place to look for project context, and every tool before it solved that problem differently, or not. Before a cross-vendor standard took hold, one agent read a file in one format, another ignored project files entirely and relied on whatever the user typed into a prompt, and a team running three different agents on the same repository had to maintain three different sets of instructions, or none. The fix was deliberately plain: a single Markdown file at the root of the repository, no required schema, nothing to install, so any compatible agent could read it and act on it immediately.

Governance of the format has since moved to the Agentic AI Foundation under the Linux Foundation, which keeps the file name stable and keeps the specification itself vendor-neutral by design. That neutrality has a practical effect: the spec stays thin, with no required frontmatter, even as individual tools keep adding their own proprietary options on top. AGENTS.md remains the portable layer that every tool-specific file builds on. A canonical-source pattern has taken hold among teams running multiple agents: write the instructions once in AGENTS.md, then in a tool-specific file like CLAUDE.md, simply tell the agent to import it. One file to maintain, every tool benefits.

That pattern matters because support isn't universal by default. As of August 2026, Claude Code still loaded CLAUDE.md rather than AGENTS.md natively. That changed with a later version update, which added AGENTS.md as a fallback when no CLAUDE.md file is present. A team that hasn't updated past that version, or set up the import manually, cannot assume out-of-the-box support for AGENTS.md across every tool in a stack.

What the research shows developers actually put in these files

Empirical analysis of real repositories shows a consistent pattern in what actually lands inside these files: testing instructions come first, followed by implementation details, then architecture, then development process and contribution guidelines, then build and run commands. Architecture does show up, but it comes third, behind two categories that are operational at their core. Most files lead with how to run the build and how to run the tests, not with how the system is shaped or why it's shaped that way.

The format's own recommended structure reinforces this bias rather than correcting it. The suggested section order, as the spec lays it out, runs through project overview, build and test commands, code style, testing instructions, security notes, and commit or PR rules. That's a sensible skeleton for getting an agent oriented quickly, but it says nothing about how to encode design boundaries, module relationships, or the constraints an architecture depends on to stay coherent. Non-functional concerns, performance budgets, security postures, architectural limits, appear rarely in existing files, a gap documented in empirical studies of how these files are actually structured.

None of this reflects carelessness on the part of the developers writing these files. People default to operational plumbing because the template emphasizes it, and because operational instructions are easier to write down than architectural ones. Agents get precise instructions for how to run a test suite and sparse guidance on what the architecture is and what it needs to remain, and that architectural information matters most once a codebase grows past the size a single person can hold in their head.

How the Operational-Instruction Model Breaks Down as a Codebase Grows

You can describe a small codebase in one prompt, and that's what a flat context file built around build commands and testing steps was designed for. Vasilopoulos (2026) makes this point directly: a small prototype can be fully described in one prompt, but a large production system cannot be, and single-file manifests simply don't scale past modest codebases. The content that works for a ten-thousand-line side project stops matching the problem once the repository reaches ten times that size with multiple teams touching different modules on different schedules.

The scale of the underlying shift makes the mismatch sharper. Agents can now handle entire implementation workflows on their own, and agent-generated pull requests have grown from a negligible share of GitHub activity to a substantial portion of it within a little more than a year. The volume of agent-authored code has outpaced the context infrastructure meant to guide it, and that creates a specific failure mode: architectural drift. When an agent can touch dozens of files in a single run and multiple agents are submitting pull requests concurrently, each individual change looks small enough to pass review without concern. The structural damage accumulates across many small, individually reasonable-looking changes, which is harder to catch than one large change that obviously breaks a pattern.

Review tooling hasn't closed that gap. AI coding agents have increased output substantially, but most review tools don't address the widening quality problem that comes with it: larger pull requests, architectural drift, and inconsistent standards across repositories that span multiple teams. LinearB's 2026 Software Engineering Benchmarks Report found that agentic AI pull requests have a pickup time 5.3 times longer than unassisted pull requests, a gap that points to a structural bottleneck in review. Reviewers are slower to start reviewing agent-authored work, plausibly because there's more of it, it's less predictable in shape, and it carries less signal about why a given change was made.

A flat AGENTS.md filled with build commands and style rules doesn't touch any of this. It gives an agent operational rails, how to compile, how to run the linter, how to format a commit, but no map of the architecture it's operating inside. An agent following those rails moves fast, but it still drifts into decisions that conflict with the system's actual design. At scale, the content inside the file starts to matter less than its structure: what gets loaded, when, for which part of the codebase, and whether what's loaded actually encodes the constraints the architecture depends on.

What the ETH Zurich findings actually show, and what they don't

The strongest empirical challenge to the idea that AGENTS.md files help agents comes from a study run by Gloaguen, Mündler, Müller, Raychev, and Vechev (arXiv:2602.11988). The study doesn't claim that context files are useless. It makes a sharper claim: poorly structured context files perform worse than no context file at all, which is a more specific and more useful finding than a blanket verdict on the format.

The researchers ran four agent-and-model pairs across a wide set of tasks drawn from a range of repositories, and they chose niche repositories over popular ones, since popular repositories tend to be unrepresentative of the average codebase and often lack developer-committed context files to begin with. Three conditions were tested: no context file, an LLM-generated context file, and a human-written one. LLM-generated files cut task success rates compared with having no context file at all, and they also drove up inference costs. Human-written files produced a marginal gain in success rate, but that gain came with a parallel rise in cost, so even the better condition wasn't free.

The trace analysis behind these numbers points to a specific mechanism. Agents followed the instructions they were given faithfully: they ran more tests, read more files, executed more searches. Much of that extra activity wasn't necessary for the specific task in front of the agent, and it forced reasoning models to spend more effort without producing better patches. Broad instructions produced broad, unfocused compliance.

One structural finding in the study deserves particular attention: including an architectural overview or a repository structure explanation did not help agents find the relevant files any faster. Agents given no context and agents given a file tree converged on the same files at roughly the same speed. Agents are already good at discovering directory structure on their own, so a static file tree in AGENTS.md, a common template practice, can actually slow down the search for the relevant source file rather than speed it up.

What the study cannot measure matters as much as what it found. The study tests single-task completion on isolated issues, so it cannot capture what AGENTS.md is worth across multi-session, multi-agent workflows, where consistency between sessions affects outcomes across every task in the sequence, not just one. The authors' own prescription follows directly from the mechanism they identified: skip LLM-generated context files entirely, and limit human-written instructions to details the model genuinely cannot infer from the codebase itself, specific tooling quirks, custom build commands, domain knowledge that isn't visible in the code. That prescription doesn't undercut the case for context files. It sharpens what a good one actually contains.

The Efficiency Evidence for Focused Context Files

A separate study answers a narrower and more practical question: when a context file is focused rather than comprehensive, does it help, and by how much? Lulla et al. (arXiv:2601.20404v2, March 2026) studied 10 repositories and 124 pull requests, running agents under two conditions, with an AGENTS.md file and without one. The presence of the file was associated with a reduction in median runtime and a reduction in output token consumption, while task completion behavior stayed comparable between the two conditions.

The gains here sit on the operational side of agent behavior. Agents waste less time exploring a codebase and recovering from wrong assumptions when the file in front of them encodes project-specific knowledge the agent could not otherwise infer. That's a different claim from the one the ETH Zurich study tested, and the two results reconcile once the difference in what was actually studied becomes clear. Lulla et al. looked at focused files operating inside real pull request workflows. The ETH Zurich team looked at a mix of file qualities, including LLM-generated ones, and applied them to isolated tasks. Read together, the two studies are compatible: quality and targeting determine the outcome.

Both studies point toward the same operating principle: serve only the instruction subset relevant to the current type of task. A file loaded for writing tests doesn't need the same content as a file loaded for restructuring a module, and if you narrow what gets loaded, you cut context size while holding accuracy steady or improving it. That gives content strategy a concrete test to apply line by line: can the agent reliably infer this from the codebase itself? If it can, the line doesn't belong in the file. If it can't, that's exactly the kind of detail AGENTS.md should carry, custom tooling behavior, a build step that isn't obvious from the repository's structure, a domain rule with no trace in the code.

The hierarchy that makes context files work in large and multi-module repos

Once the content strategy is narrowed to what's genuinely non-inferable, the next question is how to organize it across a codebase too large for one file to cover well. In a large repository, the unit of context isn't the file, it's the module, and the structure of the context system should mirror the structure of the codebase itself. A module-level file carries what's specific to that boundary: which other modules are allowed to import from it, which patterns are canonical inside it, what the local test strategy looks like. A root-level file stays small and stable, and it holds only the conventions that apply everywhere.

Vasilopoulos (2026) built a three-component system on a 108,000-line C# distributed system that shows what this looks like at the far end of the scale. The first layer is a hot-memory constitution: conventions, retrieval hooks, and orchestration protocols that stay loaded at all times, functioning like a lean root AGENTS.md with pointers out to everything else. The second layer is a set of 19 specialized domain-expert agents, each carrying project-specific knowledge for one bounded area; this is what deeply informed module-level files look like when built as agent personas rather than static text. The third layer is a cold-memory knowledge base of 34 on-demand specification documents, loaded only when a task actually calls for them, putting the tiered-injection principle from the efficiency research directly into practice. Across 283 development sessions, the infrastructure grew to roughly 26,000 lines, more than ten times the size of a typical manifest, which marks the upper bound of what this pattern can demand rather than a baseline to aim for.

A comparable pattern has emerged independently elsewhere. Google's Conductor approach for Gemini CLI uses a set of persistent Markdown files and enforces a structured lifecycle, moving through context, then spec and plan, then implementation, which addresses the same underlying problem through its own route and arrives at a similar tiered design.

What unifies both cases is a single design principle: keep the always-loaded conventions small, stable, and architectural, and keep the on-demand specifications separate, larger, and task-specific. Loading everything on every task, the static file tree, the exhaustive build instructions, the full history of style decisions, is the exact failure mode the ETH Zurich study diagnosed when it found that broad context led agents to do more work without producing better results. A context system built in tiers avoids that failure by design: the agent carries only what it needs for the task in front of it, and the architecture of the codebase stays visible at the level where it actually governs behavior, the module, rather than buried in a single file meant to describe everything at once.

Sources

  1. Codified Context: Infrastructure for AI Agents in a Complex Codebase
  2. On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents
  3. 2026 Agentic Coding Trends Report How coding agents are reshaping

More in Agentic Coding Workflows