APIs, integration & security — in depth

Context Window Limits and Codebase Chunking Strategies

Larger context windows won't fix codebase interactions without proper chunking strategy.

Staff Writer · · 11 min read
Cover illustration for “Context Window Limits and Codebase Chunking Strategies”
Agentic Coding Workflows · October 10, 2026 · 11 min read · 2,432 words

A bigger context window will not fix a broken codebase interaction, because the failure was never about how many tokens the model can hold. It is about how that content is organized once it gets there. Three mechanisms compound to produce this failure, and none of them improve just because the window grows. First, attention diffuses: an agent handed the full repository still hallucinates cross-file dependencies, because sheer volume does not sharpen focus on the files that actually matter to the task at hand. Second, output caps bind before input caps do. A model that can read hundreds of thousands of tokens still writes only a fraction of that per turn, so a large refactor demands multiple round trips, and each round trip re-sends a growing history that eats into the budget meant for new work. Third, naive truncation destroys coherence: when context gets cut arbitrarily to make room, agents lose track of what they already changed and repeat steps they already completed. Supermemory's analysis of agent failures separates these modes with some precision: an agent may retrieve an old revision, miss a caller, confuse two services, or treat an abandoned proposal as a live decision. Context-window size is one possible cause among those, and often not the one actually responsible for the mistake.

How arbitrary splitting breaks structural information

When context limits hit, people usually just split code into chunks by character count or line count, which makes the underlying problem worse. A chunk boundary drawn without regard for what the code actually contains tends to land in the worst possible places. A function sliced in half leaves an agent staring at a body with no signature, or a signature with no body attached to it. A class definition separated from its methods strips away the ability to reason about how the object actually behaves. An import stripped from the symbol it enables destroys the provenance of every call that depends on it, so the agent loses the thread connecting a piece of code to where it actually comes from. Parsing to an abstract syntax tree solves part of this: an AST can identify functions, classes, and other structural boundaries, and chunking around those boundaries makes a retrieved passage far easier to interpret on its own terms. But structural boundaries inside a single file do not add up to a complete cross-service dependency graph, particularly once dynamic calls or generated code enter the picture. That gap matters because knowing a function exists is a different kind of knowledge than knowing who calls it, what it imports, and which services depend on its contract. Answering a question about callers, imports, or contracts requires symbol search, a language server, or static analysis, not just a cleaner split boundary. Kinde's engineering guide states the practical consequence directly: without full structural context, an AI may suggest changes that break other parts of the application, introduce bugs, misread business logic, or apply outdated patterns, all because the chunk it received was internally consistent but architecturally incomplete.

Structure-aware chunking: AST, call graphs, and dependency boundaries

Structure-aware chunking treats the codebase's own structural units, functions, classes, modules, call relationships, import graphs, as the natural boundaries for a chunk, so that every piece handed to an agent is self-contained and relationally coherent. This approach has layers, and each one does work the others can't. The AST layer parses to structural boundaries instead of character counts: functions stay whole, class hierarchies stay intact, and imports travel together with the symbols they enable. The call-graph layer adds reachability. The Context Compiler approach traces explicit imports first, then falls back to a repo-wide symbol table, expanding breadth-first out to a hop limit, so that a file is included only if it can actually be reached from the edit target, not simply because it happens to sit nearby in the directory tree. VITAL-RAG frames the same problem as an invariance race: allocation of context budget should stay stable under redundant renderings of the same code object but shift when a fragment adds information the task actually needs, grouping fragments by canonical code object and rendering them under per-object and global token budgets. The interface-extraction layer compresses dependencies down to their contracts. Pass 2 of the Context Compiler strips files that are reachable but not being edited down to function signatures and docstrings, replacing every function body with a single placeholder, which shrinks file size dramatically while preserving everything the editing agent needs to know about that file's contract. This is not lossy compression in the sense of degraded information. It follows the same discipline a compiler applies at any intermediate stage: keep what the next stage needs, drop what it does not.

Why retrieval must be just-in-time rather than bulk-loaded

Chunking the codebase correctly solves only half the problem, because feeding an agent every relevant chunk upfront is structurally equivalent to feeding it the whole repository. It loads irrelevant material into the window, inflates the cost of every turn, and triggers the same attention-diffusion and distractor-interference effects that structural chunking was supposed to eliminate. The alternative is retrieval scoped to the moment: an agent requests the chunk it needs at the point in the task where it needs it, not before. Retrieval-augmented generation formalizes this pattern. A query converts to an embedding, a vector database returns the chunks most relevant to it, and those chunks augment the prompt without the agent ever holding the full index in its context window. A hybrid pattern gaining adoption pairs BM25 lexical lookup with dense vector embeddings and fuses the two scores, often through Reciprocal Rank Fusion, catching both keyword-level matches and semantic-level relevance that a single method would miss on its own. Tool output deserves particular attention here, because it is the silent killer of context budgets that structural chunking alone does nothing to prevent. A browser snapshot from a tool like Playwright, a dumped GitHub issue thread, raw shell output: each of these accumulates in context the moment an agent uses it, and then sits there consuming window space the agent will never reference again. Just-in-time retrieval has to govern this kind of output as deliberately as it governs code chunks, evicting what has already served its purpose. The freshness of the underlying index matters just as much as the retrieval method itself. Each indexed chunk should carry its repository, its branch or commit, its path, and its source revision as metadata, so that staleness becomes detectable at the moment of retrieval rather than discovered after an agent has already acted on outdated code.

The hierarchical context model: what stays loaded, what loads per task, what gets dropped

Diagram: Three-Layer Context Model: What Stays, What Loads, What Gets Dropped. Visualizes: Visualize the hierarchical three-layer context model described in the article.

A single flat context window is the wrong unit of analysis for any production agent workflow, because it treats architectural knowledge, task-specific code, and throwaway tool output as interchangeable when they need entirely different loading and eviction rules. The working model has three layers. Project-level context stays loaded at all times: architectural decisions, conventions, constraints, the design policy every agent turn must respect regardless of which file it happens to be touching. Supermemory points to maintained instruction files, CLAUDE.md and similar rules files, as the native implementation of this layer, and their value depends entirely on whether they stay maintained, how they load, and whether they remain relevant, not merely on the fact that they exist somewhere in the repo. The Bun rewrite's PORTING.md, paired with a separate lifetime inventory, stands as the clearest production example of this layer working at scale: thousands of local decisions converted into organization-wide policy, loaded as persistent context for every agent turn across the full duration of the project and across dozens of concurrent agents. Feature-level context loads per task: the structurally relevant chunks for the current edit, retrieved just-in-time using the call-graph and AST boundaries described above. The Codified Context research frames this as codified infrastructure: context built systematically for an agent's current task scope rather than assembled by hand or dumped in bulk ahead of time. Ephemeral context belongs to the current turn only: tool outputs, shell results, intermediate reasoning, all of it evicted aggressively once it has served its purpose. Tool Attention's lazy schema loading illustrates the enforcement mechanism in practice: schemas load only on invocation rather than being preloaded, and output persists only as long as the agent actually references it. Teams building against this model should plan, in advance, a substantial share of the window for accumulated work: system prompt and project context, tool results, reasoning and output, rather than discovering the squeeze only once the window fills mid-task. KVMem tackles the same problem from the infrastructure side: it extends the effective working set beyond what fits in a single GPU's KV cache and beyond the model's native context window, so a layered model like this can run on consumer hardware. What this layered structure prevents is the session-reset trap: without persistent project-level context, every new session forces a developer to re-explain the architecture from scratch, a tax that recurs every time work resumes.

Bun's agent-driven rewrite: the power and limits of context-as-policy

The Bun rewrite stands as the strongest publicly documented proof that documentation-as-policy can substitute for shared architectural context across massive concurrent agent work, but it also proves just as strongly that policy alone verifies nothing. In May 2026, the Bun team rewrote its entire runtime from Zig to Rust in a matter of days using AI agents, and a large Zig codebase became a substantially larger Rust codebase, with nearly all of the existing test suite passing on the very first AI-generated port. The orchestration behind that result followed a four-role loop applied per file: one implementer agent, two adversarial reviewer agents instructed to assume the code in front of them was wrong, and one fixer agent to apply whatever the reviewers surfaced, run at high peak concurrency across multiple git worktrees. Constraints were enforced structurally rather than left to agent discretion: git stash and git reset were banned outright, and every commit had to be atomic and per-file. The artifact that made this coordination possible was PORTING.md, a document mapping Zig types, memory patterns, error handling, metaprogramming, and project conventions to their Rust equivalents, paired with a separate lifetime inventory. Together they converted thousands of local decisions into something retrievable, persistent, and shared across the whole effort. Test parity was not sufficient verification of the result. The resulting Rust code contained 13,044 unsafe blocks, where hand-written Rust projects of similar size average approximately 73. Zig's creator, Andrew Kelley, called the output "unreviewed slop." Zig's manual memory management does not map cleanly onto Rust's ownership model, which caused agents to consistently reach for unsafe instead of redesigning the approach, because unsafe satisfies the immediate task of making code compile and pass tests without violating anything written in PORTING.md. PORTING.md described how to port individual constructs. It said nothing about what the resulting architecture needed to look like as a whole, so agents optimized for local correctness with no mechanism enforcing global structural constraints. Documentation-as-policy scales horizontally across large numbers of concurrent agents, but it cannot substitute for machine-enforceable architectural constraints that fire the moment policy is violated.

Diagram: Bun Rewrite: Policy Scaled, Verification Didn't. Visualizes: Show the tension between the Bun rewrite's scale (entire Zig runtime rewritten to Rust in days, nearly all existing tests passing on first AI-generated port, high peak…

Architectural constraints must be machine-enforced in CI

An agent carries no inherent understanding of a codebase's architecture, and no amount of context engineering supplies one. The only reliable backstop is a machine-enforced constraint that exits non-zero when the architecture is violated, blocking the merge before the code ever ships. The enforcement posture taking shape across the industry runs in three stages, and each one catches what the stage before it lets through. Pre-generation enforcement embeds architectural constraints directly into the system prompt and the project-level context, the same pattern PORTING.md represents, so it catches misunderstandings before an agent generates anything, though it cannot catch a choice that satisfies the prompt's letter while violating the system's global structure. Post-generation enforcement runs validators immediately after agent output and before a pull request ever opens, and you can see this pattern in approaches where agents submit work to a sandboxed review environment. CI validation then runs comprehensive architecture checks before any merge: tools like ArchUnit and ESLint alongside custom scripts, policy engines enforcing hard constraints, and minimum confidence thresholds that agent actions must clear before executing without human approval. Research into structured agentic workflows supports this layered posture: specialist agents working within defined roles and explicit constraints produce more structurally consistent output than generalist agents given broad latitude, because the workflow structure itself functions as an architectural constraint. The shape of code review is changing to match. Teams are reviewing less code line by line and more of the intent behind it, so the artifact under review is now the prompt, the constraints, and the reasoning chain the agent used to arrive at that diff. CI enforcement is what makes reviewing intent an actionable practice.

Treating the structural map of a codebase as the primary unit of context

The right unit of context for an AI agent is the structural map of the codebase: the graph of modules, call relationships, and dependency boundaries that makes every chunking decision and every retrieval decision architecture-aware. That graph is what defines natural chunk boundaries at the function, class, and module level, keeping each chunk semantically coherent on its own terms. It is what defines reachability for retrieval, determining which files a given edit actually touches so just-in-time retrieval pulls the chunks the task needs. It is the content of the persistent, project-level context layer, the architectural knowledge that survives session resets and stays shared across every concurrent agent working the codebase at once. And it is the machine-readable policy that CI enforcement checks generated code against, the artifact that exits non-zero the moment an agent's output drifts from the architecture it was supposed to respect. The graph has to live in the repository, version alongside the code, and stay enforceable in CI. A diagram drawn once and forgotten, or a PORTING.md an agent can satisfy locally while violating the system globally, does not meet that bar. The Codified Context research's framing of documentation as infrastructure and VITAL-RAG's invariance-aware allocation both converge on the same underlying claim: the structural invariants of a codebase are the one stable signal worth anchoring context engineering to, because they remain true at every scale the codebase grows into. What changes for the developer under this model is the job itself. Instead of reading diffs and hoping nothing critical slipped through, you navigate and steer from a living map, directing agent work against a shared architectural model and letting CI confirm that what came back actually matches it.

Sources

  1. VITAL-RAG: Invariance Race for Context Allocation in Coding Agents
  2. Codified Context: Infrastructure for AI Agents in a Complex Codebase
  3. Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows
  4. KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU
  5. Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution

More in Agentic Coding Workflows