AI on a 5-million-line codebase
A coding agent never holds a 5-million-line codebase in its context: it searches the repository, reads the few files that matter, makes a change and lets the build and tests say what to read next — the context window holds that working set, not the repository.
Why the whole codebase does not fit
Start with the size. Source code comes to roughly 10 tokens per line (estimate — it varies with language, indentation and tokenizer).
| Tokens | Lines of code | Share of the codebase | |
|---|---|---|---|
| A 5,000,000-line codebase | ~50,000,000 (estimate: 5,000,000 × 10) | 5,000,000 | 100% |
| Qwen3.8 27B's context window | 262,144 (vendor figure) | ~26,000 (estimate: 262,144 ÷ 10) | ~0.5% (calculated: 262,144 ÷ 50,000,000) |
| A 1,000,000-token window, as several vendors advertise | 1,000,000 (vendor figure) | ~100,000 (estimate: 1,000,000 ÷ 10) | 2% (calculated: 1,000,000 ÷ 50,000,000) |
Size is only the first problem. Even a window big enough to hold everything would run into three more:
- Memory. For every token in its context a model keeps a key and a value per KV head in every attention layer — the KV cache. For Llama 3.1 8B (32 layers, 8 KV heads, head_dim 128, 16-bit values) that is 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes, or 128 KiB per token (calculated from the model's published config). Its full 131,072-token context (vendor figure) takes 16 GiB for one request (calculated: 131,072 × 128 KiB). Fifty million tokens would need about 6.55 TB (calculated: 50,000,000 × 131,072 bytes) — roughly 50 times a DGX Spark's 128 GB of unified memory (vendor figure), before any weights — and the model cannot attend past 131,072 tokens anyway. Hybrid-attention models keep a growing cache in only some layers: Qwen3.8 27B has one in 16 of its 64 layers (4 KV heads, head_dim 256), so 2 × 16 × 4 × 256 × 2 bytes = 65,536 bytes, or 64 KiB per token, plus a fixed-size state in the other 48 layers (calculated from the model's published config). That is half of Llama 3.1 8B's figure per token, and it still grows with every token.
- Time. Before the first output token the model processes every token of the prompt (prefill). Twice the prompt is at least twice the work, and in standard attention the attention part grows with the square of the length. Prefix caching in vLLM and llama.cpp skips recomputing a prompt start that repeats exactly — useful across an agent's turns — but the first pass over a giant prompt is paid in full, and a metered API bills input tokens on every call.
- Quality. More text is not more understanding. Published research found that models use information at the start and end of a long input more reliably than in the middle ("Lost in the Middle", Liu et al., 2023), and the RULER benchmark (Hsieh et al., 2024) found the effective context of many models shorter than their advertised window.
And the repository changes with every merge. A search runs against the tree as it is now; a pasted prompt is stale by the next commit.
What an agent uses instead
No engineer reads 5 million lines either. They search, jump to definitions, read history and run tests. A coding agent gets the same tools as tool calls: the model asks, the harness runs the command, and the result comes back as text in the context. How that loop works in general is in How coding agents work; this page is about what changes at scale.
| Tool | Answers | Good at | Blind spot |
|---|---|---|---|
Text search: rg, git grep | Where does this name or string appear? | Exact identifiers, error messages, config keys; no setup; fast on millions of lines | Synonyms; code whose names you do not know |
| Symbol index: universal-ctags, a language server | Where is this defined? Who calls it? What implements it? | Precise navigation across modules (ctags lists definitions; callers and implementations need a language server) | Needs an indexer per language; reflection, dynamic dispatch, generated code |
Git history: git log -S, git blame | When and why did this change, and what changed with it? | Intent, recent regressions, the tests that came with a change | Says nothing about code that has not changed |
| Embeddings retrieval | Which code is about this concept? | Questions in plain language; unknown names | Returns similar code, not correct code; the index must be kept current |
| Build and tests | Is the change right? What else broke? | Ground truth; the error names the next file and line | Only as good as the coverage; slow suites stall the loop |
| Repository notes: README, an agent instructions file | How is this repository laid out? How do I build one module? | Orientation on the first turn | Goes stale unless someone maintains it |
Most agents lean on text search and file reads: Codex CLI, Claude Code and Cline run ripgrep, directly or through a built-in search tool, and read whole files or line ranges. aider works differently: it sends a compact "repo map" of the repository's symbols, built with tree-sitter, and asks to add the files the model wants to see. Some editors keep an embeddings index of the workspace as well.
The context window is a working set
- Every tool result is appended to the conversation, and the whole conversation is sent again on every model call.
- Harnesses trim old tool output, summarise the history, or hand a sub-task to a fresh context.
- Small results keep the loop going: file names before matching lines (
rg -l), line ranges before whole files, failing tests before the full log.
The loop
task: ticket, stack trace or failing test
|
v
+-- 1 LOCATE ------------------------------------+
| rg / git grep exact names, error text |
| symbol index definitions and callers |
| embeddings "code that does X" |
| git log / blame when and why it changed |
+------------------------------------------------+
| paths and line numbers
v
+-- 2 READ --------------------------------------+
| only those ranges enter the context |
+------------------------------------------------+
|
v
+-- 3 CHANGE ------------------------------------+
| the model writes an edit |
+------------------------------------------------+
|
v
+-- 4 CHECK -------------------------------------+
| compile, type-check, the module's tests |
+------------------------------------------------+
| |
fails: the error names passes
the next file and line |
| v
+--> back to 1 a diff for human review
The repository stays on disk. Each pass through the loop puts a few files into the context, and a failed check decides what the next pass looks for.
Retrieval with embeddings
Text search finds names you know. Embeddings find code by meaning — "where are invoice totals drawn on the PDF?" — when you do not know the names. The usual shape:
- Chunk along syntax. One function or class per chunk (tree-sitter or a language server gives the boundaries), each stored with its path and line range.
- Embed once, update on commit. Store each vector next to its path and range; re-embed only the files a commit changed.
- Query. Embed the question, rank chunks by cosine similarity, and hand the agent the top few as paths and line ranges to open. Run a text search for any identifiers in the question too: embeddings are weak on exact names, text search misses synonyms.
import math, os
from openai import OpenAI
client = OpenAI(base_url="https://api.axforge.ai/v1", api_key=os.environ["AXFORGE_API_KEY"])
# One chunk per function or class, keyed by where it lives.
# A real index stores the vectors and re-embeds only changed files.
chunks = {
"billing/pdf/InvoicePdfRenderer.java:180-240": "void renderTotals(Invoice inv) { ... }",
"billing/core/RetryPolicy.java:12-80": "class RetryPolicy { ... }",
"web/checkout/CartSummary.tsx:1-95": "export function CartSummary(props) { ... }",
}
query = "where are invoice totals drawn on the PDF?"
r = client.embeddings.create(model="qwen3-embed", input=[query, *chunks.values()])
q, *vecs = [d.embedding for d in sorted(r.data, key=lambda d: d.index)]
def cosine(a, b):
return sum(x * y for x, y in zip(a, b)) / (math.hypot(*a) * math.hypot(*b))
ranked = sorted(zip(chunks, vecs), key=lambda p: cosine(q, p[1]), reverse=True)
for where, _ in ranked[:2]:
print(where) # the agent opens these line ranges, then checks them
// An ES module (rank.mjs), so top-level await works
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.axforge.ai/v1", apiKey: process.env.AXFORGE_API_KEY });
// One chunk per function or class, keyed by where it lives.
const chunks = {
"billing/pdf/InvoicePdfRenderer.java:180-240": "void renderTotals(Invoice inv) { ... }",
"billing/core/RetryPolicy.java:12-80": "class RetryPolicy { ... }",
"web/checkout/CartSummary.tsx:1-95": "export function CartSummary(props) { ... }",
};
const query = "where are invoice totals drawn on the PDF?";
const r = await client.embeddings.create({ model: "qwen3-embed", input: [query, ...Object.values(chunks)] });
const [q, ...vecs] = r.data.sort((a, b) => a.index - b.index).map((d) => d.embedding);
const dot = (a, b) => a.reduce((s, x, i) => s + x * b[i], 0);
const cosine = (a, b) => dot(a, b) / Math.sqrt(dot(a, a) * dot(b, b));
Object.keys(chunks)
.map((where, i) => [where, cosine(q, vecs[i])])
.sort((a, b) => b[1] - a[1])
.slice(0, 2)
.forEach(([where]) => console.log(where)); // the agent opens these line ranges, then checks them
Retrieval returns candidates, not answers. Treat the top results like search hits: the agent opens them and checks. Request shape, batching and limits are in Embeddings.
Worked example: one bug in a 5-million-line repository
A monorepo with Java services and a TypeScript front end. The ticket: "Exporting an invoice with no line items fails with a NullPointerException in InvoicePdfRenderer.renderTotals." The agent's commands, roughly in order (an illustration — the repository is made up):
# 1. Locate the class named in the top stack frame
rg -n "class InvoicePdfRenderer" -t java
# 2. Find every caller of the failing method (-F: fixed string, not a regex)
rg -n -F "renderTotals(" -t java
# 3. Which commit added the assumption? (-S: commits that add or remove the string)
git log -S "lineItems.get(0)" --oneline -- billing/pdf
# 4. Check with the module's tests, not the whole repository's
./gradlew :billing:pdf:test --tests "*InvoicePdfRenderer*"
# 1. Locate the class named in the top stack frame
rg -n "class InvoicePdfRenderer" -t java
# 2. Find every caller of the failing method (-F: fixed string, not a regex)
rg -n -F "renderTotals(" -t java
# 3. Which commit added the assumption? (-S: commits that add or remove the string)
git log -S "lineItems.get(0)" --oneline -- billing/pdf
# 4. Check with the module's tests, not the whole repository's
.\gradlew :billing:pdf:test --tests "*InvoicePdfRenderer*"
What enters the context at each step:
| Step | What enters the context | Tokens (estimate) |
|---|---|---|
| Task | Ticket text and a 30-line stack trace | ~1,000 |
| Locate | rg hits: one class, three call sites, with paths | ~400 |
| Read | renderTotals and the two callers on the export path, ~300 lines from three files | ~3,000 |
| History | The commit that added lineItems.get(0): its message and a 60-line diff | ~1,000 |
| Change | The patch, which changes renderTotals' signature, and a new test for an empty invoice | ~800 |
| Check | Compile error: the third caller, which the agent did not read, still uses the old signature — file and line given | ~300 |
| Read + change | ~80 lines around that caller and a one-line fix | ~900 |
| Check | Tests pass: the summary line | ~200 |
| Total | About 0.015% of the codebase's ~50,000,000 tokens (calculated: 7,600 ÷ 50,000,000) | ~7,600 |
Add the agent's own instructions and tool definitions — say 8,000 tokens (estimate; it differs a lot between agents) — and the final context is about 15,600 tokens: under half of a 32,768-token window (calculated: 15,600 ÷ 32,768 = 48%). Because the whole conversation is sent again on every model call, the input adds up over the task: about a dozen calls averaging ~12,000 tokens is ~150,000 input tokens (estimate). Most of that is a repeated start that prefix caching need not recompute, and all of it is 0.3% of the codebase (calculated: 150,000 ÷ 50,000,000).
None of what made it work was context length: the stack trace gave an exact name, history explained the assumption, and a module-level build-and-test command took the agent from the patch to the caller it had missed.
Where long context still helps
A long window is useful once search has found the right place. It lets the working set be bigger:
- A whole module for a cross-cutting change. A 15,000-line package is about 150,000 tokens (estimate: 15,000 × 10) — inside a 262,144-token window, with room left for the conversation.
- One long file that does not split cleanly. A 4,000-line legacy class is about 40,000 tokens (estimate: 4,000 × 10).
- Logs, traces and CI output. There are no symbols to index, and the cause is often far from the last error line.
- A specification next to the code. An API spec, a protocol document or a migration guide beside the code it governs.
- Fewer round trips. Ten files in one read instead of ten tool calls.
A context length is a capability, not a target. Whatever goes into the window must also fit in memory as KV cache and is processed on every call — see Context windows and the KV cache.
The common mistake: buying context length instead of building search
The pattern: pick the model, or the GPU, with the biggest context window, then concatenate directories into the prompt. It fails for four reasons:
- It still is not the codebase. A 1,000,000-token window holds about 2% of it (calculated above). Choosing which 2% by directory is a cruder search, not no search.
- Every call pays for everything in it: prefill time before the first token, and input tokens on a metered API (pricing).
- Memory per request grows with the prompt, so fewer requests fit side by side on the same hardware — see Model size, VRAM and CPU offload.
- Irrelevant code is noise. The model has to look past it, and recall is weakest in the middle of a long input.
What to build instead is mostly cheap:
- Fast text search. ripgrep skips what
.gitignorelists; an.ignorefile keeps generated and vendored directories out of the results too. - A quick build-and-test command per module, written down where the agent will find it.
- An instructions file at the root: what lives where, how to build, how to test one module. Each agent has its own name for it (
AGENTS.md,CLAUDE.md,.clinerules). - A symbol index: universal-ctags for definitions, a language server for callers and implementations.
- An embeddings index when names are inconsistent or the questions are "code that does X".
Then spend the long context on what it is good at: the module, the long file, the log.
On AxForge
- Connect the agent you use: aider, Cline, Continue, Codex and Claude Code — all on Connect.
- Build a retrieval index with Qwen3 Embedding: Embeddings.
- Context windows per model: Models and the language model catalogue.
- Yes/no relevance checks with a probability you can threshold: Xev.
- Your own index and model on a rented machine: GPU rentals and the GPUs.
Next
- How coding agents work — the harness, the tools and the bug-fix loop.
- Context windows and the KV cache — what a long window costs in memory.
- Model size, VRAM and CPU offload — weights, KV cache and overhead on one GPU.