How coding agents work
A coding agent is a loop: a harness program (aider, Cline, Continue, Codex CLI, Claude Code) runs tools on your repository — search, read, edit, shell, git, tests — and sends the results to a model, which decides each next step from the text it has been shown.
Two parts: the harness and the model
People say “the agent” as if it were one thing. It is two, and they fail in different ways.
| Harness (the agent program) | Model | |
|---|---|---|
| What it is | A program on your machine: a terminal tool or an editor extension | A model behind an API, on your machine or someone else’s |
| What it can touch | Your files, your shell, git, the test runner | Nothing. It only sees the text the harness sends |
| What it decides | Which tools exist, how their output is cut, what goes into each request, when old turns are summarised, what needs your approval | Which tool to call next and with what arguments, what the output means, what to change, when to stop |
| Examples | aider, Cline, Continue, Codex CLI, Claude Code | Qwen3.8 27B, or any model behind an OpenAI-compatible API |
The model never reads your disk. Everything it knows about your code is text the harness put into the request: files it asked for, command output, its own earlier steps. Change what the harness sends and you change what the model can do.
The loop, drawn
you: "tests/test_totals.py fails, fix it"
|
v
+----------------------+ 1. history + tools +----------+
| HARNESS | -------------------> | MODEL |
| aider, Cline, | | (an API) |
| Continue, Codex CLI, | <------------------- | |
| Claude Code | 2. text, or a tool +----------+
+----------------------+ call
| ^
| 3. run | 4. output, trimmed
v |
+----------------------+
| TOOLS |
| search read edit |
| shell git tests |
+----------------------+
| ^
v |
+----------------------+
| REPOSITORY |
| files, git history, |
| tests |
+----------------------+
Steps 1-4 repeat until the model answers with text
instead of a tool call. The model sees only text;
only the harness touches the repository.
One turn is one API request. The harness sends the conversation so far plus the list of tools. The model answers with text (done, or a question for you) or with a tool call. The harness runs the call, appends the output and sends again. A typical bug fix takes 10 to 40 turns (estimate; it depends on how far the cause is from the symptom and how good the tests are).
Tool calls: acting on files the model cannot see
A tool is a name, a description and a JSON schema for its arguments. A model trained for tool use answers with a structured request such as {"name": "shell", "arguments": {"command": "pytest -q"}} instead of prose. On an OpenAI-compatible API the tools go in the request’s tools field, the call comes back in tool_calls (with arguments as a JSON string), and the harness returns the result as a tool message; Chat completions shows one full round trip.
| Tool | What the harness runs | Why it exists |
|---|---|---|
| Search | ripgrep or grep, file globs, symbol lookup | Find where something lives without reading everything |
| Read | A file, or a line range of one | Bring the exact code into context |
| Edit | A search-and-replace block or a patch | Change only the lines that need changing |
| Shell | Any command, usually after your approval | Build, install, reproduce |
| Git | status, diff, log, blame | See what changed, when and why |
| Tests | The project’s test command | The ground truth: did the change work |
Not every harness uses native tool calls. aider sends a map of the repository plus the files you added, the model replies with edit blocks in a text format, aider applies them, lints the edited files and can run your tests (--test-cmd, --auto-test). The loop is the same: act, observe, decide.
A minimal agent fits in about 40 lines. This one has a single tool, a shell:
import json
import os
import subprocess
from openai import OpenAI
client = OpenAI(base_url="https://api.axforge.ai/v1", api_key=os.environ["AXFORGE_API_KEY"])
tools = [{
"type": "function",
"function": {
"name": "shell",
"description": "Run a shell command in the repository root and return its output",
"parameters": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"],
},
},
}]
def shell(command):
# a real harness sandboxes this and asks before risky commands
try:
p = subprocess.run(command, shell=True, capture_output=True, text=True, errors="replace", timeout=300)
except subprocess.TimeoutExpired:
return "timed out after 300 s"
return (p.stdout + p.stderr)[-8000:] # keep the tail: test failures end there
messages = [
{"role": "system", "content": "Fix the bug. Search before reading whole files. Run the tests after every edit."},
{"role": "user", "content": "tests/test_totals.py fails since the last refactor."},
]
for step in range(30): # a step budget, so a confused run still ends
r = client.chat.completions.create(model="qwen3.8-27b-nvfp4", messages=messages, tools=tools)
msg = r.choices[0].message
messages.append(msg)
if not msg.tool_calls: # text instead of a tool call: the model is done
print(msg.content)
break
for call in msg.tool_calls:
args = json.loads(call.function.arguments)
messages.append({"role": "tool", "tool_call_id": call.id, "content": shell(args["command"])})
import OpenAI from "openai";
import { spawnSync } from "node:child_process";
const client = new OpenAI({ baseURL: "https://api.axforge.ai/v1", apiKey: process.env.AXFORGE_API_KEY });
const tools = [{
type: "function",
function: {
name: "shell",
description: "Run a shell command in the repository root and return its output",
parameters: {
type: "object",
properties: { command: { type: "string" } },
required: ["command"],
},
},
}];
function shell(command) {
// a real harness sandboxes this and asks before risky commands
const p = spawnSync(command, { shell: true, encoding: "utf8", timeout: 300_000 });
return ((p.stdout ?? "") + (p.stderr ?? "")).slice(-8000); // keep the tail: test failures end there
}
const messages = [
{ role: "system", content: "Fix the bug. Search before reading whole files. Run the tests after every edit." },
{ role: "user", content: "tests/test_totals.py fails since the last refactor." },
];
for (let step = 0; step < 30; step++) { // a step budget, so a confused run still ends
const r = await client.chat.completions.create({ model: "qwen3.8-27b-nvfp4", messages, tools });
const msg = r.choices[0].message;
messages.push(msg);
if (!msg.tool_calls?.length) { // text instead of a tool call: the model is done
console.log(msg.content);
break;
}
for (const call of msg.tool_calls) {
const { command } = JSON.parse(call.function.arguments);
messages.push({ role: "tool", tool_call_id: call.id, content: shell(command) });
}
}
It works, and it shows what the rest of a harness is for. With only a shell, the model has to edit through sed or heredocs, reads a long file whole with cat, and nothing stops a destructive command. Real harnesses add a precise edit tool, ranged reads, output limits, approval prompts and context management. Those additions are most of what separates one agent from another.
Exploring, not ingesting
An agent works the way you work in an unfamiliar codebase: you do not read it, you follow evidence.
- Orient — list the tree, read the README or the project’s instruction file.
- Locate — search for the error text, the symbol, the route.
- Read narrowly — the function, its callers, its test.
- Check — run the test or a reproducer.
- Change and re-check — edit, run again, read the new failure.
Each output picks the next step, so the same request can take a handful of steps on one run and several times as many on the next: one search that misses sends the exploration elsewhere, and sampling adds its own variation.
Harnesses speed up the locate step in two ways. A repo map: aider ranks the repository’s files and key symbols by how they reference each other and cuts the summary to a token budget (--map-tokens, 1k tokens by default — vendor figure, aider’s documentation). A retrieval index: an embedding model turns chunks of code into vectors, so a plain-language question finds likely chunks. An index is good at “where is retry handled?”; grep is better at exact names.
Managing the context window
Everything the model knows during a task sits in its context window, and with a stateless API the harness resends all of it on every turn:
- the harness’s system prompt and tool definitions — often several thousand tokens (estimate; varies a lot by harness),
- your project’s instruction file,
- your request,
- every tool call and every tool output so far.
So the context grows with each step, and the input adds up. Example (estimate): a session starts at 8,000 tokens and each step adds 1,500. Step 20 sends 8,000 + 19 × 1,500 = 36,500 tokens. The 20 requests together send 20 × 8,000 + 1,500 × (0 + 1 + … + 19) = 160,000 + 285,000 = 445,000 input tokens. Runtimes with prefix caching can skip recomputing an unchanged start of the conversation, but those tokens still fill the window.
What harnesses do about it:
- Narrow reads and searches — line ranges and match lists, not whole files.
- Truncated tool output — the end of a test run, not the whole log.
- Compaction — older turns replaced by a summary when the window fills, automatically or on a command.
- Sub-agents — a side investigation in its own context that hands back only its conclusion.
- An instruction file —
AGENTS.md(Codex CLI and others),CLAUDE.md(Claude Code),.clinerules(Cline): layout, test command, conventions. It is loaded every session, so keep it short.
Two things are yours to set. Tell the harness the model’s real context window, because it plans truncation and compaction from that number (the Cline setup shows where). And start a fresh session per task. Qwen3.8 27B’s 262,144-token window (vendor figure) is a capability, not a target: the model accepts that much, but every turn gets slower as the context grows, and models use information in the middle of a long context less reliably than at its start or end (Liu et al., 2023, “Lost in the Middle”, arXiv 2307.03172).
Why the model still matters
The harness decides what the model can do. The model decides what gets done. No harness can do these for it:
- turn a vague bug report into the right search,
- read a stack trace and fix the cause, not the line that threw,
- make the smallest correct edit instead of rewriting the file,
- keep the plan straight across dozens of steps and tens of thousands of tokens of history,
- emit valid tool-call arguments every time — each malformed call costs a turn, and a run of them can end the session,
- stop honestly: run the tests before claiming success, and never edit a test to make it pass.
Smaller models fail these first, typically by repeating the same search, editing the wrong file or declaring victory early. Agent benchmarks such as SWE-bench measure this whole loop — resolve a real GitHub issue so that the project’s tests pass — which makes them a better guide for agent work than single-answer coding tests. Read every such score together with the harness it was run in, because the harness is part of the result.
Why the harness changes results
Same model, different harness, different result. The SWE-agent paper isolated this: GPT-4 Turbo resolved 11.0% of SWE-bench Lite issues with a plain Linux shell and 18.0% with purpose-built commands for searching, viewing and editing files (published research: Yang et al., 2024, “SWE-agent”, arXiv 2405.15793). The model was the same; the interface changed.
| The harness decides | Why it matters |
|---|---|
| Tool set and tool descriptions | The model can only do what is offered, and picks tools by their descriptions |
| Edit format | Whole file, search-and-replace or diff — models differ in which they apply reliably; aider sets a default format per model |
| Output limits | Too little hides the error, too much buries it |
| Repo map or index | How quickly the model finds the right file |
| Compaction | What gets forgotten when the window fills |
| Test running | Whether failures come back without the model asking |
| Approval and sandbox | What can happen without you seeing it |
So compare harnesses with the same model on tasks from your own repository, and compare models inside the same harness.
If you serve the model yourself
A harness that uses native tool calls needs the server to turn the model’s tool-call text back into structured tool_calls. Both common open runtimes do it. In vLLM, enable automatic tool choice and pick the parser for the model family (hermes for Qwen2.5, per vLLM’s tool-calling docs):
vllm serve Qwen/Qwen2.5-7B-Instruct --max-model-len 32768 --enable-auto-tool-choice --tool-call-parser hermes
llama.cpp’s llama-server parses tool calls through the model’s Jinja chat template. Current builds turn Jinja on by default; --jinja makes it explicit and older builds need it:
llama-server -m Qwen2.5-7B-Instruct-Q4_K_M.gguf -c 32768 -ngl 99 --jinja
The two lines need different hardware. vLLM loads the 16-bit weights: 7.61 billion parameters × 2 bytes = 15.2 GB (calculated from the published parameter count), more than one 12 GB RTX 3060 holds. A Q4_K_M GGUF of it is about 4.7 GB (vendor figure: Qwen’s own Q4_K_M files total 4.68 GB). Either way the KV cache comes on top: 2 (K and V) × 28 layers × 4 KV heads × 128 × 2 bytes = 57,344 bytes per token, × 32,768 tokens = 1.75 GiB (calculated from the model’s published config). Set the context (--max-model-len, -c) to what your sessions need and the memory allows, enter the same number in the harness, and check the fit with Model size, GPU memory and CPU offload. If a harness reports that the model does not support tools, or the model prints its tool call as plain text, check these flags first.
Worked example: one bug-fix loop
A refactor moved one line in an invoice total. With a subtotal of 100.00, a fixed discount of 10.00 and 25% VAT, the right total is (100.00 − 10.00) × 1.25 = 112.50. The code now computes 100.00 × 1.25 − 10.00 = 115.00. Your request: “tests/test_totals.py fails since the last refactor. Fix it.” An illustrative session; token counts are estimates.
| # | The model asks for | What comes back | Context added |
|---|---|---|---|
| 1 | Search: rg -n "discount" src/billing | 9 matching lines in 3 files | ~250 tokens |
| 2 | Read: src/billing/totals.py, lines 1–80 | The total() function | ~900 |
| 3 | Tests: pytest tests/test_totals.py -q | 1 failed: assert 115.0 == 112.5 | ~600 |
| 4 | Git: git log -p -2 -- src/billing/totals.py | The refactor moved the discount below the VAT line | ~1,200 |
| 5 | Edit: subtract the discount before VAT | 2 lines changed | ~200 |
| 6 | Tests: the same command | 1 failed: assert -12.5 == 0.0 (discount 20.00 on a subtotal of 10.00) | ~600 |
| 7 | Edit: max(subtotal - discount, 0) | 1 line changed | ~150 |
| 8 | Tests: pytest -q, the whole suite | All pass | ~100 |
| 9 | Git: git diff | 3 changed lines | ~250 |
| 10 | No tool: a text answer | Cause, fix and diff, for you to review | — |
The tool traffic adds 250 + 900 + 600 + 1,200 + 200 + 600 + 150 + 100 + 250 = 4,250 tokens. With 6,000 tokens of harness prompt and tool definitions (estimate) the session ends near 10,250 tokens, about 4% of a 262,144-token window (10,250 / 262,144 = 3.9%).
What decided the outcome:
- The model’s judgment: search before reading (step 1), and ask git why the line moved (step 4) instead of guessing.
- The tests: the first edit dropped the clamp the old line carried (step 6). Running the tests caught it, not reading more code. Without a test command the session ends at step 5 with a wrong fix.
- The harness: it cut each test run to its tail, asked before running commands, and put the diff in front of you.
The common mistake: pasting the whole repository
Pasting the entire codebase into the prompt feels thorough. It makes the result worse:
- It often does not fit. A 300,000-line codebase at about 10 tokens per line (estimate; depends on language and tokenizer) is about 3,000,000 tokens, more than 11 times a 262,144-token window.
- When it fits, it is resent every turn. A 200,000-token paste over 20 steps is at least 20 × 200,000 = 4,000,000 input tokens, against about 445,000 for the whole exploring session estimated above.
- It costs memory when you serve the model. Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128. In FP16 its KV cache takes 2 (K and V) × 32 × 8 × 128 × 2 bytes = 131,072 bytes (128 KiB) per token, so its full 131,072-token context needs 131,072 × 128 KiB = 16 GiB (calculated from the model’s published config). That is more than a 12 GB RTX 3060 holds before a single weight is loaded.
- It dilutes attention. The lines that matter sit somewhere in 200,000 tokens, most likely in the middle.
- It goes stale. After the first edit the pasted copy is wrong; a read tool always sees the current file.
Give the agent what a senior colleague would want instead: the failing test or the error, where to start, the test command, and what not to touch. Put the durable parts in the instruction file.
On AxForge
- Connect an agent to the API: aider, Cline, Continue, Codex CLI & Claude Code; every tool is on Connect.
- Tool calls on the chat endpoint, one full round trip: Chat completions.
- Model ids and context windows to enter in the harness: Models.
- Serve your own model for an agent: GPU rentals, the RTX 3060 and DGX Spark pages, and the language models in the catalogue.
- A key and a first call: Quickstart and the console.
Next
- AI on a 5-million-line codebase — the same loop when nothing fits and search is everything.
- Context windows and the KV cache — what a long agent session costs in memory and speed.
- llama.cpp vs vLLM — choosing a runtime to serve an agent yourself.