# How coding agents work


# How coding agents work

A coding agent is a loop: a harness program (aider, Cline, Continue, Codex CLI, Claude Code) runs tools on your repository &mdash; search, read, edit, shell, git, tests &mdash; and sends the results to a model, which decides each next step from the text it has been shown.

## Two parts: the harness and the model

People say &ldquo;the agent&rdquo; as if it were one thing. It is two, and they fail in different ways.

  |  | Harness (the agent program) | Model |  |

    | **What it is** | A program on your machine: a terminal tool or an editor extension | A model behind an API, on your machine or someone else&rsquo;s |  |

    | **What it can touch** | Your files, your shell, git, the test runner | Nothing. It only sees the text the harness sends |  |

    | **What it decides** | Which tools exist, how their output is cut, what goes into each request, when old turns are summarised, what needs your approval | Which tool to call next and with what arguments, what the output means, what to change, when to stop |  |

    | **Examples** | aider, Cline, Continue, Codex CLI, Claude Code | Qwen3.8 27B, or any model behind an OpenAI-compatible API |  |

The model never reads your disk. Everything it knows about your code is text the harness put into the request: files it asked for, command output, its own earlier steps. Change what the harness sends and you change what the model can do.

## The loop, drawn

```
 you: "tests/test_totals.py fails, fix it"
          |
          v
 +----------------------+  1. history + tools  +----------+
 | HARNESS              | -------------------> |  MODEL   |
 | aider, Cline,        |                      | (an API) |
 | Continue, Codex CLI, | <------------------- |          |
 | Claude Code          |  2. text, or a tool  +----------+
 +----------------------+     call
      |            ^
      | 3. run     | 4. output, trimmed
      v            |
 +----------------------+
 | TOOLS                |
 | search  read   edit  |
 | shell   git    tests |
 +----------------------+
      |            ^
      v            |
 +----------------------+
 | REPOSITORY           |
 | files, git history,  |
 | tests                |
 +----------------------+

 Steps 1-4 repeat until the model answers with text
 instead of a tool call. The model sees only text;
 only the harness touches the repository.
```

One turn is one API request. The harness sends the conversation so far plus the list of tools. The model answers with text (done, or a question for you) or with a tool call. The harness runs the call, appends the output and sends again. A typical bug fix takes 10 to 40 turns (estimate; it depends on how far the cause is from the symptom and how good the tests are).

## Tool calls: acting on files the model cannot see

A tool is a name, a description and a JSON schema for its arguments. A model trained for tool use answers with a structured request such as `{"name": "shell", "arguments": {"command": "pytest -q"}}` instead of prose. On an OpenAI-compatible API the tools go in the request&rsquo;s `tools` field, the call comes back in `tool_calls` (with `arguments` as a JSON string), and the harness returns the result as a `tool` message; Chat completions shows one full round trip.

  | Tool | What the harness runs | Why it exists |  |

    | **Search** | ripgrep or grep, file globs, symbol lookup | Find where something lives without reading everything |  |

    | **Read** | A file, or a line range of one | Bring the exact code into context |  |

    | **Edit** | A search-and-replace block or a patch | Change only the lines that need changing |  |

    | **Shell** | Any command, usually after your approval | Build, install, reproduce |  |

    | **Git** | `status`, `diff`, `log`, `blame` | See what changed, when and why |  |

    | **Tests** | The project&rsquo;s test command | The ground truth: did the change work |  |

Not every harness uses native tool calls. aider sends a map of the repository plus the files you added, the model replies with edit blocks in a text format, aider applies them, lints the edited files and can run your tests (`--test-cmd`, `--auto-test`). The loop is the same: act, observe, decide.

A minimal agent fits in about 40 lines. This one has a single tool, a shell:

```
import json
import os
import subprocess
from openai import OpenAI

client = OpenAI(base_url="https://api.axforge.ai/v1", api_key=os.environ["AXFORGE_API_KEY"])

tools = [{
    "type": "function",
    "function": {
        "name": "shell",
        "description": "Run a shell command in the repository root and return its output",
        "parameters": {
            "type": "object",
            "properties": {"command": {"type": "string"}},
            "required": ["command"],
        },
    },
}]

def shell(command):
    # a real harness sandboxes this and asks before risky commands
    try:
        p = subprocess.run(command, shell=True, capture_output=True, text=True, errors="replace", timeout=300)
    except subprocess.TimeoutExpired:
        return "timed out after 300 s"
    return (p.stdout + p.stderr)[-8000:]   # keep the tail: test failures end there

messages = [
    {"role": "system", "content": "Fix the bug. Search before reading whole files. Run the tests after every edit."},
    {"role": "user", "content": "tests/test_totals.py fails since the last refactor."},
]

for step in range(30):                      # a step budget, so a confused run still ends
    r = client.chat.completions.create(model="qwen3.8-27b-nvfp4", messages=messages, tools=tools)
    msg = r.choices[0].message
    messages.append(msg)
    if not msg.tool_calls:                  # text instead of a tool call: the model is done
        print(msg.content)
        break
    for call in msg.tool_calls:
        args = json.loads(call.function.arguments)
        messages.append({"role": "tool", "tool_call_id": call.id, "content": shell(args["command"])})
```

```
import OpenAI from "openai";
import { spawnSync } from "node:child_process";

const client = new OpenAI({ baseURL: "https://api.axforge.ai/v1", apiKey: process.env.AXFORGE_API_KEY });

const tools = [{
  type: "function",
  function: {
    name: "shell",
    description: "Run a shell command in the repository root and return its output",
    parameters: {
      type: "object",
      properties: { command: { type: "string" } },
      required: ["command"],
    },
  },
}];

function shell(command) {
  // a real harness sandboxes this and asks before risky commands
  const p = spawnSync(command, { shell: true, encoding: "utf8", timeout: 300_000 });
  return ((p.stdout ?? "") + (p.stderr ?? "")).slice(-8000);   // keep the tail: test failures end there
}

const messages = [
  { role: "system", content: "Fix the bug. Search before reading whole files. Run the tests after every edit." },
  { role: "user", content: "tests/test_totals.py fails since the last refactor." },
];

for (let step = 0; step < 30; step++) {     // a step budget, so a confused run still ends
  const r = await client.chat.completions.create({ model: "qwen3.8-27b-nvfp4", messages, tools });
  const msg = r.choices[0].message;
  messages.push(msg);
  if (!msg.tool_calls?.length) {            // text instead of a tool call: the model is done
    console.log(msg.content);
    break;
  }
  for (const call of msg.tool_calls) {
    const { command } = JSON.parse(call.function.arguments);
    messages.push({ role: "tool", tool_call_id: call.id, content: shell(command) });
  }
}
```

It works, and it shows what the rest of a harness is for. With only a shell, the model has to edit through `sed` or heredocs, reads a long file whole with `cat`, and nothing stops a destructive command. Real harnesses add a precise edit tool, ranged reads, output limits, approval prompts and context management. Those additions are most of what separates one agent from another.

## Exploring, not ingesting

An agent works the way you work in an unfamiliar codebase: you do not read it, you follow evidence.

  - **Orient** &mdash; list the tree, read the README or the project&rsquo;s instruction file.

  - **Locate** &mdash; search for the error text, the symbol, the route.

  - **Read narrowly** &mdash; the function, its callers, its test.

  - **Check** &mdash; run the test or a reproducer.

  - **Change and re-check** &mdash; edit, run again, read the new failure.

Each output picks the next step, so the same request can take a handful of steps on one run and several times as many on the next: one search that misses sends the exploration elsewhere, and sampling adds its own variation.

Harnesses speed up the locate step in two ways. A **repo map**: aider ranks the repository&rsquo;s files and key symbols by how they reference each other and cuts the summary to a token budget (`--map-tokens`, 1k tokens by default &mdash; vendor figure, aider&rsquo;s documentation). A **retrieval index**: an embedding model turns chunks of code into vectors, so a plain-language question finds likely chunks. An index is good at &ldquo;where is retry handled?&rdquo;; grep is better at exact names.

## Managing the context window

Everything the model knows during a task sits in its context window, and with a stateless API the harness resends all of it on every turn:

  - the harness&rsquo;s system prompt and tool definitions &mdash; often several thousand tokens (estimate; varies a lot by harness),

  - your project&rsquo;s instruction file,

  - your request,

  - every tool call and every tool output so far.

So the context grows with each step, and the input adds up. Example (estimate): a session starts at 8,000 tokens and each step adds 1,500. Step 20 sends 8,000 + 19 &times; 1,500 = 36,500 tokens. The 20 requests together send 20 &times; 8,000 + 1,500 &times; (0 + 1 + &hellip; + 19) = 160,000 + 285,000 = 445,000 input tokens. Runtimes with prefix caching can skip recomputing an unchanged start of the conversation, but those tokens still fill the window.

What harnesses do about it:

  - **Narrow reads and searches** &mdash; line ranges and match lists, not whole files.

  - **Truncated tool output** &mdash; the end of a test run, not the whole log.

  - **Compaction** &mdash; older turns replaced by a summary when the window fills, automatically or on a command.

  - **Sub-agents** &mdash; a side investigation in its own context that hands back only its conclusion.

  - **An instruction file** &mdash; `AGENTS.md` (Codex CLI and others), `CLAUDE.md` (Claude Code), `.clinerules` (Cline): layout, test command, conventions. It is loaded every session, so keep it short.

Two things are yours to set. Tell the harness the model&rsquo;s real context window, because it plans truncation and compaction from that number (the Cline setup shows where). And start a fresh session per task. Qwen3.8 27B&rsquo;s 262,144-token window (vendor figure) is a capability, not a target: the model accepts that much, but every turn gets slower as the context grows, and models use information in the middle of a long context less reliably than at its start or end (Liu et al., 2023, &ldquo;Lost in the Middle&rdquo;, arXiv 2307.03172).

## Why the model still matters

The harness decides what the model can do. The model decides what gets done. No harness can do these for it:

  - turn a vague bug report into the right search,

  - read a stack trace and fix the cause, not the line that threw,

  - make the smallest correct edit instead of rewriting the file,

  - keep the plan straight across dozens of steps and tens of thousands of tokens of history,

  - emit valid tool-call arguments every time &mdash; each malformed call costs a turn, and a run of them can end the session,

  - stop honestly: run the tests before claiming success, and never edit a test to make it pass.

Smaller models fail these first, typically by repeating the same search, editing the wrong file or declaring victory early. Agent benchmarks such as SWE-bench measure this whole loop &mdash; resolve a real GitHub issue so that the project&rsquo;s tests pass &mdash; which makes them a better guide for agent work than single-answer coding tests. Read every such score together with the harness it was run in, because the harness is part of the result.

## Why the harness changes results

Same model, different harness, different result. The SWE-agent paper isolated this: GPT-4 Turbo resolved 11.0% of SWE-bench Lite issues with a plain Linux shell and 18.0% with purpose-built commands for searching, viewing and editing files (published research: Yang et al., 2024, &ldquo;SWE-agent&rdquo;, arXiv 2405.15793). The model was the same; the interface changed.

  | The harness decides | Why it matters |  |

    | Tool set and tool descriptions | The model can only do what is offered, and picks tools by their descriptions |  |

    | Edit format | Whole file, search-and-replace or diff &mdash; models differ in which they apply reliably; aider sets a default format per model |  |

    | Output limits | Too little hides the error, too much buries it |  |

    | Repo map or index | How quickly the model finds the right file |  |

    | Compaction | What gets forgotten when the window fills |  |

    | Test running | Whether failures come back without the model asking |  |

    | Approval and sandbox | What can happen without you seeing it |  |

So compare harnesses with the same model on tasks from your own repository, and compare models inside the same harness.

## If you serve the model yourself

A harness that uses native tool calls needs the server to turn the model&rsquo;s tool-call text back into structured `tool_calls`. Both common open runtimes do it. In vLLM, enable automatic tool choice and pick the parser for the model family (`hermes` for Qwen2.5, per vLLM&rsquo;s tool-calling docs):

```
vllm serve Qwen/Qwen2.5-7B-Instruct --max-model-len 32768 --enable-auto-tool-choice --tool-call-parser hermes
```

llama.cpp&rsquo;s `llama-server` parses tool calls through the model&rsquo;s Jinja chat template. Current builds turn Jinja on by default; `--jinja` makes it explicit and older builds need it:

```
llama-server -m Qwen2.5-7B-Instruct-Q4_K_M.gguf -c 32768 -ngl 99 --jinja
```

The two lines need different hardware. vLLM loads the 16-bit weights: 7.61 billion parameters &times; 2 bytes = 15.2 GB (calculated from the published parameter count), more than one 12 GB RTX 3060 holds. A Q4_K_M GGUF of it is about 4.7 GB (vendor figure: Qwen&rsquo;s own Q4_K_M files total 4.68 GB). Either way the KV cache comes on top: 2 (K and V) &times; 28 layers &times; 4 KV heads &times; 128 &times; 2 bytes = 57,344 bytes per token, &times; 32,768 tokens = 1.75 GiB (calculated from the model&rsquo;s published config). Set the context (`--max-model-len`, `-c`) to what your sessions need and the memory allows, enter the same number in the harness, and check the fit with Model size, GPU memory and CPU offload. If a harness reports that the model does not support tools, or the model prints its tool call as plain text, check these flags first.

## Worked example: one bug-fix loop

A refactor moved one line in an invoice total. With a subtotal of 100.00, a fixed discount of 10.00 and 25% VAT, the right total is (100.00 &minus; 10.00) &times; 1.25 = 112.50. The code now computes 100.00 &times; 1.25 &minus; 10.00 = 115.00. Your request: &ldquo;`tests/test_totals.py` fails since the last refactor. Fix it.&rdquo; An illustrative session; token counts are estimates.

  | # | The model asks for | What comes back | Context added |  |

    | 1 | Search: `rg -n "discount" src/billing` | 9 matching lines in 3 files | ~250 tokens |  |

    | 2 | Read: `src/billing/totals.py`, lines 1&ndash;80 | The `total()` function | ~900 |  |

    | 3 | Tests: `pytest tests/test_totals.py -q` | 1 failed: `assert 115.0 == 112.5` | ~600 |  |

    | 4 | Git: `git log -p -2 -- src/billing/totals.py` | The refactor moved the discount below the VAT line | ~1,200 |  |

    | 5 | Edit: subtract the discount before VAT | 2 lines changed | ~200 |  |

    | 6 | Tests: the same command | 1 failed: `assert -12.5 == 0.0` (discount 20.00 on a subtotal of 10.00) | ~600 |  |

    | 7 | Edit: `max(subtotal - discount, 0)` | 1 line changed | ~150 |  |

    | 8 | Tests: `pytest -q`, the whole suite | All pass | ~100 |  |

    | 9 | Git: `git diff` | 3 changed lines | ~250 |  |

    | 10 | No tool: a text answer | Cause, fix and diff, for you to review | &mdash; |  |

The tool traffic adds 250 + 900 + 600 + 1,200 + 200 + 600 + 150 + 100 + 250 = 4,250 tokens. With 6,000 tokens of harness prompt and tool definitions (estimate) the session ends near 10,250 tokens, about 4% of a 262,144-token window (10,250 / 262,144 = 3.9%).

What decided the outcome:

  - **The model&rsquo;s judgment**: search before reading (step 1), and ask git why the line moved (step 4) instead of guessing.

  - **The tests**: the first edit dropped the clamp the old line carried (step 6). Running the tests caught it, not reading more code. Without a test command the session ends at step 5 with a wrong fix.

  - **The harness**: it cut each test run to its tail, asked before running commands, and put the diff in front of you.

## The common mistake: pasting the whole repository

Pasting the entire codebase into the prompt feels thorough. It makes the result worse:

  - **It often does not fit.** A 300,000-line codebase at about 10 tokens per line (estimate; depends on language and tokenizer) is about 3,000,000 tokens, more than 11 times a 262,144-token window.

  - **When it fits, it is resent every turn.** A 200,000-token paste over 20 steps is at least 20 &times; 200,000 = 4,000,000 input tokens, against about 445,000 for the whole exploring session estimated above.

  - **It costs memory when you serve the model.** Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128. In FP16 its KV cache takes 2 (K and V) &times; 32 &times; 8 &times; 128 &times; 2 bytes = 131,072 bytes (128 KiB) per token, so its full 131,072-token context needs 131,072 &times; 128 KiB = 16 GiB (calculated from the model&rsquo;s published config). That is more than a 12 GB RTX 3060 holds before a single weight is loaded.

  - **It dilutes attention.** The lines that matter sit somewhere in 200,000 tokens, most likely in the middle.

  - **It goes stale.** After the first edit the pasted copy is wrong; a read tool always sees the current file.

Give the agent what a senior colleague would want instead: the failing test or the error, where to start, the test command, and what not to touch. Put the durable parts in the instruction file.

## On AxForge

  - Connect an agent to the API: aider, Cline, Continue, Codex CLI & Claude Code; every tool is on Connect.

  - Tool calls on the chat endpoint, one full round trip: Chat completions.

  - Model ids and context windows to enter in the harness: Models.

  - Serve your own model for an agent: GPU rentals, the RTX 3060 and DGX Spark pages, and the language models in the catalogue.

  - A key and a first call: Quickstart and the console.

## Next

  - AI on a 5-million-line codebase &mdash; the same loop when nothing fits and search is everything.

  - Context windows and the KV cache &mdash; what a long agent session costs in memory and speed.

  - llama.cpp vs vLLM &mdash; choosing a runtime to serve an agent yourself.



Source: https://dev.axforge.ai/ai-engineering/how-coding-agents-work/
