AI Engineering / Coding & Agents

How coding agents work

A coding agent is a loop: a harness program (aider, Cline, Continue, Codex CLI, Claude Code) runs tools on your repository — search, read, edit, shell, git, tests — and sends the results to a model, which decides each next step from the text it has been shown.

Two parts: the harness and the model

People say “the agent” as if it were one thing. It is two, and they fail in different ways.

Harness (the agent program)Model
What it isA program on your machine: a terminal tool or an editor extensionA model behind an API, on your machine or someone else’s
What it can touchYour files, your shell, git, the test runnerNothing. It only sees the text the harness sends
What it decidesWhich tools exist, how their output is cut, what goes into each request, when old turns are summarised, what needs your approvalWhich tool to call next and with what arguments, what the output means, what to change, when to stop
Examplesaider, Cline, Continue, Codex CLI, Claude CodeQwen3.8 27B, or any model behind an OpenAI-compatible API

The model never reads your disk. Everything it knows about your code is text the harness put into the request: files it asked for, command output, its own earlier steps. Change what the harness sends and you change what the model can do.

The loop, drawn

 you: "tests/test_totals.py fails, fix it"
          |
          v
 +----------------------+  1. history + tools  +----------+
 | HARNESS              | -------------------> |  MODEL   |
 | aider, Cline,        |                      | (an API) |
 | Continue, Codex CLI, | <------------------- |          |
 | Claude Code          |  2. text, or a tool  +----------+
 +----------------------+     call
      |            ^
      | 3. run     | 4. output, trimmed
      v            |
 +----------------------+
 | TOOLS                |
 | search  read   edit  |
 | shell   git    tests |
 +----------------------+
      |            ^
      v            |
 +----------------------+
 | REPOSITORY           |
 | files, git history,  |
 | tests                |
 +----------------------+

 Steps 1-4 repeat until the model answers with text
 instead of a tool call. The model sees only text;
 only the harness touches the repository.

One turn is one API request. The harness sends the conversation so far plus the list of tools. The model answers with text (done, or a question for you) or with a tool call. The harness runs the call, appends the output and sends again. A typical bug fix takes 10 to 40 turns (estimate; it depends on how far the cause is from the symptom and how good the tests are).

Tool calls: acting on files the model cannot see

A tool is a name, a description and a JSON schema for its arguments. A model trained for tool use answers with a structured request such as {"name": "shell", "arguments": {"command": "pytest -q"}} instead of prose. On an OpenAI-compatible API the tools go in the request’s tools field, the call comes back in tool_calls (with arguments as a JSON string), and the harness returns the result as a tool message; Chat completions shows one full round trip.

ToolWhat the harness runsWhy it exists
Searchripgrep or grep, file globs, symbol lookupFind where something lives without reading everything
ReadA file, or a line range of oneBring the exact code into context
EditA search-and-replace block or a patchChange only the lines that need changing
ShellAny command, usually after your approvalBuild, install, reproduce
Gitstatus, diff, log, blameSee what changed, when and why
TestsThe project’s test commandThe ground truth: did the change work

Not every harness uses native tool calls. aider sends a map of the repository plus the files you added, the model replies with edit blocks in a text format, aider applies them, lints the edited files and can run your tests (--test-cmd, --auto-test). The loop is the same: act, observe, decide.

A minimal agent fits in about 40 lines. This one has a single tool, a shell:

import json
import os
import subprocess
from openai import OpenAI

client = OpenAI(base_url="https://api.axforge.ai/v1", api_key=os.environ["AXFORGE_API_KEY"])

tools = [{
    "type": "function",
    "function": {
        "name": "shell",
        "description": "Run a shell command in the repository root and return its output",
        "parameters": {
            "type": "object",
            "properties": {"command": {"type": "string"}},
            "required": ["command"],
        },
    },
}]

def shell(command):
    # a real harness sandboxes this and asks before risky commands
    try:
        p = subprocess.run(command, shell=True, capture_output=True, text=True, errors="replace", timeout=300)
    except subprocess.TimeoutExpired:
        return "timed out after 300 s"
    return (p.stdout + p.stderr)[-8000:]   # keep the tail: test failures end there

messages = [
    {"role": "system", "content": "Fix the bug. Search before reading whole files. Run the tests after every edit."},
    {"role": "user", "content": "tests/test_totals.py fails since the last refactor."},
]

for step in range(30):                      # a step budget, so a confused run still ends
    r = client.chat.completions.create(model="qwen3.8-27b-nvfp4", messages=messages, tools=tools)
    msg = r.choices[0].message
    messages.append(msg)
    if not msg.tool_calls:                  # text instead of a tool call: the model is done
        print(msg.content)
        break
    for call in msg.tool_calls:
        args = json.loads(call.function.arguments)
        messages.append({"role": "tool", "tool_call_id": call.id, "content": shell(args["command"])})

It works, and it shows what the rest of a harness is for. With only a shell, the model has to edit through sed or heredocs, reads a long file whole with cat, and nothing stops a destructive command. Real harnesses add a precise edit tool, ranged reads, output limits, approval prompts and context management. Those additions are most of what separates one agent from another.

Exploring, not ingesting

An agent works the way you work in an unfamiliar codebase: you do not read it, you follow evidence.

  1. Orient — list the tree, read the README or the project’s instruction file.
  2. Locate — search for the error text, the symbol, the route.
  3. Read narrowly — the function, its callers, its test.
  4. Check — run the test or a reproducer.
  5. Change and re-check — edit, run again, read the new failure.

Each output picks the next step, so the same request can take a handful of steps on one run and several times as many on the next: one search that misses sends the exploration elsewhere, and sampling adds its own variation.

Harnesses speed up the locate step in two ways. A repo map: aider ranks the repository’s files and key symbols by how they reference each other and cuts the summary to a token budget (--map-tokens, 1k tokens by default — vendor figure, aider’s documentation). A retrieval index: an embedding model turns chunks of code into vectors, so a plain-language question finds likely chunks. An index is good at “where is retry handled?”; grep is better at exact names.

Managing the context window

Everything the model knows during a task sits in its context window, and with a stateless API the harness resends all of it on every turn:

  • the harness’s system prompt and tool definitions — often several thousand tokens (estimate; varies a lot by harness),
  • your project’s instruction file,
  • your request,
  • every tool call and every tool output so far.

So the context grows with each step, and the input adds up. Example (estimate): a session starts at 8,000 tokens and each step adds 1,500. Step 20 sends 8,000 + 19 × 1,500 = 36,500 tokens. The 20 requests together send 20 × 8,000 + 1,500 × (0 + 1 + … + 19) = 160,000 + 285,000 = 445,000 input tokens. Runtimes with prefix caching can skip recomputing an unchanged start of the conversation, but those tokens still fill the window.

What harnesses do about it:

  • Narrow reads and searches — line ranges and match lists, not whole files.
  • Truncated tool output — the end of a test run, not the whole log.
  • Compaction — older turns replaced by a summary when the window fills, automatically or on a command.
  • Sub-agents — a side investigation in its own context that hands back only its conclusion.
  • An instruction file — AGENTS.md (Codex CLI and others), CLAUDE.md (Claude Code), .clinerules (Cline): layout, test command, conventions. It is loaded every session, so keep it short.

Two things are yours to set. Tell the harness the model’s real context window, because it plans truncation and compaction from that number (the Cline setup shows where). And start a fresh session per task. Qwen3.8 27B’s 262,144-token window (vendor figure) is a capability, not a target: the model accepts that much, but every turn gets slower as the context grows, and models use information in the middle of a long context less reliably than at its start or end (Liu et al., 2023, “Lost in the Middle”, arXiv 2307.03172).

Why the model still matters

The harness decides what the model can do. The model decides what gets done. No harness can do these for it:

  • turn a vague bug report into the right search,
  • read a stack trace and fix the cause, not the line that threw,
  • make the smallest correct edit instead of rewriting the file,
  • keep the plan straight across dozens of steps and tens of thousands of tokens of history,
  • emit valid tool-call arguments every time — each malformed call costs a turn, and a run of them can end the session,
  • stop honestly: run the tests before claiming success, and never edit a test to make it pass.

Smaller models fail these first, typically by repeating the same search, editing the wrong file or declaring victory early. Agent benchmarks such as SWE-bench measure this whole loop — resolve a real GitHub issue so that the project’s tests pass — which makes them a better guide for agent work than single-answer coding tests. Read every such score together with the harness it was run in, because the harness is part of the result.

Why the harness changes results

Same model, different harness, different result. The SWE-agent paper isolated this: GPT-4 Turbo resolved 11.0% of SWE-bench Lite issues with a plain Linux shell and 18.0% with purpose-built commands for searching, viewing and editing files (published research: Yang et al., 2024, “SWE-agent”, arXiv 2405.15793). The model was the same; the interface changed.

The harness decidesWhy it matters
Tool set and tool descriptionsThe model can only do what is offered, and picks tools by their descriptions
Edit formatWhole file, search-and-replace or diff — models differ in which they apply reliably; aider sets a default format per model
Output limitsToo little hides the error, too much buries it
Repo map or indexHow quickly the model finds the right file
CompactionWhat gets forgotten when the window fills
Test runningWhether failures come back without the model asking
Approval and sandboxWhat can happen without you seeing it

So compare harnesses with the same model on tasks from your own repository, and compare models inside the same harness.

If you serve the model yourself

A harness that uses native tool calls needs the server to turn the model’s tool-call text back into structured tool_calls. Both common open runtimes do it. In vLLM, enable automatic tool choice and pick the parser for the model family (hermes for Qwen2.5, per vLLM’s tool-calling docs):

vllm serve Qwen/Qwen2.5-7B-Instruct --max-model-len 32768 --enable-auto-tool-choice --tool-call-parser hermes

llama.cpp’s llama-server parses tool calls through the model’s Jinja chat template. Current builds turn Jinja on by default; --jinja makes it explicit and older builds need it:

llama-server -m Qwen2.5-7B-Instruct-Q4_K_M.gguf -c 32768 -ngl 99 --jinja

The two lines need different hardware. vLLM loads the 16-bit weights: 7.61 billion parameters × 2 bytes = 15.2 GB (calculated from the published parameter count), more than one 12 GB RTX 3060 holds. A Q4_K_M GGUF of it is about 4.7 GB (vendor figure: Qwen’s own Q4_K_M files total 4.68 GB). Either way the KV cache comes on top: 2 (K and V) × 28 layers × 4 KV heads × 128 × 2 bytes = 57,344 bytes per token, × 32,768 tokens = 1.75 GiB (calculated from the model’s published config). Set the context (--max-model-len, -c) to what your sessions need and the memory allows, enter the same number in the harness, and check the fit with Model size, GPU memory and CPU offload. If a harness reports that the model does not support tools, or the model prints its tool call as plain text, check these flags first.

Worked example: one bug-fix loop

A refactor moved one line in an invoice total. With a subtotal of 100.00, a fixed discount of 10.00 and 25% VAT, the right total is (100.00 − 10.00) × 1.25 = 112.50. The code now computes 100.00 × 1.25 − 10.00 = 115.00. Your request: “tests/test_totals.py fails since the last refactor. Fix it.” An illustrative session; token counts are estimates.

#The model asks forWhat comes backContext added
1Search: rg -n "discount" src/billing9 matching lines in 3 files~250 tokens
2Read: src/billing/totals.py, lines 1–80The total() function~900
3Tests: pytest tests/test_totals.py -q1 failed: assert 115.0 == 112.5~600
4Git: git log -p -2 -- src/billing/totals.pyThe refactor moved the discount below the VAT line~1,200
5Edit: subtract the discount before VAT2 lines changed~200
6Tests: the same command1 failed: assert -12.5 == 0.0 (discount 20.00 on a subtotal of 10.00)~600
7Edit: max(subtotal - discount, 0)1 line changed~150
8Tests: pytest -q, the whole suiteAll pass~100
9Git: git diff3 changed lines~250
10No tool: a text answerCause, fix and diff, for you to review—

The tool traffic adds 250 + 900 + 600 + 1,200 + 200 + 600 + 150 + 100 + 250 = 4,250 tokens. With 6,000 tokens of harness prompt and tool definitions (estimate) the session ends near 10,250 tokens, about 4% of a 262,144-token window (10,250 / 262,144 = 3.9%).

What decided the outcome:

  • The model’s judgment: search before reading (step 1), and ask git why the line moved (step 4) instead of guessing.
  • The tests: the first edit dropped the clamp the old line carried (step 6). Running the tests caught it, not reading more code. Without a test command the session ends at step 5 with a wrong fix.
  • The harness: it cut each test run to its tail, asked before running commands, and put the diff in front of you.

The common mistake: pasting the whole repository

Pasting the entire codebase into the prompt feels thorough. It makes the result worse:

  • It often does not fit. A 300,000-line codebase at about 10 tokens per line (estimate; depends on language and tokenizer) is about 3,000,000 tokens, more than 11 times a 262,144-token window.
  • When it fits, it is resent every turn. A 200,000-token paste over 20 steps is at least 20 × 200,000 = 4,000,000 input tokens, against about 445,000 for the whole exploring session estimated above.
  • It costs memory when you serve the model. Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128. In FP16 its KV cache takes 2 (K and V) × 32 × 8 × 128 × 2 bytes = 131,072 bytes (128 KiB) per token, so its full 131,072-token context needs 131,072 × 128 KiB = 16 GiB (calculated from the model’s published config). That is more than a 12 GB RTX 3060 holds before a single weight is loaded.
  • It dilutes attention. The lines that matter sit somewhere in 200,000 tokens, most likely in the middle.
  • It goes stale. After the first edit the pasted copy is wrong; a read tool always sees the current file.

Give the agent what a senior colleague would want instead: the failing test or the error, where to start, the test command, and what not to touch. Put the durable parts in the instruction file.

On AxForge

Next

Ask on the forum Markdown For AI Updated 2026-10-01

Anything unclear on this page?

Ask on the forum — the answer helps the next person too.

Ask about this page