AI Engineering
Understand the parts that change when software includes AI — models, GPU memory, quantization, context, inference runtimes, agents and retrieval — then try them on AxForge.
What is here
Models & MemoryWhat a model needs in memory (weights, KV cache, runtime overhead, working memory) and how quantization and context length change it.
GPUsWhat one card, two cards or unified memory changes, and what it does not.
InferenceHow a request becomes tokens, and how runtimes such as llama.cpp and vLLM schedule the work.
Coding & AgentsHow coding agents read, edit and test code, and how they work on very large repositories.
AI ArchitectureRetrieval, decision models and system design. Guides are being written; the Embeddings and Xev docs are there now.
Where each part sits
your application AI Architecture, Coding & Agents
(agent loop, retrieval, tools)
|
| HTTP, OpenAI-compatible API
v
inference runtime Inference
(llama.cpp, vLLM, ...)
|
v
model weights + KV cache Models & Memory
|
v
GPU memory and compute GPUs
(CPU RAM if offloaded: much slower)
The lower three layers decide whether a model runs at all and how fast: whether it fits in memory, how much memory the context takes, and how the runtime schedules requests. The guides take each layer in turn.
All guides
| Models & Memory | Model size, VRAM and CPU offload · Quantization: FP16, FP8, NVFP4, Q4 · Context and the KV cache |
|---|---|
| GPUs | One GPU or several · RTX 3060 · DGX Spark |
| Inference | How inference fits together · llama.cpp or vLLM? |
| Coding & Agents | How coding agents work · AI on a 5-million-line codebase |
| AI Architecture | Coming: retrieval, decision models, system design. Now: Embeddings · Xev decisions |
Common questions
- Why is my model using the CPU?
- Will this model fit on my GPU?
- FP16 vs FP8 vs NVFP4 vs Q4
- How much context do I need?
- 1 GPU vs 2 GPUs
- llama.cpp or vLLM?
- How do coding agents work?
- How can AI work with a 5-million-line codebase?
On AxForge
- Quickstart: one call to the API at
https://api.axforge.ai/v1 - Model catalogue: language, embedding and image models
- GPUs and how rentals work: run a runtime yourself
- Connect a coding tool: aider, Cline, Continue, Codex CLI and Claude Code
Next
- How inference fits together: the map the other guides build on
- Model size, VRAM and CPU offload: the question most people arrive with
- How coding agents work: if you came here for the tools