# Overview


# AI Engineering

Understand the parts that change when software includes AI — models, GPU memory, quantization, context, inference runtimes, agents and retrieval — then try them on AxForge.

## What is here

**Models & Memory**What a model needs in memory (weights, KV cache, runtime overhead, working memory) and how quantization and context length change it.
**GPUs**What one card, two cards or unified memory changes, and what it does not.
**Inference**How a request becomes tokens, and how runtimes such as llama.cpp and vLLM schedule the work.
**Coding & Agents**How coding agents read, edit and test code, and how they work on very large repositories.
**AI Architecture**Retrieval, decision models and system design. Guides are being written; the Embeddings and Xev docs are there now.

## Where each part sits

```
 your application                    AI Architecture, Coding & Agents
 (agent loop, retrieval, tools)
        |
        |  HTTP, OpenAI-compatible API
        v
 inference runtime                   Inference
 (llama.cpp, vLLM, ...)
        |
        v
 model weights + KV cache            Models & Memory
        |
        v
 GPU memory and compute              GPUs
 (CPU RAM if offloaded: much slower)
```

The lower three layers decide whether a model runs at all and how fast: whether it fits in memory, how much memory the context takes, and how the runtime schedules requests. The guides take each layer in turn.

## All guides

| Models & Memory | Model size, VRAM and CPU offload · Quantization: FP16, FP8, NVFP4, Q4 · Context and the KV cache |  |

| GPUs | One GPU or several · RTX 3060 · DGX Spark |  |

| Inference | How inference fits together · llama.cpp or vLLM? |  |

| Coding & Agents | How coding agents work · AI on a 5-million-line codebase |  |

| AI Architecture | Coming: retrieval, decision models, system design. Now: Embeddings · Xev decisions |  |

## Common questions

- Why is my model using the CPU?

- Will this model fit on my GPU?

- FP16 vs FP8 vs NVFP4 vs Q4

- How much context do I need?

- 1 GPU vs 2 GPUs

- llama.cpp or vLLM?

- How do coding agents work?

- How can AI work with a 5-million-line codebase?

## On AxForge

- Quickstart: one call to the API at `https://api.axforge.ai/v1`

- Model catalogue: language, embedding and image models

- GPUs and how rentals work: run a runtime yourself

- Connect a coding tool: aider, Cline, Continue, Codex CLI and Claude Code

## Next

- How inference fits together: the map the other guides build on

- Model size, VRAM and CPU offload: the question most people arrive with

- How coding agents work: if you came here for the tools



Source: https://dev.axforge.ai/ai-engineering/
