AI Engineering / Start

AI Engineering

Understand the parts that change when software includes AI — models, GPU memory, quantization, context, inference runtimes, agents and retrieval — then try them on AxForge.

What is here

Where each part sits

 your application                    AI Architecture, Coding & Agents
 (agent loop, retrieval, tools)
        |
        |  HTTP, OpenAI-compatible API
        v
 inference runtime                   Inference
 (llama.cpp, vLLM, ...)
        |
        v
 model weights + KV cache            Models & Memory
        |
        v
 GPU memory and compute              GPUs
 (CPU RAM if offloaded: much slower)

The lower three layers decide whether a model runs at all and how fast: whether it fits in memory, how much memory the context takes, and how the runtime schedules requests. The guides take each layer in turn.

All guides

Common questions

On AxForge

Next

Ask on the forum Markdown For AI Updated 2026-10-01

Anything unclear on this page?

Ask on the forum — the answer helps the next person too.

Ask about this page