We wrote an AI Guide π
We just added an AI Guide to the developer area.
When you start running models, the same questions seem to come up.
Will this model or that model, fit on my GPU?
Why is it suddenly running on the CPU?
llama.cpp or vLLM?
Weird, i thought i already updated LMStudio?
The answers are out there. They're just scattered across model cards, GitHub issues, Reddit threads and somebody's blog post from last year or in someones head. Actually, i bet a lot of stuff is in some ppls head.
So we wrote them down, in one place.
The first eight guides:
- Model size, GPU memory and CPU offload
- Quantization: BF16, FP8, NVFP4, INT4 and GGUF
- Context windows and the KV cache
- One GPU vs multiple GPUs
- How AI inference works
- llama.cpp vs vLLM
- How coding agents work
- AI on a 5-million-line codebase
Each guide explains the idea in general first, then shows where you can try it on AxForge.
Retrieval, decision models and system design are next.
If something is unclear, wrong or missing, every guide has an Ask about this page button at the bottom. Use it. That's how the next guide gets written. But you can also just join in on the forum and write something right here.
It's a bit different from the regular help pages and documentation. And i feel it's a good fit because i get a lot of questions like this. And also, this is the first one so a lot of it is summarized things by AI that i've written before and we're still working a bit on the format. So this is a first version of these pages.
Eight guides in. Plenty more to write. Let's see what you ask first.