Models & pricing
One key serves every model below, from Stockholm, Sweden (eu-se-1). Pass the API model name in the
model field of the matching endpoint. Prices are per-use, in
EUR, excluding VAT. Every new account currently includes 5M serverless tokens/month at launch.
Available on the serverless API
| Model | Endpoint | API model name | Context | Price |
|---|---|---|---|---|
| Qwen3.8 27B | /v1/chat/completions | qwen3.8-27b-nvfp4 | 262,144 | €0.29 / 1M in · €1.77 / 1M out |
| Qwen3 Embedding | /v1/embeddings | qwen3-embed | 32,768 | €0.015 / 1M |
| ERNIE Image Turbo | /v1/images/generations | ernie-image-turbo | — | €0.20 / image |
| FLUX.2 Klein 4B | /v1/images/edits | flux2-klein-4b | — | €0.20 / image |
| Whisper | /v1/audio/transcriptions | whisper | — | €0.005 / minute |
| Piper | /v1/audio/speech | piper | — | €2.95 / 1M characters |
| MiniMax Music 3 | /v1/audio/music | minimax-music3 | — | €0.09 / track |
Chat prices are launch pricing. Each linked page carries the model's specs, benchmarks and data-handling details; the API pages in these docs (chat, embeddings, images, audio) carry the request and response shapes.
How we price
We price inference aggressively low: our target is 80% of the lowest current price from a proper inference supplier for the same model, tracked continuously. Found better pricing at a proper inference supplier? Tell us and we'll look into lowering ours.
Where the tracked market has no comparable listing — Whisper, Piper, image editing, music — prices are launch prices anchored to the closest public list price. Committed volume: talk to an engineer.
Available as managed deployment
These models are validated on AxForge hardware and deployed on a dedicated NVIDIA DGX Spark for your traffic only — an OpenAI-compatible endpoint on your own machine, operated by AxForge. Hardware from €0.55/hour; managed service quoted per deployment.
| Model | Type | Context |
|---|---|---|
| Qwen3.6 35B A3B | LLM (MoE, 3B active) | 65,536 |
| Qwen3 30B A3B | LLM (MoE) | 65,536 |
| Gemma-4 26B | LLM, vision-capable | 262,144 |
| Mistral Small 3.2 | LLM | 32,768 |
| SDXL | Image generation | — |
| Qwen-Image | Image generation & editing | — |
Want one of these, or a different open model on dedicated hardware? Request deployment — or rent the DGX Spark yourself with full SSH access.