What should I use for vision models with lots of images?
What should I use for vision models with lots of images?
What should I use for vision models with lots of images?
That depends on the vision model and how many images you process together.
If the workload fits in 32 GB, RTX 5090 is attractive because of its speed. If you need a large language model, vision encoder and large context resident together, GB10 gives you much more memory headroom.
For a production recommendation we would need the exact model, image sizes, images per request and target concurrency.
Sign in to reply.
Posting guidelines
Accounts that break these can lose forum access — paid plans included. Read the full guidelines →
Related topics