A library of open models, each matched to the right GPU and ready to deploy on dedicated compute — pricing on demand.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| Qwen3-8B | Fast chatbots, lightweight agents and prototyping | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| Qwen3-14B | Balanced chat quality and operating cost | RTX 5090 | B70 quantized, RTX PRO 6000 |
| Qwen3-30B-A3B-Instruct | Business assistants, agents and multilingual applications | RTX 5090 | B70 quantized, RTX PRO 6000 |
| Qwen3-32B | Higher-quality business AI, RAG and multilingual work | RTX PRO 6000 | RTX 5090 quantized, B70 quantized/experimental |
| Qwen3-235B-A22B-Instruct | Premium enterprise and high-quality agent workloads | Multi-GPU RTX PRO 6000 | Not recommended on one 5090 or B70 |
Qwen3 includes dense and mixture-of-experts models and supports reasoning and normal response modes. The Qwen3 family is publicly released under Apache 2.0.
Best launch choices: Qwen3-8B, Qwen3-30B-A3B and Qwen3-32B.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| DeepSeek-R1-Distill-Qwen-14B | Affordable reasoning, math and research | RTX 5090 | B70 quantized, RTX PRO 6000 |
| DeepSeek-R1-Distill-Qwen-32B | Strong reasoning and mathematical work | RTX PRO 6000 | RTX 5090 quantized, B70 experimental |
| DeepSeek-R1-Distill-Llama-70B | More demanding reasoning workloads | RTX PRO 6000 | Not ideal on one 5090 or B70 |
| Qwen3-32B reasoning mode | Combined business chat and reasoning | RTX PRO 6000 | RTX 5090 quantized |
| Full DeepSeek-R1 | Large-scale reasoning and synthetic-data generation | Multi-GPU RTX PRO 6000 | Not recommended on one card |
Best launch choice: DeepSeek-R1-Distill-Qwen-32B.
The 14B model is better for affordable single-GPU jobs. The 32B version is better positioned as the standard reasoning product, while the 70B version should be a premium RTX PRO 6000 deployment.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| Qwen3-Coder-30B-A3B-Instruct | Coding assistants, tool use and software agents | RTX 5090 | RTX PRO 6000, B70 experimental |
| Qwen3-Coder-30B-A3B-Instruct FP8 | Faster coding inference and longer contexts | RTX 5090 | RTX PRO 6000 |
| Qwen3-Coder-Next | Autonomous coding agents and local development | RTX PRO 6000 | RTX 5090 quantized |
| Qwen2.5-Coder-32B-Instruct | Stable coding APIs and existing integrations | RTX PRO 6000 | RTX 5090 quantized, B70 experimental |
| DeepSeek-Coder-V2-Lite-Instruct | Faster code completion and lower-cost assistants | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| Qwen3-Coder-480B-A35B-Instruct | High-end autonomous coding agents | Multi-GPU RTX PRO 6000 | Not recommended on one card |
Qwen3-Coder-30B-A3B is the practical single-GPU version, whereas the 480B model is intended for much larger deployments. Qwen3-Coder-Next is an 80B-parameter mixture-of-experts model activating approximately 3B parameters during inference, but its total model memory still makes the RTX PRO 6000 the safer deployment choice.
Best launch choice: Qwen3-Coder-30B-A3B-Instruct FP8.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| Qwen3-VL-8B-Instruct | Images, PDFs, charts, screenshots and OCR-like analysis | RTX 5090 | B70 experimental, RTX PRO 6000 |
| Qwen3-VL-8B-Instruct FP8 | High-throughput visual and document processing | RTX 5090 | RTX PRO 6000 |
| Qwen3-VL-30B-A3B-Instruct | Advanced visual reasoning and long documents | RTX PRO 6000 | RTX 5090 quantized |
| Qwen3-VL-32B | High-quality image, document and video analysis | RTX PRO 6000 | RTX 5090 quantized |
| Qwen3-VL-235B-A22B-Instruct | Large-scale multimodal and enterprise workflows | Multi-GPU RTX PRO 6000 | Not recommended on one card |
Qwen3-VL supports text, images and video, with dense and mixture-of-experts variants. The model family supports long multimodal contexts, making it useful for document collections and video analysis.
Best launch choices: Qwen3-VL-8B FP8 and Qwen3-VL-30B-A3B.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| Stable Diffusion XL 1.0 | General image generation, LoRAs and ControlNet | RTX 5090 | B70 with supported software, RTX PRO 6000 |
| SDXL Turbo | Fast previews and high-volume image generation | RTX 5090 | B70 with supported software, RTX PRO 6000 |
| PixArt-Sigma | High-resolution text-to-image generation | RTX 5090 | RTX PRO 6000, B70 experimental |
| Customer-supplied SD/SDXL checkpoint | Custom styles, LoRAs and production workflows | RTX 5090 | RTX PRO 6000 |
| Stable Diffusion 3.5 Medium | Modern image generation with moderate requirements | RTX 5090 | RTX PRO 6000 |
| Stable Diffusion 3.5 Large | More demanding image generation and typography | RTX PRO 6000 | RTX 5090 with optimized workflow |
Best hardware: RTX 5090. Image-generation workloads usually do not need 96 GB. The RTX 5090 gives customers high performance without charging them for unused RTX PRO 6000 memory. Use the RTX PRO 6000 for large batches, multiple simultaneous workflows or unusually large pipelines. Offer these through ComfyUI, including persistent storage for customer checkpoints, LoRAs, ControlNets and workflows.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| Wan2.2 TI2V-5B | Efficient text-to-video and image-to-video | RTX 5090 | RTX PRO 6000, B70 experimental |
| Wan2.1 T2V-14B | Higher-quality text-to-video | RTX PRO 6000 | RTX 5090 with optimization |
| Wan2.1 I2V-14B 720p | Image animation, advertising and creative production | RTX PRO 6000 | RTX 5090 with optimization |
| Wan2.2 Animate-14B | Character animation and motion transfer | RTX PRO 6000 | RTX 5090 with optimized workflow |
Best launch choice: Wan2.2 TI2V-5B on RTX 5090.
Use the RTX PRO 6000 for longer videos, higher resolutions, heavier ComfyUI graphs, larger batches or concurrent users. The Intel B70 may eventually be viable, but video-generation pipelines commonly depend on CUDA-specific components, so it should not be your default B70 product until fully validated.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| Whisper Large V3 Turbo | Fast multilingual transcription | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| Whisper Large V3 | Maximum transcription accuracy | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| French Whisper fine-tune | French calls, interviews and media | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| Batched Whisper endpoint | High-volume transcription service | Intel Arc Pro B70 | RTX 5090 |
Best hardware: Intel Arc Pro B70. Whisper does not require 96 GB of VRAM. A B70 is a better economic match, provided your selected inference runtime is stable on Intel. The RTX 5090 is the safer fallback for maximum software compatibility. For shared transcription, sell by audio minute rather than reserving an entire GPU.
| Model | Best use | Best GPU | Other suitable GPUs |
|---|---|---|---|
| BGE-M3 | Multilingual semantic search and RAG | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| BGE Reranker V2 M3 | Improving retrieved-result quality | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| Nomic Embed Text V1.5 | Document embeddings and retrieval | Intel Arc Pro B70 | RTX 5090, RTX PRO 6000 |
| BGE Small EN V1.5 | Fast English embeddings | Intel Arc Pro B70 | Any of the three |