Models library

A library of open models, each matched to the right GPU and ready to deploy on dedicated compute — pricing on demand.

1General chat & business AI

ModelBest useBest GPUOther suitable GPUs
Qwen3-8BFast chatbots, lightweight agents and prototypingIntel Arc Pro B70RTX 5090, RTX PRO 6000
Qwen3-14BBalanced chat quality and operating costRTX 5090B70 quantized, RTX PRO 6000
Qwen3-30B-A3B-InstructBusiness assistants, agents and multilingual applicationsRTX 5090B70 quantized, RTX PRO 6000
Qwen3-32BHigher-quality business AI, RAG and multilingual workRTX PRO 6000RTX 5090 quantized, B70 quantized/experimental
Qwen3-235B-A22B-InstructPremium enterprise and high-quality agent workloadsMulti-GPU RTX PRO 6000Not recommended on one 5090 or B70

Qwen3 includes dense and mixture-of-experts models and supports reasoning and normal response modes. The Qwen3 family is publicly released under Apache 2.0.

Best launch choices: Qwen3-8B, Qwen3-30B-A3B and Qwen3-32B.

2Reasoning models

ModelBest useBest GPUOther suitable GPUs
DeepSeek-R1-Distill-Qwen-14BAffordable reasoning, math and researchRTX 5090B70 quantized, RTX PRO 6000
DeepSeek-R1-Distill-Qwen-32BStrong reasoning and mathematical workRTX PRO 6000RTX 5090 quantized, B70 experimental
DeepSeek-R1-Distill-Llama-70BMore demanding reasoning workloadsRTX PRO 6000Not ideal on one 5090 or B70
Qwen3-32B reasoning modeCombined business chat and reasoningRTX PRO 6000RTX 5090 quantized
Full DeepSeek-R1Large-scale reasoning and synthetic-data generationMulti-GPU RTX PRO 6000Not recommended on one card

Best launch choice: DeepSeek-R1-Distill-Qwen-32B.

The 14B model is better for affordable single-GPU jobs. The 32B version is better positioned as the standard reasoning product, while the 70B version should be a premium RTX PRO 6000 deployment.

3Coding & software agents

ModelBest useBest GPUOther suitable GPUs
Qwen3-Coder-30B-A3B-InstructCoding assistants, tool use and software agentsRTX 5090RTX PRO 6000, B70 experimental
Qwen3-Coder-30B-A3B-Instruct FP8Faster coding inference and longer contextsRTX 5090RTX PRO 6000
Qwen3-Coder-NextAutonomous coding agents and local developmentRTX PRO 6000RTX 5090 quantized
Qwen2.5-Coder-32B-InstructStable coding APIs and existing integrationsRTX PRO 6000RTX 5090 quantized, B70 experimental
DeepSeek-Coder-V2-Lite-InstructFaster code completion and lower-cost assistantsIntel Arc Pro B70RTX 5090, RTX PRO 6000
Qwen3-Coder-480B-A35B-InstructHigh-end autonomous coding agentsMulti-GPU RTX PRO 6000Not recommended on one card

Qwen3-Coder-30B-A3B is the practical single-GPU version, whereas the 480B model is intended for much larger deployments. Qwen3-Coder-Next is an 80B-parameter mixture-of-experts model activating approximately 3B parameters during inference, but its total model memory still makes the RTX PRO 6000 the safer deployment choice.

Best launch choice: Qwen3-Coder-30B-A3B-Instruct FP8.

4Vision & document understanding

ModelBest useBest GPUOther suitable GPUs
Qwen3-VL-8B-InstructImages, PDFs, charts, screenshots and OCR-like analysisRTX 5090B70 experimental, RTX PRO 6000
Qwen3-VL-8B-Instruct FP8High-throughput visual and document processingRTX 5090RTX PRO 6000
Qwen3-VL-30B-A3B-InstructAdvanced visual reasoning and long documentsRTX PRO 6000RTX 5090 quantized
Qwen3-VL-32BHigh-quality image, document and video analysisRTX PRO 6000RTX 5090 quantized
Qwen3-VL-235B-A22B-InstructLarge-scale multimodal and enterprise workflowsMulti-GPU RTX PRO 6000Not recommended on one card

Qwen3-VL supports text, images and video, with dense and mixture-of-experts variants. The model family supports long multimodal contexts, making it useful for document collections and video analysis.

Best launch choices: Qwen3-VL-8B FP8 and Qwen3-VL-30B-A3B.

5Image generation

ModelBest useBest GPUOther suitable GPUs
Stable Diffusion XL 1.0General image generation, LoRAs and ControlNetRTX 5090B70 with supported software, RTX PRO 6000
SDXL TurboFast previews and high-volume image generationRTX 5090B70 with supported software, RTX PRO 6000
PixArt-SigmaHigh-resolution text-to-image generationRTX 5090RTX PRO 6000, B70 experimental
Customer-supplied SD/SDXL checkpointCustom styles, LoRAs and production workflowsRTX 5090RTX PRO 6000
Stable Diffusion 3.5 MediumModern image generation with moderate requirementsRTX 5090RTX PRO 6000
Stable Diffusion 3.5 LargeMore demanding image generation and typographyRTX PRO 6000RTX 5090 with optimized workflow

Best hardware: RTX 5090. Image-generation workloads usually do not need 96 GB. The RTX 5090 gives customers high performance without charging them for unused RTX PRO 6000 memory. Use the RTX PRO 6000 for large batches, multiple simultaneous workflows or unusually large pipelines. Offer these through ComfyUI, including persistent storage for customer checkpoints, LoRAs, ControlNets and workflows.

6Video generation

ModelBest useBest GPUOther suitable GPUs
Wan2.2 TI2V-5BEfficient text-to-video and image-to-videoRTX 5090RTX PRO 6000, B70 experimental
Wan2.1 T2V-14BHigher-quality text-to-videoRTX PRO 6000RTX 5090 with optimization
Wan2.1 I2V-14B 720pImage animation, advertising and creative productionRTX PRO 6000RTX 5090 with optimization
Wan2.2 Animate-14BCharacter animation and motion transferRTX PRO 6000RTX 5090 with optimized workflow

Best launch choice: Wan2.2 TI2V-5B on RTX 5090.

Use the RTX PRO 6000 for longer videos, higher resolutions, heavier ComfyUI graphs, larger batches or concurrent users. The Intel B70 may eventually be viable, but video-generation pipelines commonly depend on CUDA-specific components, so it should not be your default B70 product until fully validated.

7Speech & transcription

ModelBest useBest GPUOther suitable GPUs
Whisper Large V3 TurboFast multilingual transcriptionIntel Arc Pro B70RTX 5090, RTX PRO 6000
Whisper Large V3Maximum transcription accuracyIntel Arc Pro B70RTX 5090, RTX PRO 6000
French Whisper fine-tuneFrench calls, interviews and mediaIntel Arc Pro B70RTX 5090, RTX PRO 6000
Batched Whisper endpointHigh-volume transcription serviceIntel Arc Pro B70RTX 5090

Best hardware: Intel Arc Pro B70. Whisper does not require 96 GB of VRAM. A B70 is a better economic match, provided your selected inference runtime is stable on Intel. The RTX 5090 is the safer fallback for maximum software compatibility. For shared transcription, sell by audio minute rather than reserving an entire GPU.

8Embeddings & RAG

ModelBest useBest GPUOther suitable GPUs
BGE-M3Multilingual semantic search and RAGIntel Arc Pro B70RTX 5090, RTX PRO 6000
BGE Reranker V2 M3Improving retrieved-result qualityIntel Arc Pro B70RTX 5090, RTX PRO 6000
Nomic Embed Text V1.5Document embeddings and retrievalIntel Arc Pro B70RTX 5090, RTX PRO 6000
BGE Small EN V1.5Fast English embeddingsIntel Arc Pro B70Any of the three
Don't see your model? We can deploy most open-weight models on request. Ask us about a specific model and we'll scope the GPU and pricing for your workload.