Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job — personal learning notes, no official statements. Verify critical steps yourself.

Selecting the Right GPU for Qwen3 Inference

A practical playbook for LLM inference on GPUs: roofline thinking, bandwidth vs. compute, precision and utilization by model size.

6 min readflozi00
aidatacentergpudeep-learninghardwareguide

This overview extends the calculations from LLM Inference Math: From Theory to Hardware and applies them to concrete hardware + model pairings. The goal is to make it obvious when the NVIDIA RTX PRO 6000, H200, or the shipping DGX Station delivers the best efficiency for Qwen3-class workloads ranging from 4B to 32B active parameters.

Important: Every number in this playbook is a theoretical ceiling derived from vendor specs and simplified roofline math. Real deployments often land lower because kernels are imperfect, host↔device pipelines add friction, and GPUs rarely sustain 100% efficiency across an entire decode pass.

Recap: the 60-second ops:byte checklist

  1. Compute the GPU's ops:byte ratio (peak FLOPS ÷ memory bandwidth).
  2. Compute the model's arithmetic intensity (≈ d_head ÷ 2 for attention-heavy steps).
  3. If model_intensity < gpu_intensity, the workload is memory-bound; focus on bandwidth and VRAM.
  4. Estimate time per output token with model_size_bytes ÷ memory_bandwidth_bytes_per_second to derive a theoretical tokens/s ceiling for batch size 1.

GPU capability snapshots

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

  • 96 GB GDDR7 ECC VRAM fed by 1,792 GB/s bandwidth and 503.8 TFLOPS FP16/BF16 Tensor (dense) compute — the ratio that matters for inference — which yields an ops:byte ratio near 281. (The ~125 TFLOPS non-tensor FP32 shader rate would misleadingly suggest ~70; see the footnote in the inference math guide.)
  • Peak board power of 600 W enables deskside deployments where noise + thermals matter but rack power is limited.
  • Ideal when you need local fine-tuning or multi-modal prototyping (vision, audio) with up to ~64 GB of weights plus a useful KV cache budget.1

NVIDIA H200 Tensor Core GPU (SXM + NVL)

  • First Hopper-based accelerator with 141 GB of HBM3e and 4.8 TB/s of memory bandwidth; BF16 Tensor performance reaches 989 TFLOPS dense (1,979 TFLOPS with sparsity), so ops:byte climbs to roughly 206 (dense) / 412 (sparse).
  • Ships with hardware MIG slicing (7 instances) and optional NVL configurations for air-cooled racks.
  • Best suited for 14B+ dense models or MoE deployments where both capacity and streaming bandwidth dominate cost.2

NVIDIA DGX Station (Grace Blackwell Ultra)

  • Desktop supercomputer that combines one Blackwell-Ultra GPU (252 GB HBM3e @ 7.1 TB/s per the shipping DGX Station datasheet) with a 72-core Grace CPU and 496 GB of LPDDR5X in a coherent 748 GB memory pool.
  • NVLink-C2C delivers 900 GB/s between CPU and GPU, so large retrieval datasets can stay resident without PCIe penalties.
  • The datasheet rates the GPU at 20 PFLOPS FP4 / 5 PFLOPS BF16 (sparse; 15 / 2.5 dense) — the same Blackwell Ultra silicon as HGX B300 servers, whose 8-GPU baseboard NVIDIA rates at 18 PFLOPS dense BF16 (~2.25 PFLOPS per GPU).34
  • Targets multi-user labs that need on-prem autonomy for iterative training, MoE routing experiments, and agent stacks before shipping them to a cluster.

Qwen3 model footprints (batch size 1, BF16)

Small, well-scoped translation chunks keep the translation pipeline happy. Each model reuses the GQA-aware KV-cache formula (2 * layers * num_kv_heads * head_dim * 2 bytes; all Qwen3 models below use 8 KV heads with head_dim 128).

ModelActive paramsHidden size / layersWeights (GB)KV cache per token (MB)Notes
Qwen3-4B-Instruct~4B2,560 / 36~80.15Sliding-window ready; great for CPU offload experiments.5
Qwen3-VL-8B-Instruct~8B4,096 / 36~160.15Vision-language encoder adds ~1152-dim vision tower.6
Qwen3-14B~14B5,120 / 40~280.1640-layer stack with 1M rope theta for 40k context.7
Qwen3-32B~32B5,120 / 64~640.2664 decoder layers; same d_head so arithmetic intensity stays ≈64 ops/byte.8

Matching scenarios

1. Workstation prototyping (RTX PRO 6000)

  • Recommended models: Qwen3-4B, Qwen3-VL-8B.
  • Why: Both weights (8–16 GB) plus KV cache for 4K tokens still leave >70 GB VRAM for batching, LoRA adapters, or vision embeddings.
  • Throughput: Theoretical tokens/s = 1.792 TB/s ÷ weights. Expect ~220 tok/s (4B) or ~110 tok/s (8B) before compute saturation, so latency is dominated by memory fetch, not tensor ops.
  • Tip: Stay compute-balanced by pushing batch size to 4 whenever latency budget allows; OBS or Whisper sidecars barely dent VRAM.

2. Enterprise copilots (H200 SXM)

  • Recommended models: Qwen3-14B dense, Qwen3-32B MoE routing with batch 1–2.
  • Why: 141 GB HBM3e accommodates the 64 GB model plus >70 GB for long-context KV caches. The ~206 ops:byte ratio (dense BF16) means arithmetic intensity (64) keeps you memory-bound, so the 4.8 TB/s feed confers ~170 tok/s on 14B and ~75 tok/s on 32B without tensor parallelism.
  • Tip: Split attention + feed-forward layers across MIG instances when serving multiple tenants; each MIG slice still gets ≥18 GB.

3. Lab-scale supercomputer (DGX Station)

  • Recommended models: Any Qwen3 variant plus stacked tools (RAG, VLM agents) thanks to the 748 GB coherent pool.
  • Why: 252 GB of on-package HBM3e means you can pin two dense models or a dense+MoE pair simultaneously while the Grace CPU handles data preprocessing at 396 GB/s. NVLink-C2C eliminates PCIe resharding when streaming documents from RAM into KV caches.
  • Throughput: At 7.1 TB/s feeding 64 GB of BF16 weights, the roofline stays memory-bound: plan for ~111 tok/s on a 32B dense model (7.1 TB/s ÷ 64 GB) and scale with batch size until the compute ceiling (~2.5 PFLOPS dense BF16 ÷ 2·active params) is approached.4
  • Tip: Use MIG (7 slices) to dedicate small partitions to telemetry or guardrail models without interrupting the main VLM job.

Quick pairing matrix

ScenarioModelGPUEst. tokens/s (batch 1)Primary bottleneckNotes
Edge copilotsQwen3-4BRTX PRO 6000~220Memory BWPlenty of VRAM left for RAG embeddings.
Vision agent demosQwen3-VL-8BRTX PRO 6000~110Memory BWVision tower benefits from 96 GB VRAM for image batches.
Customer support copilotsQwen3-14BH200~170Memory BWMIG lets you mirror-prod topology in dev.
Technical assistant / codegenQwen3-32BH200~75Memory BWRequires tensor parallel if batching >2.
Multi-agent sandboxQwen3-32B + toolsDGX Station~111Memory BW748 GB pool hosts RAG corpora in-memory.

Interactive calculator

Use the LLM inference planner to stress-test context windows, batch sizes, and precision assumptions against each GPU profile — it computes VRAM fit, decode throughput, prefill latency and the break-even batch B* live for every model and GPU in this playbook.

Note: Tokens per second plateau once the workload hits the compute roofline—the calculator compares both limits and reports the stricter one.

Footnotes

  1. NVIDIA RTX PRO 6000 Blackwell Workstation Edition specifications, NVIDIA. ↩

  2. NVIDIA H200 Tensor Core GPU specifications, NVIDIA. ↩

  3. NVIDIA DGX Station (Grace Blackwell Ultra) specifications, NVIDIA. ↩

  4. NVIDIA HGX Platform and Blackwell Ultra specifications (HGX B300), NVIDIA. ↩ ↩2

  5. Qwen/Qwen3-4B-Instruct-2507 model card, Hugging Face. ↩

  6. Qwen/Qwen3-VL-8B-Instruct model card, Hugging Face. ↩

  7. Qwen/Qwen3-14B model card, Hugging Face. ↩

  8. Qwen/Qwen3-32B model card, Hugging Face. ↩