GPU Platform: NVIDIA GB300
NVIDIA GB300: Rack-Scale Grace Blackwell for the Largest Clusters
GB300 pairs the Grace CPU with the Blackwell Ultra GPU in NVIDIA's NVL72 rack-scale architecture — 72 GPUs and 36 Grace CPUs acting as one NVLink-coherent domain per rack. It's the platform for teams pushing the largest training runs and highest-throughput inference, where rack-scale memory and bandwidth are the constraint.
- GPUs / Rack
- 72 (NVL72)
- Cooling
- Direct liquid
- Architecture
- Grace CPU + Blackwell Ultra GPU
- Domain
- Rack-scale NVLink
What the GB300 platform is built for
NVL72's entire premise is treating a full rack as one coherent accelerator — every GPU sees every other GPU's memory over NVLink at rack scale, rather than being bottlenecked at the 8-GPU node boundary. That matters most for frontier-scale training and for inference workloads serving the largest models with the lowest latency.
It's a liquid-cooled system, which means it needs a site built or retrofitted for direct-liquid-cooling — one of the specific things we evaluate when matching a GB300 deployment to a powered site.
What we can get you on GB300
Same two paths as our other platforms, sized to rack-scale deployments:
- GPUs-as-a-Service — rack-scale capacity on infrastructure we've sourced, financed, and already built for direct-liquid-cooling
- Dedicated deployment — a GB300 buildout on a powered site matched and financed for your specific scale, if you're deploying multiple racks or need a dedicated footprint
Why liquid cooling changes the site conversation
Not every powered site can take a GB300 deployment without retrofit — direct-liquid-cooling plant, CDUs, and piping are real infrastructure asks. This is exactly the kind of site-viability question we run before matching a deployment, not something discovered mid-buildout.
GB300 vs. B300: which one you actually need
If your training runs are bottlenecked by rack-scale memory and bandwidth rather than raw GPU count — the largest frontier-scale training jobs and the highest-throughput, lowest-latency inference for the biggest models — NVL72's coherent 72-GPU domain is built specifically for that problem. It's also the platform to reach for if you're deploying into a facility already built or being built for direct liquid cooling.
If your workloads fit within an 8-GPU NVLink domain and your site runs conventional air-cooled infrastructure, B300 gets you the same Blackwell Ultra generation without the cooling-plant requirement — see the B300 page for the fuller comparison.
What a typical timeline looks like
Rack-scale deployments carry more site-side variables than an 8-GPU node — cooling infrastructure chief among them — so timelines depend more heavily on whether the target site already has direct-liquid-cooling capability or needs it built. GPUs-as-a-Service capacity on infrastructure we've already built for liquid cooling is the faster path when your timeline is tight; a dedicated buildout on a matched site follows our standard qualify-assess-match-execute process.
What we need from you to scope it
Target scale (racks or GPU count), deployment type, timeline, and region preference. If you already know your cooling and power requirements, tell us — it speeds up matching considerably.