Engineered on NVIDIA’s cutting-edge Blackwell architecture and coupled with the Arm-based Grace CPU, the GB200 NVL72 sets a new standard for AI compute.
It delivers up to 30x faster real-time LLM inference, cuts Total Cost of Ownership (TCO) by 25x, and consumes 25x less energy compared to previous generations.
Pricing details will be announced upon official commercial availability.
The NVIDIA DGX B300 serves as a unified AI supercomputing platform designed to empower developers and data science teams.
By dramatically accelerating time-to-insight, it helps organizations unlock the full strategic and economic potential of enterprise AI.
Engineered for massive-scale AI workloads, the NVIDIA GB200 NVL72 effortlessly processes datasets containing trillions of parameters.
Combining the state-of-the-art Blackwell architecture with the high-performance Arm Grace CPU, this advanced superchip delivers unprecedented computing power while slashing both operational costs (TCO) and energy consumption by up to 25x relative to prior-generation architectures.
NVIDIA: LLM inference and energy efficiency: Time to First Token (TTFT) = 50 ms real-time, Full-Time Latency (FTL) = 5 s, with 32,768 input and 1,024 output tokens.
Benchmark compares NVIDIA HGX™ H100 scaled over InfiniBand (IB) against the GB200 NVL72; 1.8T MoE model training evaluates 4,096x HGX H100 scaled via IB versus 456x GB200 NVL72 scaled via IB within a 32,768-cluster size.
Database join and aggregation performance with Snappy/Deflate compression is derived from the TPC-H Q4 query, featuring custom query implementations across x86, a single H100 GPU, and a single GPU from the GB200 NVL72 vs.
Intel Xeon 8480+.
Projected performance subject to change.
| GB200 NVL72 | GB200 Superchip | |
|---|---|---|
| Configuration | 36x Grace CPU, 72x B200 GPU | 1x Grace CPU, 2x B200 GPU |
| FP4 Tensor Core* | 1,440 PFLOPS | 40 PFLOPS |
| FP8 / FP6 Tensor Core* | 720 PFLOPS | 20 PFLOPS |
| INT8 Tensor Core* | 720 POPS | 20 POPS |
| FP16 / BF16 Tensor Core* | 360 PFLOPS | 10 PFLOPS |
| TF32 Tensor Core* | 180 PFLOPS | 5 PFLOPS |
| FP64 Tensor Core | 3,240 TFLOPS | 90 TFLOPS |
| GPU Memory | Up to 13.5 TB HBM3e, 576 TBps | Up to 384 GB HBM3e, 16 TBps |
| NVLink Bandwidth | 130 TBps | 3.6 TBps |
| CPU Cores | 2,952 Arm Neoverse V2 Cores | 72 Arm Neoverse V2 Cores |
| CPU Memory | Up to 17 TB LPDDR5X, Up to 18.4 TBps | Up to 480 GB LPDDR5X, Up to 18.4 TB/s |
| Product info | Datasheet | |
Designed for ultra-high-throughput environments, the NVIDIA GB200 NVL72 supports next-generation networking capabilities with bandwidth reaching up to 800 Gb/s.
To maximize AI throughput and eliminate performance bottlenecks, it integrates seamlessly with the latest NVIDIA Quantum-X800 InfiniBand and Spectrum™-X800 Ethernet fabrics.
Furthermore, embedded NVIDIA BlueField-3 DPUs power hyper-scalable AI environments by enabling elastic GPU computing, zero-trust security frameworks, composable storage architectures, and optimized cloud networking.
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.
Your email address will not be published. Required fields are marked *