GPU FABRIC

GPU Cluster Network Architecture — Narrated Tutorial

RoCEv2 · RDMA · Rail-Optimized Stripes · Dragonfly+ DLB · PFC / ECN / DCQCN

NVIDIA B300 · 800G
RoCEv2 packet
Switch egress queue
0%
0%ECN 70%PFC 90%100%
🎙
Ready to begin
Press Play to hear the narration for this step, or press Next to advance.
0.5× 2× 1.0×
Step 1 / 1
Step 0 / 0
KNOWLEDGE CHECK
🎉
Lesson complete
🎓
Course complete
You finished all 7 lessons of the GPU Cluster Network Architecture tutorial: RDMA and RoCEv2 · rail-optimized routing · ECN and DCQCN · PFC · Dragonfly+ DLB · adaptive and Valiant routing.
⚡
GPU platform lab
Pick an accelerator on each side, run the same fabric model on both, and compare network performance and cost
SIDE A — also drives the topology
SIDE B — comparison
A
B
CLUSTER
32 nodes · 256 GPUs
405B parameters
VERDICT
AT A GLANCE
Assumptions. Ring all-reduce over BF16 gradients; 4M-token global batch; 40% peak MFU; 85% RoCEv2 line-rate efficiency; fabric efficiency 90% rail-optimized, 94% DLB; 70% of compute time available to hide comms behind. 64-port leaves with 32 host and 32 fabric ports, spine count leaves ÷ 2. $75k per node for CPU, chassis and storage. Power at $0.10/kWh, PUE 1.3, 3-year window.

Read the comparison, not the absolutes. This models a naive data-parallel all-reduce over the full gradient set. Real training shards optimiser state and overlaps far more aggressively, so step times here are an upper bound. The ratio between two platforms on the same workload is the useful output.

Prices are public street-price estimates and move constantly — read them as order-of-magnitude, not as a quote. Specs are vendor-published.