Skip to main content
Optiscale
Log inContact us
Solutions
HPC Cluster Design · CPU supercomputersAI Supercomputer · GPU training & inference clustersParallel Storage · Spectrum Scale · Lustre · BeeGFS · VASTInterconnect · InfiniBand · RoCE fabricsSchedulers · Slurm · KubernetesApplication Workload · Reference diagrams by workloadSizing Calculator · 5-minute sizing + 5-year TCO
Products
Luxe Vision · Unified monitoring across heterogeneous resourcesLuxe Ray · Heterogeneous resource operations & unified controlLuxe Vantage · Usage & billing built on the AI/HPC schedulerLuxe Orbit · Heterogeneous software provisioning & parallel managementLuxe Series Overview · The unified 4-layer storyRequest a Live Demo · Vision · Ray · VantageRequest a Closed PoC · Orbit — 1:1 environment setup
Services
Architecture Consulting · Workload definition → specification designImplementation & PM · Vendor integration & project managementPerformance Tuning · Architecture & configuration optimization plus code-level performance gainsManaged Operations · Multi-year operations contracts
Resources
TCO Calculator · On-premises vs. the three major cloudsSelf-Check · 8-question workload assessmentApplication Workload Library · 6 workload diagramsWhite Papers · In-depth technical white papersBlog · Tech Notes · Engineering blogNewsletter · Biweekly infrastructure updates

Solutions

Interconnect

InfiniBand · RoCE Fabrics

InfiniBand XDR/NDR, RoCEv2, and high-density switching — 95%+ efficiency versus spec on MPI/NCCL collectives.

The challenges clients bring to us

  • Miscalculated fat-tree blocking ratios
  • Difficulty tuning PFC/ECN for RoCEv2
  • Routing collisions in multi-hop topologies
  • Operational complexity when mixing IB + Ethernet at the same site

Reference Architecture

Non-blocking fat-tree (InfiniBand XDR 800Gbps / NDR 400Gbps · RoCEv2 comparison recommended)SPINE-1SPINE-2SPINE-3SPINE-4LEAF-1LEAF-2LEAF-3LEAF-4LEAF-5LEAF-6Rack 1x32 hostsRack 2x32 hostsRack 3x32 hostsRack 4x32 hostsRack 5x32 hostsRack 6x32 hosts
Networking — non-blocking fat-tree (4 spine × 6 leaf, adaptive routing recommended)
02DESIGN DECISIONS

Key design decisions

Why we chose each component, and the assumptions and thresholds behind it — we publish our design decisions as they are.

Q.01

What are the criteria for choosing between IB and RoCEv2?

The decision rests on three axes. (1) Workload latency sensitivity — if a 1 μs difference in MPI AllReduce translates directly to 1% performance, IB comes first. (2) Operations-team RDMA experience — RoCEv2 has many tuning parameters (PFC headroom, ECN threshold, DCQCN, etc.) and needs 1–2 quarters of stabilization on first adoption. (3) Cost — RoCEv2 is roughly 20–30% cheaper at the same bandwidth, but once operations staffing costs are included, IB often has the TCO advantage at large scale. A design mixing IB + Ethernet on the same fabric (separating storage and management networks) is also possible.

Q.02

How do you set the fat-tree blocking ratio?

It is determined by the share of collective traffic (AllReduce, etc.) that traverses the spine. AI training (60–80% NCCL collectives) requires non-blocking (1:1); HPC MPI can tolerate 1:2 (oversubscribed). Simple inference/serving workloads are fine down to 1:4. We validate with osu_bw and osu_alltoall measurements during the PoC phase.

Q.03

What are the criteria for adopting NDR (200/400 Gbps)?

To fully utilize a single GPU node’s PCIe Gen5 ×16 bandwidth (~64 GB/s, about 512 Gbps), 2× NDR 400Gbps NICs are recommended. For H100 / B300 × 8 GPU nodes, NDR 1 port = 400 Gbps is the minimum recommended spec. Expanding an existing EDR (100Gbps) site is also possible in stages via NDR down-converter mode.

03PERFORMANCE

Validated performance numbers

Estimated ranges from internal PoCs and customer acceptance tests (HPL · STREAM · IOR · MDTest · osu_*). Real-site results may vary by around ±10% depending on workload and topology.

MetricValueCondition
osu_latency0.93 μsNDR 400Gbps, 8B payload
osu_bw bidirectional194 Gbps64 MB, NDR single port (97% of the theoretical 200 Gbps)
osu_alltoall88% peakNDR fat-tree, 64 nodes
NCCL AllReduce94–97% peakRail-Optimized topology, 512 GPUs
04VENDOR MATRIX

Vendor matrix

Vendors validated at real sites — recommended to match workload characteristics and your operations team's expertise.

  • NVIDIA Mellanox
  • Cornelis
  • Arista
  • Cisco

Have a project in this domain?

Talk to an engineer
Interconnect | Optiscale