Skip to main content
Optiscale
Log inContact us
Solutions
HPC Cluster Design · CPU supercomputersAI Supercomputer · GPU training & inference clustersParallel Storage · Spectrum Scale · Lustre · BeeGFS · VASTInterconnect · InfiniBand · RoCE fabricsSchedulers · Slurm · KubernetesApplication Workload · Reference diagrams by workloadSizing Calculator · 5-minute sizing + 5-year TCO
Products
Luxe Vision · Unified monitoring across heterogeneous resourcesLuxe Ray · Heterogeneous resource operations & unified controlLuxe Vantage · Usage & billing built on the AI/HPC schedulerLuxe Orbit · Heterogeneous software provisioning & parallel managementLuxe Series Overview · The unified 4-layer storyRequest a Live Demo · Vision · Ray · VantageRequest a Closed PoC · Orbit — 1:1 environment setup
Services
Architecture Consulting · Workload definition → specification designImplementation & PM · Vendor integration & project managementPerformance Tuning · Architecture & configuration optimization plus code-level performance gainsManaged Operations · Multi-year operations contracts
Resources
TCO Calculator · On-premises vs. the three major cloudsSelf-Check · 8-question workload assessmentApplication Workload Library · 6 workload diagramsWhite Papers · In-depth technical white papersBlog · Tech Notes · Engineering blogNewsletter · Biweekly infrastructure updates

Solutions

AI Supercomputer

GPU Training & Inference Clusters

H100 / H200 / B200 / B300 + InfiniBand + high-performance storage — designed to run LLM training, multimodal, and reinforcement learning workloads all the way through.

The challenges clients bring to us

  • GPU utilization stuck around 30%
  • NCCL collective performance at 60% of spec
  • Checkpoint I/O consuming 20% of training time
  • Multi-tenancy — resource disputes between departments

Reference Architecture

Rail-Optimized Fat-Tree (InfiniBand XDR 800Gbps / NDR 400Gbps) — adaptive routingSPINE-1SPINE-2SPINE-3SPINE-4LEAF-1LEAF-2LEAF-3LEAF-4LEAF-5LEAF-6B300x8Rack 1B300x8Rack 2B300x8Rack 3B300x8Rack 4B300x8Rack 5B300x8Rack 6Parallel Storage + NVMe-oF (checkpoint IO)Spectrum Scale · Lustre · BeeGFS · VAST · DAOS12 NSD nodes · 320 GB/s aggregate · 12 PBSeparate metadata / data NSDsHead + Mgmt(Luxe Orbit HA)
AI Supercomputer Reference Architecture (Rail-Optimized) — B300x8 × N nodes, InfiniBand XDR 800Gbps / NDR 400Gbps fat-tree, NVMe-oF + Spectrum Scale · Lustre · BeeGFS · VAST, Luxe Orbit management-server HA
02DESIGN DECISIONS

Key design decisions

Why we chose each component, and the assumptions and thresholds behind it — we publish our design decisions as they are.

Q.01

Why a Rail-Optimized topology?

NCCL AllReduce traffic across N 8-GPU nodes is most efficient when it is aggregated within the same rail (GPUs at the same index). In a standard fat-tree, ECMP hash collisions scatter same-rail traffic across different spines, dropping effective bandwidth to 60–70% of spec. A Rail-Optimized topology + adaptive routing keeps same-rail traffic on the same path, recovering 95%+ efficiency. The benefit is greatest at ≥ 64 8-GPU nodes (≥ 512 GPUs).

Q.02

Why NVMe-oF (NVMe over Fabric)?

LLM training produces burst I/O every epoch or step, as N GPUs write checkpoints simultaneously. Standard NFS/SMB storage cannot absorb these bursts, leaving GPUs idle waiting on I/O. The combination of NVMe-oF (RDMA-based) + a Spectrum Scale backend hides inter-step I/O to 0% of GPU time, raising utilization to 78–85% in 70B-model training. We also evaluate checkpoint compression and shard distribution.

Q.03

Why a B300 × 8-GPU node configuration?

The NVIDIA Blackwell B300 delivers ~20 PFLOPs FP8 per GPU with 288 GB of HBM3e — more than enough for both LLM training and inference. Eight GPUs per node secures 1.8 TB/s of bidirectional inter-GPU bandwidth via NVLink/NVSwitch, while striking the balance point for single-node PCIe, memory channels, and power (within liquid-cooling limits). Some sites also run H100 × 8 / H200 × 8 alongside, separating pools by workload.

03PERFORMANCE

Validated performance numbers

Estimated ranges from internal PoCs and customer acceptance tests (HPL · STREAM · IOR · MDTest · osu_*). Real-site results may vary by around ±10% depending on workload and topology.

MetricValueCondition
NCCL AllReduce efficiency94–97% peakB300 × 64 nodes (512 GPUs), NDR InfiniBand, Rail-Optimized
LLM training GPU utilization78–85%70B parameters, BF16, gradient checkpointing
Checkpoint I/O (aggregate)320 GB/s64 nodes writing concurrently, Spectrum Scale + NVMe-oF
osu_latency (intra-rack)0.93 μsNDR 400Gbps, 8B payload
04VENDOR MATRIX

Vendor matrix

Vendors validated at real sites — recommended to match workload characteristics and your operations team's expertise.

  • NVIDIA
  • Dell
  • Supermicro
  • HPE
  • Lenovo
  • WekaIO
  • IBM Spectrum Scale

Operations automation — Luxe Series

Our own products that automate operations in this domain

FAQ

  • Q. Do you support AMD MI300X / Intel Gaudi?

    We have PoC experience with both. We recommend validating workload fit on the ROCm / Habana SynapseAI stack before adoption.

Have a project in this domain?

Talk to an engineer
AI Supercomputer | Optiscale