Skip to main content
Optiscale
Log inContact us
Solutions
HPC Cluster Design · CPU supercomputersAI Supercomputer · GPU training & inference clustersParallel Storage · Spectrum Scale · Lustre · BeeGFS · VASTInterconnect · InfiniBand · RoCE fabricsSchedulers · Slurm · KubernetesApplication Workload · Reference diagrams by workloadSizing Calculator · 5-minute sizing + 5-year TCO
Products
Luxe Vision · Unified monitoring across heterogeneous resourcesLuxe Ray · Heterogeneous resource operations & unified controlLuxe Vantage · Usage & billing built on the AI/HPC schedulerLuxe Orbit · Heterogeneous software provisioning & parallel managementLuxe Series Overview · The unified 4-layer storyRequest a Live Demo · Vision · Ray · VantageRequest a Closed PoC · Orbit — 1:1 environment setup
Services
Architecture Consulting · Workload definition → specification designImplementation & PM · Vendor integration & project managementPerformance Tuning · Architecture & configuration optimization plus code-level performance gainsManaged Operations · Multi-year operations contracts
Resources
TCO Calculator · On-premises vs. the three major cloudsSelf-Check · 8-question workload assessmentApplication Workload Library · 6 workload diagramsWhite Papers · In-depth technical white papersBlog · Tech Notes · Engineering blogNewsletter · Biweekly infrastructure updates
CASEEnterprise AI Division
TIER T2

300-node H100 cluster deployment

Workload
LLM Training (~70B)
Scale
300 H100 nodes / NDR InfiniBand / 12PB storage
Performance vs. spec
96%
01Challenge

The situation

The prior PoC delivered only 60% NCCL all-reduce efficiency. Checkpoint I/O consumed 20% of training time. Multi-tenancy was unresolved.

02Approach

The approach

  1. 01Rail-Optimized topology + adaptive routing
  2. 02NVMe-oF + Spectrum Scale with separated metadata/data NSDs
  3. 03Slurm fair-share + priority queues
  4. 04Luxe Orbit management server redundancy + automated OS deployment
03Results

Results

NCCL all-reduce
94-97% peak
Checkpoint I/O share
< 2%
Deployment timeline
14 weeks
04Suite Used

Operations automation — Luxe Series products used

  • Luxe Orbit

Have a similar case?

Just tell us your workload type and scale, and we'll give you a first diagnosis within 30 minutes.

Request a consultation
300-node H100 cluster deployment | Optiscale