Skip to main content
Optiscale
Log inContact us
Solutions
HPC Cluster Design · CPU supercomputersAI Supercomputer · GPU training & inference clustersParallel Storage · Spectrum Scale · Lustre · BeeGFS · VASTInterconnect · InfiniBand · RoCE fabricsSchedulers · Slurm · KubernetesApplication Workload · Reference diagrams by workloadSizing Calculator · 5-minute sizing + 5-year TCO
Products
Luxe Vision · Unified monitoring across heterogeneous resourcesLuxe Ray · Heterogeneous resource operations & unified controlLuxe Vantage · Usage & billing built on the AI/HPC schedulerLuxe Orbit · Heterogeneous software provisioning & parallel managementLuxe Series Overview · The unified 4-layer storyRequest a Live Demo · Vision · Ray · VantageRequest a Closed PoC · Orbit — 1:1 environment setup
Services
Architecture Consulting · Workload definition → specification designImplementation & PM · Vendor integration & project managementPerformance Tuning · Architecture & configuration optimization plus code-level performance gainsManaged Operations · Multi-year operations contracts
Resources
TCO Calculator · On-premises vs. the three major cloudsSelf-Check · 8-question workload assessmentApplication Workload Library · 6 workload diagramsWhite Papers · In-depth technical white papersBlog · Tech Notes · Engineering blogNewsletter · Biweekly infrastructure updates

Solutions

HPC Cluster Design

CPU Supercomputers

CPU-based supercomputers — CFD, molecular dynamics, materials science, weather, and astrophysics. Designed to run single jobs at the million-core-hour scale, reliably.

The challenges clients bring to us

  • Inter-node communication bottlenecks in MPI workloads
  • Miscalculated parallel filesystem IOPS
  • Scheduler queue policies misaligned with usage patterns
  • No clear owner accountable for the full service lifecycle

Reference Architecture

CPU Supercomputer (FFT/CFD/MD) — InfiniBand XDR 800Gbps / NDR 400GbpsLogin Nodesx2 · LDAP/SlurmMgmt (HA)Luxe Orbit · ipmiInfiniBand XDR 800Gbps / NDR 400Gbps fat-tree (non-blocking)cn-01B300x8cn-02B300x8cn-03B300x8cn-04B300x8cn-05B300x8cn-06B300x8cn-07B300x8cn-08B300x8cn-09B300x8cn-10B300x8cn-11B300x8cn-12B300x8ParallelStorageSpectrum ScaleLustre · BeeGFSVAST · DAOSNSD x8240 GB/s
HPC Cluster Reference Architecture — Slurm + Spectrum Scale · Lustre · BeeGFS · VAST + InfiniBand XDR/NDR, management-server HA, 12+ compute-node seed
02DESIGN DECISIONS

Key design decisions

Why we chose each component, and the assumptions and thresholds behind it — we publish our design decisions as they are.

Q.01

Why InfiniBand?

In large-scale MPI collectives (AllReduce, AllGather, etc.), a 1 μs difference in network latency translates to a 5–15% difference in total training/simulation time. InfiniBand XDR 800Gbps / NDR 400Gbps guarantees end-to-end latency around 200ns via an RDMA hardware path + adaptive routing, and — unlike RoCEv2 — requires no separate PFC/ECN tuning. Running the same workload on RoCEv2 leaf-spine ECMP adds routing-collision and congestion-control work, so IB is the more stable choice unless the operations team has sufficiently deep RDMA expertise.

Q.02

Why IBM Spectrum Scale?

Large-scale HPC workloads must handle three things well simultaneously: (1) metadata hot-spots, (2) failure-group (FG) isolation on disk failure, and (3) multi-site DR. Spectrum Scale 5.2 has been validated with distributed metadata NSDs + Spectrum Scale failure groups + AFM (Active File Management) caching together, giving it a clear edge in operational stability over Lustre for large writes. Lustre carries the risk of a single-node MDS bottleneck and tricky ldiskfs/ZFS backend tuning, which is dangerous at understaffed sites.

Q.03

Why Slurm?

The real-world usage patterns of research institutes, universities, and government supercomputing sites (per-department queue separation, priorities, backfill, fair-share) fit Slurm’s policy model better than PBS Pro or LSF. Slurm runs on an active community with a plugin architecture (SPANK), sacct accounting data, and burst-buffer integration, and it also integrates well at the usage-data (metrics) level with K8s-side schedulers such as Run:ai and Volcano.

03PERFORMANCE

Validated performance numbers

Estimated ranges from internal PoCs and customer acceptance tests (HPL · STREAM · IOR · MDTest · osu_*). Real-site results may vary by around ±10% depending on workload and topology.

MetricValueCondition
HPL (Linpack)95–97% theoreticalSapphire Rapids 8480+ × 256 nodes, IB XDR/NDR
STREAM Triad380–410 GB/sper node, DDR5-4800, 8 channel × 2 socket
IOR write (sequential)120 GB/sSpectrum Scale 8 NSD nodes, 64 MB block
MPI osu_latency0.93–1.10 μsNDR 400Gbps, 8B payload, intra-rack
04VENDOR MATRIX

Vendor matrix

Vendors validated at real sites — recommended to match workload characteristics and your operations team's expertise.

  • Intel
  • AMD
  • HPE
  • Dell
  • Lenovo
  • Supermicro
  • IBM
  • NVIDIA Mellanox

Operations automation — Luxe Series

Our own products that automate operations in this domain

FAQ

  • Q. Can domestic CPUs (still in engineering validation) be considered?

    Yes, provided BLAS and MPI profiling is done first. At this stage we recommend a single-workload PoC.

Have a project in this domain?

Talk to an engineer
HPC Cluster Design | Optiscale