Skip to main content
Optiscale
Log inContact us
Solutions
HPC Cluster Design · CPU supercomputersAI Supercomputer · GPU training & inference clustersParallel Storage · Spectrum Scale · Lustre · BeeGFS · VASTInterconnect · InfiniBand · RoCE fabricsSchedulers · Slurm · KubernetesApplication Workload · Reference diagrams by workloadSizing Calculator · 5-minute sizing + 5-year TCO
Products
Luxe Vision · Unified monitoring across heterogeneous resourcesLuxe Ray · Heterogeneous resource operations & unified controlLuxe Vantage · Usage & billing built on the AI/HPC schedulerLuxe Orbit · Heterogeneous software provisioning & parallel managementLuxe Series Overview · The unified 4-layer storyRequest a Live Demo · Vision · Ray · VantageRequest a Closed PoC · Orbit — 1:1 environment setup
Services
Architecture Consulting · Workload definition → specification designImplementation & PM · Vendor integration & project managementPerformance Tuning · Architecture & configuration optimization plus code-level performance gainsManaged Operations · Multi-year operations contracts
Resources
TCO Calculator · On-premises vs. the three major cloudsSelf-Check · 8-question workload assessmentApplication Workload Library · 6 workload diagramsWhite Papers · In-depth technical white papersBlog · Tech Notes · Engineering blogNewsletter · Biweekly infrastructure updates
00Resources · Reference Architectures

A validated architecture library.

We publish per-workload diagrams alongside their measurement assumptions. A starting point for writing your RFP and making decisions.

v0 seed (Phase 1A) — interactive SVGs and IaC modules land in Phase 1B.

01RA
v0 published

AI Supercomputer

Rail-Optimized fat-tree + NDR + NVMe-oF — 96% NCCL efficiency across 64 B300×8 nodes

Solutions page
Rail-Optimized Fat-Tree (InfiniBand XDR 800Gbps / NDR 400Gbps) — adaptive routingSPINE-1SPINE-2SPINE-3SPINE-4LEAF-1LEAF-2LEAF-3LEAF-4LEAF-5LEAF-6B300x8Rack 1B300x8Rack 2B300x8Rack 3B300x8Rack 4B300x8Rack 5B300x8Rack 6Parallel Storage + NVMe-oF (checkpoint IO)Spectrum Scale · Lustre · BeeGFS · VAST · DAOS12 NSD nodes · 320 GB/s aggregate · 12 PBSeparate metadata / data NSDsHead + Mgmt(Luxe Orbit HA)
AI Supercomputer Reference Architecture (Rail-Optimized) — B300x8 × N nodes, InfiniBand XDR 800Gbps / NDR 400Gbps fat-tree, NVMe-oF + Spectrum Scale · Lustre · BeeGFS · VAST, Luxe Orbit management-server HA
Performance · Validated figures
  • NCCL AllReduce efficiency

    94–97% peak

    B300 × 64 nodes (512 GPUs), NDR InfiniBand, Rail-Optimized

  • LLM training GPU utilization

    78–85%

    70B parameters, BF16, gradient checkpointing

  • Checkpoint I/O (aggregate)

    320 GB/s

    64 nodes writing concurrently, Spectrum Scale + NVMe-oF

02RA
v0 published

HPC Cluster Design

Slurm + Spectrum Scale · Lustre · BeeGFS · VAST + InfiniBand XDR/NDR — 95%+ HPL across 256 B300×8 nodes

Solutions page
CPU Supercomputer (FFT/CFD/MD) — InfiniBand XDR 800Gbps / NDR 400GbpsLogin Nodesx2 · LDAP/SlurmMgmt (HA)Luxe Orbit · ipmiInfiniBand XDR 800Gbps / NDR 400Gbps fat-tree (non-blocking)cn-01B300x8cn-02B300x8cn-03B300x8cn-04B300x8cn-05B300x8cn-06B300x8cn-07B300x8cn-08B300x8cn-09B300x8cn-10B300x8cn-11B300x8cn-12B300x8ParallelStorageSpectrum ScaleLustre · BeeGFSVAST · DAOSNSD x8240 GB/s
HPC Cluster Reference Architecture — Slurm + Spectrum Scale · Lustre · BeeGFS · VAST + InfiniBand XDR/NDR, management-server HA, 12+ compute-node seed
Performance · Validated figures
  • HPL (Linpack)

    95–97% theoretical

    Sapphire Rapids 8480+ × 256 nodes, IB XDR/NDR

  • STREAM Triad

    380–410 GB/s

    per node, DDR5-4800, 8 channel × 2 socket

  • IOR write (sequential)

    120 GB/s

    Spectrum Scale 8 NSD nodes, 64 MB block

03RA
Seed

Parallel Storage

Spectrum Scale metadata/data NSD separation + AFM DR — 240 GB/s write · 320 GB/s read across 12 NSDs

Solutions page
Spectrum Scale · Lustre · BeeGFS · VAST — distributed metadata + separate data NSDsHPC / AI Clients(200+ nodes via IB)Metadata NSDs (NVMe)x4 · 1.2M ops/sMetadata IO isolationData NSDs (HDD/SSD tier)x12 NSD · 180 GB/s sequentialHot / Warm / Cold tier (HSM)AFM Cache (Remote)DR / multi-site
Parallel Storage — Spectrum Scale · Lustre · BeeGFS · VAST · DAOS with separate metadata / data NSDs, AFM DR cache
Performance · Validated figures
  • IOR write (sequential, 1 MB)

    240 GB/s

    12 NSD nodes, 64 client write

  • IOR read (sequential, 1 MB)

    320 GB/s

    12 NSD nodes, 64 client read

  • IOR random (4 KB)

    48 GB/s

    12 NSD nodes, 64 client, queue depth 64

04RA
Seed

Interconnect

Non-blocking fat-tree (4 spine × 6 leaf) — adaptive routing recommended

Solutions page
Non-blocking fat-tree (InfiniBand XDR 800Gbps / NDR 400Gbps · RoCEv2 comparison recommended)SPINE-1SPINE-2SPINE-3SPINE-4LEAF-1LEAF-2LEAF-3LEAF-4LEAF-5LEAF-6Rack 1x32 hostsRack 2x32 hostsRack 3x32 hostsRack 4x32 hostsRack 5x32 hostsRack 6x32 hosts
Networking — non-blocking fat-tree (4 spine × 6 leaf, adaptive routing recommended)
Performance · Validated figures
  • osu_latency

    0.93 μs

    NDR 400Gbps, 8B payload

  • osu_bw bidirectional

    194 Gbps

    64 MB, NDR single port (97% of the theoretical 200 Gbps)

  • osu_alltoall

    88% peak

    NDR fat-tree, 64 nodes

05RA
Seed

Schedulers

Slurm + Kubernetes time-slicing a shared GPU pool — unified Vantage metrics

Solutions page
Slurm + Kubernetes hybrid — time-of-day partitioning of a shared GPU poolSlurm (HPC / Batch)Backfill + fair-share + queue separation· queue:debug· queue:short· queue:gpu· queue:longNightly training jobs prioritizedKubernetes (Online / Inference)Volcano + GPU Operator· ns:inference· ns:notebooks· ns:dev· ns:internalDaytime services prioritizedShared GPU Pool (B300x8 × N)Vantage unifies metrics across both Slurm and K8sTime-of-day partitioning maximizes utilization
Workload Schedulers — Slurm + Kubernetes time-of-day partitioning of a shared GPU pool (Luxe Vantage unified metrics)
06RA
Seed

Library Overview

Overview of all 6 workloads + the Phase 1B roadmap for completing the interactive SVGs

Reference Architecture LibraryValidated interactive diagrams by workloadAI SupercomputerGPU + Rail-OptimizedHPC ClusterCPU + Slurm + IBParallel StorageSpectrum Scale · LustreBeeGFS · VAST · DAOSNetworkingIB (XDR/NDR) · RoCEv2SchedulersSlurm + K8sHybrid (Mixed)Phase 1B
Reference Architecture Library — 6 workloads (interactive SVGs completed in Phase 1B)

We design the architecture that fits your workload, together.

The library is only a starting point. We take responsibility for designing and tuning it to your actual environment.

Architecture consultation Get the RFP template
Reference Architectures — A Validated Architecture Library | Optiscale