Skip to main content
Optiscale
Log inContact us
Solutions
HPC Cluster Design · CPU supercomputersAI Supercomputer · GPU training & inference clustersParallel Storage · Spectrum Scale · Lustre · BeeGFS · VASTInterconnect · InfiniBand · RoCE fabricsSchedulers · Slurm · KubernetesApplication Workload · Reference diagrams by workloadSizing Calculator · 5-minute sizing + 5-year TCO
Products
Luxe Vision · Unified monitoring across heterogeneous resourcesLuxe Ray · Heterogeneous resource operations & unified controlLuxe Vantage · Usage & billing built on the AI/HPC schedulerLuxe Orbit · Heterogeneous software provisioning & parallel managementLuxe Series Overview · The unified 4-layer storyRequest a Live Demo · Vision · Ray · VantageRequest a Closed PoC · Orbit — 1:1 environment setup
Services
Architecture Consulting · Workload definition → specification designImplementation & PM · Vendor integration & project managementPerformance Tuning · Architecture & configuration optimization plus code-level performance gainsManaged Operations · Multi-year operations contracts
Resources
TCO Calculator · On-premises vs. the three major cloudsSelf-Check · 8-question workload assessmentApplication Workload Library · 6 workload diagramsWhite Papers · In-depth technical white papersBlog · Tech Notes · Engineering blogNewsletter · Biweekly infrastructure updates
BLOGOps

GPU Utilization from 30% to 78% — What We Did Over Six Months

How we more than doubled GPU utilization on a university AI cluster by redesigning queue policy and giving users visibility.

Engineering Team··9 min read

The Starting Point — Users Come Looking for Answers

Sixty people were running training jobs simultaneously on a cluster of 80 H100s. Average GPU utilization was stuck in the low 30s, and user satisfaction sat at 2.3 out of 5.

The operations team already had monitoring tools. Twelve Grafana dashboards, in fact. Yet no one could quickly find the answer they actually needed.

Step 1 — Measure Workload Patterns for Four Weeks

The first thing we did was measure. We pulled four weeks of job logs from Slurm sacct and analyzed:

- Job size distribution (1-GPU jobs vs. 8-GPU jobs) - Job length distribution (under 1 hour vs. over 24 hours) - Per-user resource share - Queue wait time distribution

The finding: with a single priority queue, whenever a 24-hour training job grabbed priority, short debugging jobs were pushed back indefinitely.

Step 2 — Redesign the Queue Policy

We applied three changes at once:

1. Enable backfill — slot small jobs into the schedule without shifting the start times of large jobs. 2. Fair-share — factor each user's 14-day resource usage into priority as a dynamic weight. 3. Queue separation — a debugging queue (2-hour limit, priority +500), a main training queue, and a one-off interactive queue.

Step 3 — Use Vision to Answer Users Directly

Half the problem came from users walking over to the ops team to ask, "When does my job start?" We rolled out the user view in Luxe Vision.

- Queue wait prediction (based on sacct history) - Real-time "My Jobs" status - Per-department usage (finance team view)

Queue-related inquiries to the ops team dropped 80% immediately.

Six Months Later — The Numbers

| Metric | Before | After |
|------|--------|-------|
| GPU utilization (weighted average) | 31% | 78% |
| Average job wait time | 4.2 hours | 1.2 hours |
| Ops team queue inquiries | 12/day avg | 2/day avg |
| User satisfaction | 2.3 / 5 | 4.4 / 5 |

What We Learned

- Queue policy isn't a tool; it's a function of user behavior. Design it without measurement and you'll miss the mark. - Visibility reduces the burden on the ops team. When users find answers themselves, everyone is happier. - Six months is not a long time. Look at the data once a week, and fine-tune the policy once a week.

The detailed case page for this project: /case-studies/research-cluster-utilization.

GPU Utilization from 30% to 78% — What We Did Over Six Months | Optiscale