The Starting Point — Users Come Looking for Answers
Sixty people were running training jobs simultaneously on a cluster of 80 H100s. Average GPU utilization was stuck in the low 30s, and user satisfaction sat at 2.3 out of 5.
The operations team already had monitoring tools. Twelve Grafana dashboards, in fact. Yet no one could quickly find the answer they actually needed.
Step 1 — Measure Workload Patterns for Four Weeks
The first thing we did was measure. We pulled four weeks of job logs from Slurm sacct and analyzed:
- Job size distribution (1-GPU jobs vs. 8-GPU jobs) - Job length distribution (under 1 hour vs. over 24 hours) - Per-user resource share - Queue wait time distribution
The finding: with a single priority queue, whenever a 24-hour training job grabbed priority, short debugging jobs were pushed back indefinitely.
Step 2 — Redesign the Queue Policy
We applied three changes at once:
1. Enable backfill — slot small jobs into the schedule without shifting the start times of large jobs. 2. Fair-share — factor each user's 14-day resource usage into priority as a dynamic weight. 3. Queue separation — a debugging queue (2-hour limit, priority +500), a main training queue, and a one-off interactive queue.
Step 3 — Use Vision to Answer Users Directly
Half the problem came from users walking over to the ops team to ask, "When does my job start?" We rolled out the user view in Luxe Vision.
- Queue wait prediction (based on sacct history) - Real-time "My Jobs" status - Per-department usage (finance team view)
Queue-related inquiries to the ops team dropped 80% immediately.
Six Months Later — The Numbers
| Metric | Before | After | |------|--------|-------| | GPU utilization (weighted average) | 31% | 78% | | Average job wait time | 4.2 hours | 1.2 hours | | Ops team queue inquiries | 12/day avg | 2/day avg | | User satisfaction | 2.3 / 5 | 4.4 / 5 |
What We Learned
- Queue policy isn't a tool; it's a function of user behavior. Design it without measurement and you'll miss the mark. - Visibility reduces the burden on the ops team. When users find answers themselves, everyone is happier. - Six months is not a long time. Look at the data once a week, and fine-tune the policy once a week.
The detailed case page for this project: /case-studies/research-cluster-utilization.