Senior AI Infrastructure Engineer
Design, deployment, and tuning of GPU clusters. Own the full path from capacity sizing to 95%+ NCCL efficiency.
Responsibilities
- Analyze AI training/inference workloads and size capacity
- Design Rail-Optimized topologies and apply adaptive routing
- Integrate NVMe-oF with Spectrum Scale and isolate IO
- Partner with the deployment PM to validate acceptance criteria
- Facilitate RFP authoring workshops
Qualifications
- 5+ years deploying or operating GPU clusters
- Hands-on experience with NCCL / MPI / Slurm
- Experience troubleshooting InfiniBand or RoCEv2
- Proficient in writing technical documentation in English
Preferred
- PoC experience at 64+ H100 node scale
- Experience tuning Spectrum Scale / Lustre metadata
- Experience operating Bright Cluster Manager / Warewulf / xCAT