Source Hugging Face · Published · Open the original ↗
Lab notes
Impactful scheduling for GPU clusters
A Blog post by Ai2 on Hugging Face
Building a cluster scheduler to prioritize high-impact research while maintaining full occupancy
On the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads. We think about this task as a pyramid of four metrics that build on each other.
The foundation is availability : how often the hardware is healthy and ready for work. Above this is occupancy : the fraction of available time assigned to a specific workload. Next is impact : how often the most valuable workloads are chosen to receive resources. The capstone of the pyramid is utilization : the fraction of GPU capacity used over the lifetime of a workload.
The opening of the post, quoted unchanged from Hugging Face. The full text continues at the source.