Back
Ai2 Replaces Priority-Based GPU Scheduler With Budget and Fair-Share System
SiTech AI Team3 min read

Ai2 Replaces Priority-Based GPU Scheduler With Budget and Fair-Share System

The Allen Institute for AI swapped its priority-based cluster scheduler for GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract, cutting debug job waits from hours to seconds.

The Allen Institute for AI (Ai2) has replaced the priority-based scheduler on its GPU clusters with a system built around GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The change, rolled out cluster by cluster starting at the end of July, aims to prioritize high-impact research while keeping thousands of Nvidia H100, B200, and B300 GPUs fully occupied.

From priorities to budgets

Ai2 manages clusters ranging from 88 to 1024 GPUs for roughly 150 internal researchers working on large language model and vision-language model training, robotics reinforcement learning, and post-training for scientific agentic use cases. Demand for GPU time exceeds supply by a factor of two to three at any given moment.

The old scheduler relied on priority levels and optional preemption. That produced predictable pathologies: users parked no-op squatting workloads to keep GPUs available for debugging, priority inflation left 100 percent of scheduled workloads at HIGH priority, and on-call engineers spent most of their ticket time negotiating manual shutdowns of non-preemptable jobs on hosts needing maintenance.

The new model treats GPU time as an investment. Managers proportionally allocate budgets to projects and researchers before workloads exist, translating program strategy into a guaranteed share of capacity. Every request must be funded by a budget or it is not protected from preemption, which makes squatting workloads directly expensive for the team doing it.

Fair-share scheduling and the time-slicing contract

A hierarchical fair-share scheduler, an approach with roots in the Hadoop Fair Scheduler from 2009 and used today in SLURM and YARN, tracks occupancy over a sliding seven-day lookback window and sorts workloads from under-utilized allocations above over-utilized ones. The tree mirrors the research program structure, and the weights are the budgets set by managers rather than static quotas.

The scheduler distinguishes allocated occupancy, which is charged to a budget and protected during a minimum runtime window, from unallocated occupancy, which is free but preemptible. A scheduling contract requires each workload to declare a minimum runtime, capped at 8 hours, during which it cannot be preempted. After that window, resumable workloads can be automatically requeued so fair-share can converge and unhealthy hosts can drain for automated repair.

Results and open challenges

Over a 30-day test period, teams received 98 percent of the GPU hours they were owed, with 13 of 15 team allocations getting 95 percent or more and the worst case at 90 percent. Cluster occupancy held steady at 98 percent, with 18 percent of delivered GPU time unallocated to keep hardware busy when funded workloads were not ready.

Debug workloads saw p90 queue wait times fall from 2 hours to 30 seconds, beating the simulated prediction of 6 hours to 5 minutes. On the largest H100 cluster, median queue wait fell from 5 minutes to 24 seconds and p90 fell from 2.8 hours to 1.8 hours. Repairs requiring a human in the loop dropped by 74 percent.

Not every use case improved. Interactive data analysis sessions, which researchers previously held for up to a week, became preemptible after the 8-hour protected runtime cap, forcing users to rebuild volatile state by hand. Ai2 is now investing in a CPU-only cluster for development sessions and plans restorable sessions so preempted work can resume elsewhere. The team is also investigating capacity fragmentation, which may increase queue waits for the largest workloads, and says its next focus is utilization: making bootstrapping, checkpointing, and training applications as efficient as possible.

Sources: Hugging Face Blog

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.