Making Sense of Job Preemption for Distributed Deep Learning Acceleration

Authors: Younghun Go, Changyong Shin, Minchul Kang, Jaehyun Hwang, Chuck Yoo, Gyeongsik Yang

Venue: DAC 2026 (63rd Chips to Systems Conference), Acceptance rate 22.3%, NRF BK IF: 3, Long Beach, California, USA, 2026. (Accepted) (2026)

Abstract: Preemptive scheduling is gaining attention in GPU scheduling for distributed training because it reduces job completion times (JCT). However, our analysis of production traces reveals that, surprisingly, it can increase JCT in practice. We identify two key contributors to this inefficiency. First, preemptions become "futile" when a job is preempted right after being loaded, before it begins execution. We find this futile preemption wastes the load time and so inflates JCT ~1.6⨉. Second, existing GPU schedulers run at fixed intervals (e.g., 360 s) rather than upon new job arrivals or when a job completes. So, new jobs must wait until the scheduler kicks in, which, we find, increases JCT ~2.6⨉. To address the problems, we introduce Lazer, a novel job scheduler that predicts when to preempt jobs based on job-specific and cluster conditions. Lazer is designed to efficiently explore the scheduling space and adapt to diverse job and GPU cluster characteristics based on Bayesian optimization. Our extensive evaluation shows that Lazer significantly outperforms state-of-the-art schedulers---reducing JCT by 1.2⨉--233.3⨉, waiting time by 2⨉--11690⨉, and futile preemptions by 2⨉--67⨉.