Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction
| Source: arXiv AI
Tags: HPC, job scheduling, CatBoost, machine learning, grid computing, runtime prediction
CatBoost-based job runtime prediction on the AuverGrid trace achieves R²=0.239 under temporal validation and reduces simulated scheduling wait time by 50.92% — while showing that random cross-validation overstates real-world accuracy by a significant margin.
Details
Predicting how long a compute job will run before it starts is a classic HPC scheduling problem. This paper revisits it using the GWA-T-4 AuverGrid dataset, but with a methodologically stricter setup than most prior work: the authors use only submission-time features (no post-execution information), apply temporal and cold-start validation splits rather than random cross-validation, and explicitly analyze runtime predictability by job length. CatBoost with native categorical handling achieves the best results under temporal validation: R²=0.239, MAE=27,019 seconds, RMSE=46,587. The gap between random-split performance and temporal performance is substantial — a key finding that most published benchmarks likely overstate real-world model accuracy because they leak future information via random splits. A minimal scheduling simulation across 69,523 held-out test jobs shows that prediction-informed Shortest Job First (SJF) reduces average wait time by 50.92% compared to First-Come-First-Served (FCFS). Long-job underestimation remains the hardest error mode and the most scheduler-relevant failure case. Accepted at ICMLA 2026. The methodological contributions around evaluation rigor (temporal splits, leakage-safe features, cold-start testing) are as valuable as the modeling results themselves for practitioners.