Resource-Aware GPU Scheduling for Large-Scale Foundation Model Training in Hybrid Cloud Environments

Authors

  • Sudhakar Murthy Molli

Abstract

Training large-scale foundation models consumes GPU capacity at a scale and duration that expose the limitations of scheduling strategies inherited from earlier, shorter-lived batch and microservice workloads. Multi-tenant GPU clusters supporting foundation model training must reconcile long-running, gang-scheduled, communication-sensitive jobs with a long tail of short interactive and hyperparameter-search workloads, while increasingly spanning on-premises capacity and public cloud burst capacity to meet demand that fixed-size clusters cannot absorb. This paper presents a resource-aware scheduling architecture that combines topology-aware placement, elastic job resizing, checkpoint-aware preemption, and deadline- and cost-aware cloud bursting within a hierarchical scheduling hierarchy spanning global, regional, rack, and node-level decision points. We evaluate the architecture against fair-share, gang-scheduling, and priority-backfill baselines across clusters ranging from 64 to 4,096 GPUs. Our results show that the proposed scheduler improves average GPU utilization from 58% to 87% relative to a Kubernetes default fair-share baseline, reduces P99 job completion time by 62%, lowers per-training-run cost by 6% relative to an on-premises-only baseline while eliminating queueing delay through selective cloud bursting, and reduces stranded GPU capacity from 34% to 4% through dynamic partitioning and defragmentation. We conclude with a discussion of open challenges in cross-cluster job migration, carbon-aware scheduling, and the interaction between scheduling policy and model-training reproducibility.

References

[1] Verma, A., Pedrosa, L., Korupolu, M., Oppenheimer, D., Tune, E., & Wilkes, J. (2015). Large-scale cluster management at Google with Borg. Proceedings of the European Conference on Computer Systems.

[2] Xiao, W., Bhardwaj, R., Ramjee, R., Sivathanu, M., Kwatra, N., Han, Z., et al. (2018). Gandiva: Introspective cluster scheduling for deep learning. Proceedings of the USENIX Symposium on Operating Systems Design and Implementation.

[3] Narayanan, D., Santhanam, K., Kazhamiaka, F., Phanishayee, A., & Zaharia, M. (2020). Heterogeneity-aware cluster scheduling policies for deep learning workloads. Proceedings of the USENIX Symposium on Operating Systems Design and Implementation.

[4] Qiao, A., Choe, S. K., Subramanya, S. J., Neiswanger, W., Ho, Q., Zhang, H., et al. (2021). Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning. Proceedings of the USENIX Symposium on Operating Systems Design and Implementation.

[5] Peng, Y., Bao, Y., Chen, Y., Wu, C., & Guo, C. (2018). Optimus: An efficient dynamic resource scheduler for deep learning clusters. Proceedings of the European Conference on Computer Systems.

[6] Gu, J., Chowdhury, M., Shin, K. G., Zhu, Y., Jeon, M., Qian, J., et al. (2019). Tiresias: A GPU cluster manager for distributed deep learning. Proceedings of the USENIX Symposium on Networked Systems Design and Implementation.

[7] Mohan, J., Phanishayee, A., Raniwala, A., & Chidambaram, V. (2021). Analyzing and mitigating data stalls in DNN training. Proceedings of the VLDB Endowment.

[8] Jeon, M., Venkataraman, S., Phanishayee, A., Qian, J., Xiao, W., & Yang, F. (2019). Analysis of large-scale multi-tenant GPU clusters for DNN training workloads. Proceedings of the USENIX Annual Technical Conference.

[9] Weng, Q., Xiao, W., Yu, Y., Wang, W., Wang, C., He, J., et al. (2022). MLaaS in the wild: Workload analysis and scheduling in large-scale heterogeneous GPU clusters. Proceedings of the USENIX Symposium on Networked Systems Design and Implementation.

[10] Mahajan, K., Balasubramanian, A., Singhvi, A., Venkataraman, S., Akella, A., Phanishayee, A., & Chawla, S. (2020). Themis: Fair and efficient GPU cluster scheduling. Proceedings of the USENIX Symposium on Networked Systems Design and Implementation.

[11] Zaharia, M., Chowdhury, M., Das, T., Dave, A., Ma, J., McCauley, M., et al. (2012). Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. Proceedings of the USENIX Symposium on Networked Systems Design and Implementation.

[12] Or, A., Zhang, H., & Freedman, M. (2020). Resource elasticity in distributed deep learning. Proceedings of Machine Learning and Systems.

[13] Thinakaran, P., Gunasekaran, J. R., Sharma, B., Kandemir, M. T., & Das, C. R. (2019). Kube-Knots: Resource harvesting through dynamic container orchestration in GPU-based datacenters. Proceedings of the IEEE International Conference on Cluster Computing.

[14] Wu, Y., Ma, K., Yan, X., Liu, Z., & Cheng, J. (2021). Elastic deep learning in multi-tenant GPU cluster. arXiv preprint.

[15] Chen, Y., Ganapathi, A., Griffith, R., & Katz, R. (2011). The case for evaluating MapReduce performance using workload suites. Proceedings of the IEEE International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems

[16] Konda, P. R. (2026). Cloud-Native AI/ML Analytics Platform for Real-Time Enterprise Data Processing and Optimization. Synergia: A Journal of Multidisciplinary Innovation, 8(8). Retrieved from https://ijcdra.us/index.php/Synergia/article/view/72

[17] Sharma, M., Vangara, Y., Sharma, P., & Konda, P. R. (2025, June). NeuroNav: A Hybrid Deep Learning Framework for Sustainable Autonomous Indoor Robot Localization and Navigation. In International Conference on Sustainable Development through Machine Learning, AI and IoT (pp. 330-349). Cham: Springer Nature Switzerland.

[18] Konda, P. R. (2025). ADVANCED ENTERPRISE DATA ENGINEERING USING MACHINE LEARNING AND SCALABLE CLOUD ARCHITECTURES. Indonasian Journal of Advanced Research & Technology , 7(7). Retrieved from https://scholarlyarticle.vncinstitute.com/index.php/IJART/article/view/71

[19] Janakiraman, A. (2026). From Generative Intelligence to Agentic Autonomy: Leveraging Large Language Models for Multi-Agent Reasoning, Planning, and Execution. Journal of Integrated Science, AI and Engineering, 2(1).

[20] Janakiraman, A. (2026). Agentic Large Language Models for Autonomous Decision-Making and Adaptive Task Orchestration in Intelligent Systems. International Journal of Sustainable Digital and Computing Systems, 3(1).

[21] Pathak, S., Balantrapu, S. S., & Janakiraman, A. (2025). Future-Proofing the Planet: AI and XR for a Sustainable Tomorrow. In Exploring the Impact of Extended Reality (XR) Technologies on Promoting Environmental Sustainability (pp. 313-332). Cham: Springer Nature Switzerland.

[22] Janakiraman, A. (2025). Explainability and Interpretability in Generative AI Agents. International Journal of Science, Technology and Convergence, 7(7).

Downloads

Published

2026-02-20

How to Cite

Molli , S. M. (2026). Resource-Aware GPU Scheduling for Large-Scale Foundation Model Training in Hybrid Cloud Environments. International Journal of Science, Technology and Convergence, 8(8). Retrieved from https://ijcdra.us/index.php/IJSTC/article/view/92

Issue

Section

Articles