Perhimpunan Mahasiswa SUTD Indonesia (PADI
Auto-Scaling Framework for Heterogeneous NPUs
Pages
17
Time to read
82 mins
Publication
Language
English
Pages
17
Time to read
82 mins
Publication
Language
English
This document is a technical report that presents NeuScale, an auto-scaling framework designed for managing heterogeneous neural processing units (NPUs) in cloud platforms, specifically for large language model (LLM) services. The report begins with a characterization study of various NPU chip generations, highlighting the challenges posed by NPU heterogeneity and the lack of system support for efficient management. NeuScale introduces a new abstraction called vPod, which allows for compatibility with existing machine learning frameworks while optimizing resource allocation. The framework employs a roofline-based analysis to determine the best-fit vPod allocations for different LLM inference requests, thereby improving energy and cost efficiency. The report details the implementation of NeuScale using a production-level NPU simulator and validates its effectiveness in enhancing service-level objective satisfaction rates. The findings indicate that NeuScale can significantly optimize the utilization of heterogeneous NPU resources, addressing the challenges faced by modern cloud platforms.