SuperMicro
Hybrid CPU and GPU Architecture for LLM Inference
Pages
23
Time to read
18 mins
Publication
Language
English
Pages
23
Time to read
18 mins
Publication
Language
English
This white paper presents an evaluation of a hybrid CPU-GPU architecture that integrates Intel Xeon 6 processors with Intel Gaudi 3 accelerators for scalable and cost-efficient large language model (LLM) inference. The document outlines the transition in the generative AI landscape towards modular multi-agent systems, where complex queries are decomposed into orchestrated pipelines. Key performance metrics such as latency, throughput, and Time to First Token (TTFT) are measured, alongside a comparison with previous-generation configurations to illustrate improvements in efficiency. The findings indicate that Intel Xeon 6 processors effectively manage high-concurrency workloads for smaller models, while Intel Gaudi 3 deployments significantly enhance large-model inference. The architecture is shown to optimize resource utilization, reduce hardware requirements, and improve total cost of ownership (TCO) by intelligently partitioning tasks between CPU and GPU resources. This approach ensures high-throughput and low-latency inference across various model sizes, validating the effectiveness of hybrid architectures in modern AI infrastructure.