LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling

Dingyan Zhang; Jinbo Han; Kaixi Zhang; Xingda Wei; Sijie Shen; Chenguang Fang; Wenyuan Yu; Jingren Zhou; Rong Chen

LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling

Dingyan Zhang, Jinbo Han, Kaixi Zhang, Xingda Wei, Sijie Shen, Chenguang Fang, Wenyuan Yu, Jingren Zhou, Rong Chen

Abstract

High-quality LLM request scheduling requires achieving two key objectives: whether the routed instance has KV$ to accelerate the request execution and whether the workload is balanced across instances. Achieving both objectives is challenging because pursuing one objective may compromise the other. Current approaches adopt various combinators (e.g., linear combinations) to compute a scheduling score combining indicators for the two objectives, which are complex in that they either require significant workload-specific hyperparameter tuning or model-hardware-aware simulator development, and could still lead to suboptimal performance. In this paper, we show that using a simple multiplication of two carefully chosen indicators-one for KV$-aware (new prefill tokens if routed to an instance) and one for load balancing-aware (current batch size of the instance)-as the scheduling score can simultaneously achieve both objectives well without any hyperparameter tuning. The key idea is that the multiplied score considers both objectives in a manner similar to a linear combination, with a nice property that the original hyperparameters are canceled out during comparison so we don't need tuning to find the best parameters. The two indicators are chosen based on our analysis of LLM characteristics, and our extensive experiments show that this simple approach can reduce TTFT by 92% and 52%, and TPOT by 21% and 20%, compared to vLLM-v1 and a production scheduler on real-world workloads covering chatbots, API calls, and coding agents. We also mathematically derive the conditions under which multiplication may fail, and find that such conditions are extremely rare in practice and can be detected (and mitigated) beforehand.

LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling

Abstract

High-quality LLM request scheduling requires achieving two key objectives: whether the routed instance has KV

-aware (new prefill tokens if routed to an instance) and one for load balancing-aware (current batch size of the instance)-as the scheduling score can simultaneously achieve both objectives well without any hyperparameter tuning. The key idea is that the multiplied score considers both objectives in a manner similar to a linear combination, with a nice property that the original hyperparameters are canceled out during comparison so we don't need tuning to find the best parameters. The two indicators are chosen based on our analysis of LLM characteristics, and our extensive experiments show that this simple approach can reduce TTFT by 92% and 52%, and TPOT by 21% and 20%, compared to vLLM-v1 and a production scheduler on real-world workloads covering chatbots, API calls, and coding agents. We also mathematically derive the conditions under which multiplication may fail, and find that such conditions are extremely rare in practice and can be detected (and mitigated) beforehand.

Paper Structure (17 sections, 5 equations, 27 figures)

This paper contains 17 sections, 5 equations, 27 figures.

Introduction
LLM Serving and Scheduling
Background: Scheduling LLM requests in a Cluster
The Analysis Framework
Characterizing LLM Request Scheduling
Characterization Methodology
Load-balancing Alone is Insufficient for LLM
KV$-awareness vs. Load balancing: The Trade-off
The Case of Linear Combination
The Case of Filter-based Combination
The Case of Simulation-based Combination
Simple Multiplication May Be All You Need
The Choice of the Indicators
Benign and Failure Cases Analysis of Multiplication-based Scheduling Score
End-to-end Evaluation
...and 2 more sections

Figures (27)

Figure 1: An illustration of generating output tokens using an LLM and the two performance metrics: time to first token (TTFT) and time per output token (TPOT).
Figure 2: The system view of how an LLM serving instance serves requests and some direct system indicators that can be collected by the global scheduler. The detailed meaning of each indicator will be described upon the first usage. BS is the abbreviation for batch size.
Figure 3: (a) System view of a cluster LLM serving system, and (b) comparison of per-request serving time and routing cost.
Figure 4: The system architecture of the LMetric metric factory and its programming model for scheduling algorithms.
Figure 5: Our studied traces that cover major scenarios in powering LLM services.
...and 22 more figures

LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling

Abstract

LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling

Authors

Abstract

Table of Contents

Figures (27)