Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution

Siqi Wang; Hailong Yang; Xuezhu Wang; Tongxuan Liu; Pengbo Wang; Xuning Liang; Kejie Ma; Tianyu Feng; Xin You; Yongjun Bao; Yi Liu; Zhongzhi Luan; Depei Qian

Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution

Siqi Wang, Hailong Yang, Xuezhu Wang, Tongxuan Liu, Pengbo Wang, Xuning Liang, Kejie Ma, Tianyu Feng, Xin You, Yongjun Bao, Yi Liu, Zhongzhi Luan, Depei Qian

TL;DR

This paper proposes Minions, an LLM inference system that accelerates LLM inference with a collective and adaptive speculative generation and proposes an adaptive mechanism to dynamically determine the optimal speculation length of SSM, which can achieve better inference performance across different models, datasets, and hyper-parameters.

Abstract

Large language models (LLM) have recently attracted surging interest due to their outstanding capabilities across various domains. However, enabling efficient LLM inference is challenging due to its autoregressive decoding that generates tokens only one at a time. Although research works apply pruning or quantization to speed up LLM inference, they typically require fine-tuning the LLM, incurring significant time and economic costs. Meanwhile, speculative decoding has been proposed to use small speculative models (SSMs) to accelerate the inference of LLM. However, the low acceptance rate of SSM and the high verification cost of LLM prohibit further performance improvement of inference. In this paper, we propose Minions, an LLM inference system that accelerates LLM inference with a collective and adaptive speculative generation. Specifically, Minions proposes a majority-voted mechanism to leverage multiple SSMs to jointly speculate the outputs of LLM, which improves the inference performance without introducing prohibitive computation costs for LLM. To better trade off the number of tokens speculated from SSM and the verification cost of LLM, Minions proposes an adaptive mechanism to dynamically determine the optimal speculation length of SSM, which can achieve better inference performance across different models, datasets, and hyper-parameters. In addition, Minions decouples the SSM decoding and LLM verification efficiently and adopts a pipelined execution mechanism to further improve the inference performance of LLM. By comparing with the state-of-the-art LLM inference systems, we demonstrate that Minions can achieve higher inference throughput and lower inference time.

Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution

TL;DR

Abstract

Paper Structure (21 sections, 4 equations, 14 figures, 1 table)

This paper contains 21 sections, 4 equations, 14 figures, 1 table.

Introduction
Background and Motivation
Speculative Decoding
Model Capability Gap Between SSM and LLM
Vast Search Space for Determining Optimal Speculation Length
Tightly-Coupled Speculative Execution Hinders Inference Performance
Methodology
Design Overview
Majority-voted Speculator
Adaptive Speculation Length Selector
Speculative Generation Pipeline
Implementation
Evaluation
Experiment Setup
End-to-end Performance
...and 6 more sections

Figures (14)

Figure 1: The acceptance rate of models with different speculation lengths and datasets.
Figure 2: The inference time of the LLM under different speculation lengths, datasets, and batch size.
Figure 3: Performance potential revealed by pipelined execution. $T$ represents the reduced inference time.
Figure 4: Design overview of Minions
Figure 5: The primary workflow of Majority-voted Speculator.
...and 9 more figures

Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution

TL;DR

Abstract

Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution

Authors

TL;DR

Abstract

Table of Contents

Figures (14)