Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning

Liang Qiao; Jun Shi; Xiaoyu Hao; Xi Fang; Sen Zhang; Minfan Zhao; Ziqi Zhu; Junshi Chen; Hong An; Xulong Tang; Bing Li; Honghui Yuan; Xinyang Wang

Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning

Liang Qiao, Jun Shi, Xiaoyu Hao, Xi Fang, Sen Zhang, Minfan Zhao, Ziqi Zhu, Junshi Chen, Hong An, Xulong Tang, Bing Li, Honghui Yuan, Xinyang Wang

TL;DR

This work tackles the inefficiency of search-based tensor program tuning caused by slow learned cost models and cross-platform online unawareness. It introduces Pruner, a Draft-then-Verify exploration mechanism with a Latent Schedule Explorer for rapid draft generation and a Pattern-aware Cost Model for accurate verification, complemented by MoA-Pruner, a momentum online adaptation strategy for cross-platform transfer. Across three GPU platforms, the approach achieves substantial speedups over state-of-the-art baselines in both online and offline tuning (e.g., average online speedups of $2.6\times$ for Pruner and $4.82\times$ for MoA-Pruner vs Ansor, and offline gains around $4.75\times$–$4.05\times$ vs TenSet/TLP with an additional $4.08\times$ over MetaSchedule on TensorCore). Implemented in the TVM framework, Pruner demonstrates strong end-to-end performance gains, robust single-operator results, and favorable compilation costs, highlighting its practical impact for efficient tensor-program tuning.

Abstract

Tensor program tuning is essential for the efficient deployment of deep neural networks. Search-based approaches have demonstrated scalability and effectiveness in automatically finding high-performance programs for specific hardware. However, the search process is often inefficient, taking hours or even days to discover optimal programs due to the exploration mechanisms guided by an accurate but slow-learned cost model. Meanwhile, the learned cost model trained on one platform cannot seamlessly adapt online to another, which we call cross-platform online unawareness. In this work, we propose Pruner and MoA-Pruner. Pruner is a "Draft-then-Verify" exploration mechanism that accelerates the schedule search process. Instead of applying the complex learned cost model to all explored candidates, Pruner drafts small-scale potential candidates by introducing a naive Symbol-based Analyzer (draft model), then identifies the best candidates by the learned cost model. MoA-Pruner introduces a Momentum online Adaptation strategy to address the cross-platform online unawareness. We incorporate Pruner into the TVM and conduct extensive experiments on three GPU-based platforms. Results show considerable speedup in schedule search time. In online tuning scenarios, Pruner and MoA-Pruner achieve an average speedup of $2.6 \times$ and $4.82 \times$ compared to Ansor. In offline tuning scenarios, Pruner achieves an average speedup of $4.75 \times$ and $4.05\times$ compared to TenSet and TLP, respectively. Furthermore, Pruner achieves an average speedup of $4.08 \times$ compared to MetaSchedule on TensorCore.

Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning

TL;DR

for Pruner and

for MoA-Pruner vs Ansor, and offline gains around

–

vs TenSet/TLP with an additional

over MetaSchedule on TensorCore). Implemented in the TVM framework, Pruner demonstrates strong end-to-end performance gains, robust single-operator results, and favorable compilation costs, highlighting its practical impact for efficient tensor-program tuning.

Abstract

and

compared to Ansor. In offline tuning scenarios, Pruner achieves an average speedup of

and

compared to TenSet and TLP, respectively. Furthermore, Pruner achieves an average speedup of

compared to MetaSchedule on TensorCore.

Paper Structure (20 sections, 3 equations, 16 figures, 13 tables, 2 algorithms)

This paper contains 20 sections, 3 equations, 16 figures, 13 tables, 2 algorithms.

Introduction
Background
Search-based deep learning compilers
Learned cost models and cross-platform transfer
Opportunities
System Design
Pruner
Draft: Latent Schedule Explorer
Verify: Pattern-aware Cost Model
MoA-Pruner
Experimental Settings
Evaluation
End-to-End Workload Benchmark
Single Operator Performance
Compilation Cost
...and 5 more sections

Figures (16)

Figure 1: The workflow of search-based DLCs. The red dashed box is the optimization workspace of the Pruner.
Figure 2: The system overview of Pruner, which contains latent schedule explorer and pattern-aware cost model. Momentum online adaptation is only activated for MoA-Pruner in online cost model tuning scenarios.
Figure 3: An illustrative example of the hardware-aware symbols extraction process for a GEMM-ReLU graph. Some schedule primitives in the schedules generation rules and hardware symbols for some statements are omitted for brevity. Prod or Sum means that the variable equals the product or sum of an array of variables.
Figure 4: (top) The pipeline of Pattern-aware Cost Model; (bottom) Extraction of temporal dataflow feature.
Figure 5: Overview of MoA, where $\mathcal{\phi}_{s}$ and ${\Delta}_{t}$ refer to the parameters of the Siamese and gradient of target model, the momentum $m=0.99$. The red arrow means offline pre-train.
...and 11 more figures

Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning

TL;DR

Abstract

Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning

Authors

TL;DR

Abstract

Table of Contents

Figures (16)