Table of Contents
Fetching ...

TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model

Bin Yu, Xinming Wang, Shijie Lian, Haotian Li, Changti Wu, Ruina Hu, Bailing Wang, Yuliang Wei, Kai Chen

TL;DR

TrajSelector introduces a lightweight process verifier that leverages the frozen sampler's latent hidden states to score stepwise reasoning in Best-of-$N$ external test-time scaling. By using a $0.6$B verifier and a data-driven training scheme with pseudo-labels and a buffer class, it eliminates the need for large PRMs and step-level annotations. Empirical results across multiple math benchmarks demonstrate consistent improvements, especially in Best-of-$32$, while maintaining substantially lower inference costs. The work highlights the practical value of latent, self-reflective signals for efficient external TTS in large reasoning models.

Abstract

Large language models (LLMs) have shown remarkable progress in complex reasoning tasks, largely enabled by test-time scaling (TTS) paradigms that allocate additional compute during inference. Among these, external TTS (particularly the Best-of-N selection paradigm) yields scalable performance improvements by selecting from multiple independently generated reasoning trajectories. However, this approach faces key limitations: (i) the high computational overhead of deploying process reward models, (ii) the underutilization of the LLM's intrinsic latent representations. We introduce TrajSelector, an efficient and effective Best-of-N framework that exploit the hidden states in the sampler LLM for process-level scoring. A lightweight verifier (with only 0.6B parameters) evaluates the quality of step-wise trajectory, and then aggregates these scores to identify the optimal reasoning trajectory. Our framework employs a fully data-driven, end-to-end training recipe that eliminates reliance on massive step-level annotations. Experiential results across five benchmarks demonstrate that TrajSelector delivers consistent performance gains. In Best-of-32 settings, it surpasses majority voting by 4.61% accuracy and outperforms existing process reward models by 4.31% to 12.21%, all while maintaining lower inference costs.

TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model

TL;DR

TrajSelector introduces a lightweight process verifier that leverages the frozen sampler's latent hidden states to score stepwise reasoning in Best-of- external test-time scaling. By using a B verifier and a data-driven training scheme with pseudo-labels and a buffer class, it eliminates the need for large PRMs and step-level annotations. Empirical results across multiple math benchmarks demonstrate consistent improvements, especially in Best-of-, while maintaining substantially lower inference costs. The work highlights the practical value of latent, self-reflective signals for efficient external TTS in large reasoning models.

Abstract

Large language models (LLMs) have shown remarkable progress in complex reasoning tasks, largely enabled by test-time scaling (TTS) paradigms that allocate additional compute during inference. Among these, external TTS (particularly the Best-of-N selection paradigm) yields scalable performance improvements by selecting from multiple independently generated reasoning trajectories. However, this approach faces key limitations: (i) the high computational overhead of deploying process reward models, (ii) the underutilization of the LLM's intrinsic latent representations. We introduce TrajSelector, an efficient and effective Best-of-N framework that exploit the hidden states in the sampler LLM for process-level scoring. A lightweight verifier (with only 0.6B parameters) evaluates the quality of step-wise trajectory, and then aggregates these scores to identify the optimal reasoning trajectory. Our framework employs a fully data-driven, end-to-end training recipe that eliminates reliance on massive step-level annotations. Experiential results across five benchmarks demonstrate that TrajSelector delivers consistent performance gains. In Best-of-32 settings, it surpasses majority voting by 4.61% accuracy and outperforms existing process reward models by 4.31% to 12.21%, all while maintaining lower inference costs.
Paper Structure (24 sections, 6 equations, 6 figures, 2 tables)

This paper contains 24 sections, 6 equations, 6 figures, 2 tables.

Figures (6)

  • Figure 1: Best-of-$N$ selection method illustration.
  • Figure 2: Best-of-$N$ Scaling Curve. The y-axis represents the average accuracy of different methods across the 5 benchmarks in the experimental section. TrajSelector achieves robust performance improvement as $N$ increases.
  • Figure 3: Overview Architecture
  • Figure 4: PyTorch style code of TrajSelector.
  • Figure 5: Ablation Study on Loss Function
  • ...and 1 more figures