Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
Xuying LI
TL;DR
The paper addresses the lack of fine-grained controllability over internal reasoning in mathematical problem solving by introducing self-optimizing thought vectors whose selections are guided by entropy-based rewards $R = -\mathcal{H}(p)$. It presents an eight-vector architecture with a three-dimensional control framework (Depth, Length, Path) and demonstrates that entropy-driven optimization can produce focused, explainable reasoning patterns without external annotations, achieving GSM8K accuracy of 90.1% and a controllability score of 0.42 on Gemma-2-9B with LoRA. Empirical analyses reveal meaningful clustering of thought vectors and low-entropy distributions under control, along with robust controllability metrics and insightful case studies. The work suggests that internal reasoning can be steered through self-supervised entropy optimization, offering a path toward more transparent and adaptable AI reasoning across domains.
Abstract
We present a novel approach for controllable mathematical reasoning that leverages self-optimizing thought vectors with entropy minimization. Our method introduces learnable thought vectors that dynamically modulate the internal reasoning process of large language models. Using Gemma-2-9B on GSM8K, we achieve 90.1% accuracy with a controllability score of 0.42, demonstrating that entropy-based rewards effectively guide focused reasoning patterns without requiring external reward annotations. Our analysis reveals distinct thought vector clusters and consistent low-entropy distributions across control conditions, validating our framework for controllable AI reasoning.
