Large Language Models Think Too Fast To Explore Effectively

Lan Pan; Hanbo Xie; Robert C. Wilson

Large Language Models Think Too Fast To Explore Effectively

Lan Pan, Hanbo Xie, Robert C. Wilson

TL;DR

The paper investigates how large language models explore in open-ended tasks using Little Alchemy 2 as a testbed. It contrasts uncertainty-driven exploration with empowerment-based exploration, employing regression analyses, thought tracing, and sparse autoencoder mappings to understand internal representations. The findings show most LLMs underperform humans, except for o1, while DeepSeek-R1 achieves human-like exploration through deeper, iterative reasoning. The work highlights a temporal mismatch in how LLMs represent empowerment and uncertainty, and argues for architecture- and training-level changes to enable more effective exploration with practical implications for adaptive AI systems.

Abstract

Large Language Models (LLMs) have emerged with many intellectual capacities. While numerous benchmarks assess their intelligence, limited attention has been given to their ability to explore--an essential capacity for discovering new information and adapting to novel environments in both natural and artificial systems. The extent to which LLMs can effectively explore, particularly in open-ended tasks, remains unclear. This study investigates whether LLMs can surpass humans in exploration during an open-ended task, using Little Alchemy 2 as a paradigm, where agents combine elements to discover new ones. Results show most LLMs underperform compared to humans, except for the o1 model, with traditional LLMs relying primarily on uncertainty-driven strategies, unlike humans who balance uncertainty and empowerment. Results indicate that traditional reasoning-focused LLMs, such as GPT-4o, exhibit a significantly faster and less detailed reasoning process, limiting their exploratory performance. In contrast, the DeepSeek reasoning model demonstrates prolonged, iterative thought processes marked by repetitive analysis of combinations and past trials, reflecting a more thorough and human-like exploration strategy. Representational analysis of the models with Sparse Autoencoders (SAE) revealed that uncertainty and choices are represented at earlier transformer blocks, while empowerment values are processed later, causing LLMs to think too fast and make premature decisions, hindering effective exploration. These findings shed light on the limitations of LLM exploration and suggest directions for improving their adaptability.

Large Language Models Think Too Fast To Explore Effectively

TL;DR

Abstract

Large Language Models Think Too Fast To Explore Effectively

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (16)