I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
Giacomo Camposampiero, Michael Hersche, Roger Wattenhofer, Abu Sebastian, Abbas Rahimi
TL;DR
I-RAVEN-X is a symbolic benchmark designed to stress-test generalization and robustness in analogical and mathematical reasoning for large language and reasoning models. It extends the prior I-RAVEN with larger operand counts, broader attribute ranges, and simulated perceptual uncertainty across four axes (Productivity, Systematicity, confounder robustness, and distribution smoothing), while remaining fully symbolic. Empirical results show that large reasoning models (LRMs) outperform language models (LLMs) on longer reasoning chains and wider attribute ranges, and can do so with less prompt engineering, particularly in arithmetic tasks; however, LRMs remain significantly challenged by reasoning under uncertainty, struggling to explore multiple probabilistic outcomes simultaneously. The study provides a principled framework and dataset for evaluating end-to-end reasoning under perceptual uncertainty, with code and data publicly available for broader reuse and extension.
Abstract
We introduce I-RAVEN-X, a symbolic benchmark designed to evaluate generalization and robustness in analogical and mathematical reasoning for Large Language Models (LLMs) and Large Reasoning Models (LRMs). I-RAVEN-X extends I-RAVEN by increasing operand complexity, attribute range, and introducing perceptual uncertainty. Compared to LLMs, empirical results show that LRMs achieve improved productivity and systematicity on longer reasoning relations and wider attribute ranges, respectively. However, LRMs are still significantly challenged by reasoning under uncertainty and cannot effectively explore multiple probabilistic outcomes.
