Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models

Yitian Li; Jidong Tian; Hao He; Yaohui Jin

Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models

Yitian Li, Jidong Tian, Hao He, Yaohui Jin

TL;DR

The paper addresses the challenge that existing prompting strategies can induce invalid or non-robust reasoning paths in large language models when solving deductive problems. It proposes Hypothesis Testing Prompting, which integrates conclusion assumptions, backward reasoning, and fact verification to guide intermediate reasoning toward correct conclusions. Empirical evaluation on RuleTaker (CWA) and ProofWriter (OWA) shows significant improvements in accuracy and the generation of more reasonable reasoning traces, including better handling of unknown conclusions. This prompting strategy offers a generalizable approach to enhance deductive reasoning in LLMs with potential applicability to broader reasoning tasks.

Abstract

Combining different forms of prompts with pre-trained large language models has yielded remarkable results on reasoning tasks (e.g. Chain-of-Thought prompting). However, along with testing on more complex reasoning, these methods also expose problems such as invalid reasoning and fictional reasoning paths. In this paper, we develop \textit{Hypothesis Testing Prompting}, which adds conclusion assumptions, backward reasoning, and fact verification during intermediate reasoning steps. \textit{Hypothesis Testing prompting} involves multiple assumptions and reverses validation of conclusions leading to its unique correct answer. Experiments on two challenging deductive reasoning datasets ProofWriter and RuleTaker show that hypothesis testing prompting not only significantly improves the effect, but also generates a more reasonable and standardized reasoning process.

Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models

TL;DR

Abstract

Paper Structure (11 sections, 4 figures)

This paper contains 11 sections, 4 figures.

Introduction
Related Work
Few-Shot Prompting
Deductive Reasoning
Hypothesis Testing Prompting
Experiment
Experimental Setup
Experimental Results
Further Analysis
Conclusion
References

Figures (4)

Figure 1: Questions in RuleTaker involve logical reasoning with facts and rules.
Figure 2: Comparison of three prompting methods: (a) Standard (b) Chain-of-Thought (c) Hypothesis Testing. Particularly, we highlight the Hypothesis testing reasoning processes. The comparative experimental results show that: Hypothesis testing prompting enables large language models to tackle complex logical reasoning.
Figure 3: Prediction accuracy results on the (a) RuleTaker and (b) ProofWriter datasets.
Figure 4: Further results on ProofWriter.

Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models

TL;DR

Abstract

Hypothesis Testing Prompting Improves Deductive Reasoning in Large Language Models

Authors

TL;DR

Abstract

Table of Contents

Figures (4)