MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Peng Ding; Jiading Fang; Peng Li; Kangrui Wang; Xiaochen Zhou; Mo Yu; Jing Li; Matthew R. Walter; Hongyuan Mei

MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Peng Ding, Jiading Fang, Peng Li, Kangrui Wang, Xiaochen Zhou, Mo Yu, Jing Li, Matthew R. Walter, Hongyuan Mei

TL;DR

MANGO introduces a text-based mapping and navigation benchmark for large language models by leveraging 53 Jericho mazes paired with hundreds of destination-finding and route-finding questions. It provides a rigorous evaluation program with answerable/easy labels and an emphasis on structured reasoning outputs, including imputed edges to extend the navigational graph beyond the walkthrough. Experimental results show GPT-4 outperforms other models but still struggles on hard questions and some mazes, while humans achieve high accuracy; analysis links maze properties to model performance and demonstrates downstream benefits in text-based navigation tasks. The benchmark offers a scalable platform for advancing LLM spatial reasoning, with data and code public to drive future research in mapping, navigation, and embodied task planning in text-only environments.

Abstract

Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks. In this paper, we propose MANGO, a benchmark to evaluate their capabilities to perform text-based mapping and navigation. Our benchmark includes 53 mazes taken from a suite of textgames: each maze is paired with a walkthrough that visits every location but does not cover all possible paths. The task is question-answering: for each maze, a large language model reads the walkthrough and answers hundreds of mapping and navigation questions such as "How should you go to Attic from West of House?" and "Where are we if we go north and east from Cellar?". Although these questions are easy to humans, it turns out that even GPT-4, the best-to-date language model, performs poorly at answering them. Further, our experiments suggest that a strong mapping and navigation ability would benefit large language models in performing relevant downstream tasks, such as playing textgames. Our MANGO benchmark will facilitate future research on methods that improve the mapping and navigation capabilities of language models. We host our leaderboard, data, code, and evaluation program at https://mango.ttic.edu and https://github.com/oaklight/mango/.

MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

TL;DR

Abstract

Paper Structure (34 sections, 9 figures, 14 tables)

This paper contains 34 sections, 9 figures, 14 tables.

Introduction
MANGO: A Benchmark for Text-Based Mapping and Navigation
Maze Collection: From Game Walkthroughs to Mazes
Generation of Question Skeletons: Traversing Mazes and Imputing Edges
Evaluation Program
Experiments
Experiment Setup
Main Results
Analysis of GPTs
What makes those mazes challenging?
Human Performance
Does Mapping and Navigation Ability Matter in Downstream Tasks?
Related Work
Conclusion
Benchmark Details
...and 19 more sections

Figures (9)

Figure 1: Map of Zork-I. Arrows denote the direction of travel during the walkthrough, while the reverse direction is unseen but may be possible. Note that it is a 3D map projected onto a 2D plane so up may not point upward in the 2D visualization (e.g., Rocky Ledge to Canyon View).
Figure 2: Success rates of the examined models on (\ref{['fig:main_df']}) DF and (\ref{['fig:main_rf']}) RF questions, averaged over all $53$ mazes. \ref{['app:results']} provides similar graphs (e.g., \ref{['fig:main_rea_acc']}) for other evaluation metrics.
Figure 3: Success rates of GPT-3.5 and GPT-4 broken down into individual games. \ref{['fig:gpt3vs4_reasoning']} in \ref{['app:exp']} provides a similar visualization of reasoning accuracy.
Figure 4: Playing minigames.
Figure 5: Success rates weighted by route length on DF (\ref{['fig:length_df']}) and RF (\ref{['fig:length_rf']}) questions, averaged over all $53$ mazes.
...and 4 more figures

MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

TL;DR

Abstract

MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Authors

TL;DR

Abstract

Table of Contents

Figures (9)