Table of Contents
Fetching ...

Repairing Tool Calls Using Post-tool Execution Reflection and RAG

Jason Tsay, Zidane Wright, Gaodan Fang, Kiran Kate, Saurabh Jha, Yara Rizk

TL;DR

This work tackles failures in LLM-driven tool calls, particularly semantic and contextual errors that only become evident after tool execution. It introduces a post-tool execution reflection system (RAG-Repair) that augments LLM reasoning with retrieval-augmented generation over official tool documentation and troubleshooting documents, focusing on kubectl in Kubernetes. Through a large-scale evaluation and targeted manual checks, RAG-Repair improves the likelihood that repaired commands execute successfully and better satisfy user queries for several models, with troubleshooting documents often providing more benefit than official docs. The approach demonstrates the value of integrating structured domain knowledge into error recovery for tool-using agents, and points to future work in extending the method to other tool domains and documentation sources.

Abstract

Agentic systems interact with external systems by calling tools such as Python functions, REST API endpoints, or command line tools such as kubectl in Kubernetes. These tool calls often fail for various syntactic and semantic reasons. Some less obvious semantic errors can only be identified and resolved after analyzing the tool's response. To repair these errors, we develop a post-tool execution reflection component that combines large language model (LLM)-based reflection with domain-specific retrieval-augmented generation (RAG) using documents describing both the specific tool being called and troubleshooting documents related to the tool. For this paper, we focus on the use case of the kubectl command line tool to manage Kubernetes, a platform for orchestrating cluster applications. Through a larger empirical study and a smaller manual evaluation, we find that our RAG-based reflection will repair kubectl commands such that they are both more likely to successfully execute (pass rate) for 55% of our models evaluated and 36% more likely to correctly answer the user query on average. We find that troubleshooting documents improve pass rate compared to official documentation by an average of 10%.

Repairing Tool Calls Using Post-tool Execution Reflection and RAG

TL;DR

This work tackles failures in LLM-driven tool calls, particularly semantic and contextual errors that only become evident after tool execution. It introduces a post-tool execution reflection system (RAG-Repair) that augments LLM reasoning with retrieval-augmented generation over official tool documentation and troubleshooting documents, focusing on kubectl in Kubernetes. Through a large-scale evaluation and targeted manual checks, RAG-Repair improves the likelihood that repaired commands execute successfully and better satisfy user queries for several models, with troubleshooting documents often providing more benefit than official docs. The approach demonstrates the value of integrating structured domain knowledge into error recovery for tool-using agents, and points to future work in extending the method to other tool domains and documentation sources.

Abstract

Agentic systems interact with external systems by calling tools such as Python functions, REST API endpoints, or command line tools such as kubectl in Kubernetes. These tool calls often fail for various syntactic and semantic reasons. Some less obvious semantic errors can only be identified and resolved after analyzing the tool's response. To repair these errors, we develop a post-tool execution reflection component that combines large language model (LLM)-based reflection with domain-specific retrieval-augmented generation (RAG) using documents describing both the specific tool being called and troubleshooting documents related to the tool. For this paper, we focus on the use case of the kubectl command line tool to manage Kubernetes, a platform for orchestrating cluster applications. Through a larger empirical study and a smaller manual evaluation, we find that our RAG-based reflection will repair kubectl commands such that they are both more likely to successfully execute (pass rate) for 55% of our models evaluated and 36% more likely to correctly answer the user query on average. We find that troubleshooting documents improve pass rate compared to official documentation by an average of 10%.
Paper Structure (16 sections, 3 figures, 5 tables)

This paper contains 16 sections, 3 figures, 5 tables.

Figures (3)

  • Figure 1: The process of RAG-based repair. On the left we have an agentic system that generates tool calls and executes them using some tool executor. The right side shows our RAG-Repair approach for post-tool execution reflection that uses the repair context to repair the failing tool cool.
  • Figure 2: Prompt for RAG-Repair
  • Figure 3: Prompt for the non-RAG baseline