Molly: Making Large Language Model Agents Solve Python Problem More Logically

Rui Xiao; Jiong Wang; Lu Han; Na Zong; Han Wu

Molly: Making Large Language Model Agents Solve Python Problem More Logically

Rui Xiao, Jiong Wang, Lu Han, Na Zong, Han Wu

TL;DR

This work addresses the challenge of making large language models act as effective programming teaching assistants for Chinese Python learners. It introduces Molly, a three-stage LLM agent that combines scenario-based intent detection, knowledge retrieval from a structured educational knowledge base, and iterative self-reflection to improve answer accuracy, expressiveness, and usefulness. A Chinese Python QA dataset (5,960 QA pairs, built from 16,247 questions) with expert annotations underpins the knowledge base, enabling pedagogy-aligned responses. Across multiple LLMs, Molly improves teaching-oriented outcomes, demonstrating potential for scalable, dialog-based programming education.

Abstract

Applying large language models (LLMs) as teaching assists has attracted much attention as an integral part of intelligent education, particularly in computing courses. To reduce the gap between the LLMs and the computer programming education expert, fine-tuning and retrieval augmented generation (RAG) are the two mainstream methods in existing researches. However, fine-tuning for specific tasks is resource-intensive and may diminish the model`s generalization capabilities. RAG can perform well on reducing the illusion of LLMs, but the generation of irrelevant factual content during reasoning can cause significant confusion for learners. To address these problems, we introduce the Molly agent, focusing on solving the proposed problem encountered by learners when learning Python programming language. Our agent automatically parse the learners' questioning intent through a scenario-based interaction, enabling precise retrieval of relevant documents from the constructed knowledge base. At generation stage, the agent reflect on the generated responses to ensure that they not only align with factual content but also effectively answer the user's queries. Extensive experimentation on a constructed Chinese Python QA dataset shows the effectiveness of the Molly agent, indicating an enhancement in its performance for providing useful responses to Python questions.

Molly: Making Large Language Model Agents Solve Python Problem More Logically

TL;DR

Abstract

Molly: Making Large Language Model Agents Solve Python Problem More Logically

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (5)