Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Ying-Chun Lin; Jennifer Neville; Jack W. Stokes; Longqi Yang; Tara Safavi; Mengting Wan; Scott Counts; Siddharth Suri; Reid Andersen; Xiaofeng Xu; Deepak Gupta; Sujay Kumar Jauhar; Xia Song; Georg Buscher; Saurabh Tiwary; Brent Hecht; Jaime Teevan

Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar, Xia Song, Georg Buscher, Saurabh Tiwary, Brent Hecht, Jaime Teevan

TL;DR

The paper introduces SPUR, a three-phase framework that leverages large language models to extract interpretable SAT/DSAT patterns from multi-turn conversations and summarize them into domain-specific rubrics. These rubrics guide a final USE estimation, producing interpretable scores while maintaining high accuracy in few-shot settings. Through extensive experiments on four diverse datasets, SPUR demonstrates superior performance to embedding-based baselines, shows rubric interpretability and cross-domain adaptability, and enables scalable deployment via knowledge distillation and rubric-as-features. The work offers a practical approach to interpretable USE for both general-purpose and task-oriented conversational systems, with clear implications for monitoring, auditing, and improving conversational AI deployments.

Abstract

Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational patterns in both general-purpose (ChatGPT and Bing Copilot) and task-oriented (customer service chatbot) conversational systems. Existing approaches based on featurized ML models or text embeddings fall short in extracting generalizable patterns and are hard to interpret. In this work, we show that LLMs can extract interpretable signals of user satisfaction from their natural language utterances more effectively than embedding-based approaches. Moreover, an LLM can be tailored for USE via an iterative prompting framework using supervision from labeled examples. The resulting method, Supervised Prompting for User satisfaction Rubrics (SPUR), not only has higher accuracy but is more interpretable as it scores user satisfaction via learned rubrics with a detailed breakdown.

Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

TL;DR

Abstract

Paper Structure (24 sections, 9 figures, 5 tables)

This paper contains 24 sections, 9 figures, 5 tables.

Introduction
Problem Definition and Related Work
SPUR
Supervised Extraction
Rubric Summarization
User Satisfaction Estimation
Evaluation
Baselines
Dataset
USE under Few-Shot Setting.
Importance of Rubric Summarization.
Rubric vs. Thumb Feedback.
Pattern Variance for Different Conversational Systems.
Knowledge Distillation.
Rubrics as Features
...and 9 more sections

Figures (9)

Figure 1: Illustration of user utterances with satisfaction patterns (green) and dissatisfaction patterns (red).
Figure 2: Illustration of SPUR approach. Step 1 corresponds to Sec. \ref{['sec:se']}, Step 2: Sec. \ref{['sec:rs']}, and Step 3: Sec. \ref{['sec:sat_score']}.
Figure 3: The average scores for each rubric item w.r.t. thumb feedback (Like or Dislike). The '*' beside each keyword indicates that the rubric item is significantly correlated with thumb feedback.
Figure 4: Satisfaction/Dissatisfaction Conversational Pattern Distributions.
Figure 5: ROC on Knowledge Distillation from GPT-4.
...and 4 more figures

Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

TL;DR

Abstract

Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Authors

TL;DR

Abstract

Table of Contents

Figures (9)