Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
Melik Ozolcer, Sang Won Bae
TL;DR
This work presents an evaluation-focused study of a web-deployed, tool-augmented LLM health coach using real users and wearables. It combines offline policy evaluation (OPE) of decisions over Tool/Style heads with a lightweight simulator that embeds hidden archetypes to test a brief early information-gain (curiosity) phase. Key findings show that a uniform heavy-tool policy can raise log-based averages but harm specific archetypes, while a small early-curiosity boost improves trait-identification speed and task success in simulation. The authors propose an evaluation-first, personalization-first framework: freeze the generator, learn archetype-aware decision heads with typed rewards, and surface per-archetype metrics to surface subgroup harms, enabling responsible deployment and future RL-based personalization.
Abstract
We study a web-deployed, tool-augmented LLM health coach with real users. In a pilot with seven users (280 rated turns), offline policy evaluation (OPE) over factorized decision heads (Tool/Style) shows that a uniform heavy-tool policy raises average value on logs but harms specific subgroups, most notably low-health-literacy/high-self-efficacy users. A lightweight simulator with hidden archetypes further shows that adding a small early information-gain bonus reliably shortens trait identification and improves goal success and pass@3. Together, these early findings indicate an evaluation-first path to personalization: freeze the generator, learn subgroup-aware decision heads on typed rewards (objective tool outcomes and satisfaction), and always report per-archetype metrics to surface subgroup harms that averages obscure.
