Table of Contents
Fetching ...

User Perceptions of Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

Xiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer, Koichi Onoue

TL;DR

This work addresses how users perceive the privacy-preservation quality and usefulness of LLM responses in privacy-sensitive scenarios and whether proxy LLMs accurately reflect those perceptions. It combines a user study (94 participants, 90 PrivacyLens scenarios) with evaluations from five proxy LLMs, each running five times per scenario, to quantify agreement and misalignment. Key findings show that while individual users often view responses as helpful and privacy-preserving, judgments are highly divergent across participants for the same scenario, whereas proxy LLMs are internally consistent and moderately consistent with each other but only weakly to moderately aligned with human judgments ($0.24 \le \rho \le 0.68$). The qualitative analysis reveals that proxy LLMs can miss contextual cues and diverge on privacy norms and preferences, underscoring the necessity of human-centered evaluation in privacy-sensitive, utility-driven tasks. The paper concludes with recommendations to improve proxy evaluations via output diversity and personalization and advocates clear taxonomies that separate objective versus subjective tasks to better guide evaluation practices and interpretation.

Abstract

Large language models (LLMs) have seen rapid adoption for tasks such as drafting emails, summarizing meetings, and answering health questions. In such uses, users may need to share private information (e.g., health records, contact details). To evaluate LLMs' ability to identify and redact such private information, prior work developed benchmarks (e.g., ConfAIde, PrivacyLens) with real-life scenarios. Using these benchmarks, researchers have found that LLMs sometimes fail to keep secrets private when responding to complex tasks (e.g., leaking employee salaries in meeting summaries). However, these evaluations rely on LLMs (proxy LLMs) to gauge compliance with privacy norms, overlooking real users' perceptions. Moreover, prior work primarily focused on the privacy-preservation quality of responses, without investigating nuanced differences in helpfulness. To understand how users perceive the privacy-preservation quality and helpfulness of LLM responses to privacy-sensitive scenarios, we conducted a user study with 94 participants using 90 scenarios from PrivacyLens. We found that, when evaluating identical responses to the same scenario, users showed low agreement with each other on the privacy-preservation quality and helpfulness of the LLM response. Further, we found high agreement among five proxy LLMs, while each individual LLM had low correlation with users' evaluations. These results indicate that the privacy and helpfulness of LLM responses are often specific to individuals, and proxy LLMs are poor estimates of how real users would perceive these responses in privacy-sensitive scenarios. Our results suggest the need to conduct user-centered studies on measuring LLMs' ability to help users while preserving privacy. Additionally, future research could investigate ways to improve the alignment between proxy LLMs and users for better estimation of users' perceived privacy and utility.

User Perceptions of Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

TL;DR

This work addresses how users perceive the privacy-preservation quality and usefulness of LLM responses in privacy-sensitive scenarios and whether proxy LLMs accurately reflect those perceptions. It combines a user study (94 participants, 90 PrivacyLens scenarios) with evaluations from five proxy LLMs, each running five times per scenario, to quantify agreement and misalignment. Key findings show that while individual users often view responses as helpful and privacy-preserving, judgments are highly divergent across participants for the same scenario, whereas proxy LLMs are internally consistent and moderately consistent with each other but only weakly to moderately aligned with human judgments (). The qualitative analysis reveals that proxy LLMs can miss contextual cues and diverge on privacy norms and preferences, underscoring the necessity of human-centered evaluation in privacy-sensitive, utility-driven tasks. The paper concludes with recommendations to improve proxy evaluations via output diversity and personalization and advocates clear taxonomies that separate objective versus subjective tasks to better guide evaluation practices and interpretation.

Abstract

Large language models (LLMs) have seen rapid adoption for tasks such as drafting emails, summarizing meetings, and answering health questions. In such uses, users may need to share private information (e.g., health records, contact details). To evaluate LLMs' ability to identify and redact such private information, prior work developed benchmarks (e.g., ConfAIde, PrivacyLens) with real-life scenarios. Using these benchmarks, researchers have found that LLMs sometimes fail to keep secrets private when responding to complex tasks (e.g., leaking employee salaries in meeting summaries). However, these evaluations rely on LLMs (proxy LLMs) to gauge compliance with privacy norms, overlooking real users' perceptions. Moreover, prior work primarily focused on the privacy-preservation quality of responses, without investigating nuanced differences in helpfulness. To understand how users perceive the privacy-preservation quality and helpfulness of LLM responses to privacy-sensitive scenarios, we conducted a user study with 94 participants using 90 scenarios from PrivacyLens. We found that, when evaluating identical responses to the same scenario, users showed low agreement with each other on the privacy-preservation quality and helpfulness of the LLM response. Further, we found high agreement among five proxy LLMs, while each individual LLM had low correlation with users' evaluations. These results indicate that the privacy and helpfulness of LLM responses are often specific to individuals, and proxy LLMs are poor estimates of how real users would perceive these responses in privacy-sensitive scenarios. Our results suggest the need to conduct user-centered studies on measuring LLMs' ability to help users while preserving privacy. Additionally, future research could investigate ways to improve the alignment between proxy LLMs and users for better estimation of users' perceived privacy and utility.
Paper Structure (34 sections, 6 figures, 3 tables)

This paper contains 34 sections, 6 figures, 3 tables.

Figures (6)

  • Figure 1: Overview of our study design.
  • Figure 2: Participants found LLM-generated responses completed the given task over 90% of the time. Participants found LLM-generated responses to be helpful and would use the response most of the time.
  • Figure 3: 78% of the time, participants found the LLM-generated response to comply with the privacy norm. Over 82% of the time, participants found the response respects their personal privacy preferences.
  • Figure 4: For each scenario, we show the range of judgements collected from participants and proxy LLMs. We compute the range across the $\geq 5$ participants and $5$ runs per proxy LLM for each scenario. Completely agree means every rating is the same on the five-point likert scale. Two points apart means one rating is strongly agree while another is neutral with the remaining participants selecting something in between.
  • Figure 5: With cumulative density curves, we show the cumulative percentage of scenarios (x-axis) with a certain standard deviation (y-axis) of the likert scale evaluation. Proxy LLMs show a lower standard deviation than participants' evaluations on more scenarios. However, the differences vary by proxy LLM and the question asked.
  • ...and 1 more figures