The Evolution of LLM Adoption in Industry Data Curation Practices

Crystal Qian; Michael Xieyang Liu; Emily Reif; Grady Simon; Nada Hussein; Nathan Clement; James Wexler; Carrie J. Cai; Michael Terry; Minsuk Kahng

The Evolution of LLM Adoption in Industry Data Curation Practices

Crystal Qian, Michael Xieyang Liu, Emily Reif, Grady Simon, Nada Hussein, Nathan Clement, James Wexler, Carrie J. Cai, Michael Terry, Minsuk Kahng

TL;DR

This study investigates how data practitioners in a large technology organization adopt LLMs for unstructured data curation, using a three-stage program—an exploratory survey ($N=84$), expert interviews ($N=10$), and a user study with two LLM-based design probes ($N=12$). The findings reveal a shift from heuristic, bottom-up data understanding to insights-first, top-down workflows supported by LLMs, complemented by the emergence of silver and super-golden datasets to improve labeling and evaluation. While early adoption was limited, the design-probe studies show potential productivity gains and broad appeal for spreadsheet- and notebook–integrated LLM tooling, alongside barriers related to reliability, scale, and unfamiliarity with new features. Collectively, the work points to a paradigm where LLMs augment data practitioners’ workflows, enabling more targeted, high-quality data curation and heralding future directions in tool development, governance, and multimodal data handling.

Abstract

As large language models (LLMs) grow increasingly adept at processing unstructured text data, they offer new opportunities to enhance data curation workflows. This paper explores the evolution of LLM adoption among practitioners at a large technology company, evaluating the impact of LLMs in data curation tasks through participants' perceptions, integration strategies, and reported usage scenarios. Through a series of surveys, interviews, and user studies, we provide a timely snapshot of how organizations are navigating a pivotal moment in LLM evolution. In Q2 2023, we conducted a survey to assess LLM adoption in industry for development tasks (N=84), and facilitated expert interviews to assess evolving data needs (N=10) in Q3 2023. In Q2 2024, we explored practitioners' current and anticipated LLM usage through a user study involving two LLM-based prototypes (N=12). While each study addressed distinct research goals, they revealed a broader narrative about evolving LLM usage in aggregate. We discovered an emerging shift in data understanding from heuristic-first, bottom-up approaches to insights-first, top-down workflows supported by LLMs. Furthermore, to respond to a more complex data landscape, data practitioners now supplement traditional subject-expert-created 'golden datasets' with LLM-generated 'silver' datasets and rigorously validated 'super golden' datasets curated by diverse experts. This research sheds light on the transformative role of LLMs in large-scale analysis of unstructured data and highlights opportunities for further tool development.

The Evolution of LLM Adoption in Industry Data Curation Practices

TL;DR

Abstract

The Evolution of LLM Adoption in Industry Data Curation Practices

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (3)