SocialQuotes: Learning Contextual Roles of Social Media Quotes on the Web
John Palowitch, Hamidreza Alvari, Mehran Kazemi, Tanvir Amin, Filip Radlinski
TL;DR
This work addresses the challenge of understanding how social media quotes function within web pages by proposing a formal framework that treats embedded posts as quotes with latent roles. It introduces an eight-role taxonomy and a four-stage model of social quotation, and builds SocialQuotes, a large-scale dataset derived from Common Crawl with topic annotations and 8.3k human-annotated roles across 32.6M quotes. The authors demonstrate that a state-of-the-art LLM can infer quote roles from page context, with improvements from few-shot, chain-of-thought, self-consistency, and persistence techniques, and they analyze cross-domain distributions of platforms, domains, and roles to reveal culture- vs. reporting-oriented patterns. The dataset and methodology pave the way for cross-platform retrieval, off-API analyses, and potential generative citation, enabling richer, context-aware studies of the web's social media landscape.
Abstract
Web authors frequently embed social media to support and enrich their content, creating the potential to derive web-based, cross-platform social media representations that can enable more effective social media retrieval systems and richer scientific analyses. As step toward such capabilities, we introduce a novel language modeling framework that enables automatic annotation of roles that social media entities play in their embedded web context. Using related communication theory, we liken social media embeddings to quotes, formalize the page context as structured natural language signals, and identify a taxonomy of roles for quotes within the page context. We release SocialQuotes, a new data set built from the Common Crawl of over 32 million social quotes, 8.3k of them with crowdsourced quote annotations. Using SocialQuotes and the accompanying annotations, we provide a role classification case study, showing reasonable performance with modern-day LLMs, and exposing explainable aspects of our framework via page content ablations. We also classify a large batch of un-annotated quotes, revealing interesting cross-domain, cross-platform role distributions on the web.
