Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
Atharvan Dogra, Soumya Suvra Ghosal, Ameet Deshpande, Ashwin Kalyan, Dinesh Manocha
TL;DR
This study investigates safety risks in humor generation by large language systems, showing that optimizing for funniness can amplify harmful content through stereotypes and toxicity. By jointly evaluating humor, stereotypes, toxicity, and incongruity across six LLMs and role-based prompts, the authors reveal a bias amplification loop where both generators and evaluators reward edgier humor. Information-theoretic analyses indicate that harmful cues expand the space of plausible punchlines and can become more expected for some systems, underscoring structural embedding in learned humor distributions. External satire tasks and human judgments corroborate that satire often increases stereotyping and toxicity, highlighting the need for multi-objective generation and evaluation to balance engagement with safety.
Abstract
Large language models are increasingly used for creative writing and engagement content, raising safety concerns about the outputs. Therefore, casting humor generation as a testbed, this work evaluates how funniness optimization in modern LLM pipelines couples with harmful content by jointly measuring humor, stereotypicality, and toxicity. This is further supplemented by analyzing incongruity signals through information-theoretic metrics. Across six models, we observe that harmful outputs receive higher humor scores which further increase under role-based prompting, indicating a bias amplification loop between generators and evaluators. Information-theoretic analyses show harmful cues widen predictive uncertainty and surprisingly, can even make harmful punchlines more expected for some models, suggesting structural embedding in learned humor distributions. External validation on an additional satire-generation task with human perceived funniness judgments shows that LLM satire increases stereotypicality and typically toxicity, including for closed models. Quantitatively, stereotypical/toxic jokes gain $10-21\%$ in mean humor score, stereotypical jokes appear $11\%$ to $28\%$ more often among the jokes marked funny by LLM-based metric and up to $10\%$ more often in generations perceived as funny by humans.
