ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy
Kazuki Kawamura, Kengo Nakai, Jun Rekimoto
TL;DR
ManzaiSet addresses the Western-centric bias in affective computing by providing the first large-scale, multimodal dataset of viewer responses to culturally specific humor. The authors collected synchronized facial video and audio from 241 Japanese viewers as they watched 10 professional manzai performances, yielding 191.8 hours of data and enabling within-subject analyses (n=228). Three core findings emerge: (i) clear viewer typologies with trait-like stability and significant variance differences; (ii) a positive viewing-order effect, contradicting fatigue hypotheses; and (iii) no robust differences across nine humor types after multiple testing corrections. These results support culturally aware emotion AI and personalized entertainment systems that account for audience heterogeneity and temporal dynamics, with implications for cross-cultural emotion recognition and humor-enabled interaction design. The dataset and analyses offer a foundation for measuring moment-to-moment engagement, building audience-specific models, and exploring delivery and timing in non-Western comedic contexts.
Abstract
We present ManzaiSet, the first large scale multimodal dataset of viewer responses to Japanese manzai comedy, capturing facial videos and audio from 241 participants watching up to 10 professional performances in randomized order (94.6 percent watched >= 8; analyses focus on n=228). This addresses the Western centric bias in affective computing. Three key findings emerge: (1) k means clustering identified three distinct viewer types: High and Stable Appreciators (72.8 percent, n=166), Low and Variable Decliners (13.2 percent, n=30), and Variable Improvers (14.0 percent, n=32), with heterogeneity of variance (Brown Forsythe p < 0.001); (2) individual level analysis revealed a positive viewing order effect (mean slope = 0.488, t(227) = 5.42, p < 0.001, permutation p < 0.001), contradicting fatigue hypotheses; (3) automated humor classification (77 instances, 131 labels) plus viewer level response modeling found no type wise differences after FDR correction. The dataset enables culturally aware emotion AI development and personalized entertainment systems tailored to non Western contexts.
