Table of Contents
Fetching ...

ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy

Kazuki Kawamura, Kengo Nakai, Jun Rekimoto

TL;DR

ManzaiSet addresses the Western-centric bias in affective computing by providing the first large-scale, multimodal dataset of viewer responses to culturally specific humor. The authors collected synchronized facial video and audio from 241 Japanese viewers as they watched 10 professional manzai performances, yielding 191.8 hours of data and enabling within-subject analyses (n=228). Three core findings emerge: (i) clear viewer typologies with trait-like stability and significant variance differences; (ii) a positive viewing-order effect, contradicting fatigue hypotheses; and (iii) no robust differences across nine humor types after multiple testing corrections. These results support culturally aware emotion AI and personalized entertainment systems that account for audience heterogeneity and temporal dynamics, with implications for cross-cultural emotion recognition and humor-enabled interaction design. The dataset and analyses offer a foundation for measuring moment-to-moment engagement, building audience-specific models, and exploring delivery and timing in non-Western comedic contexts.

Abstract

We present ManzaiSet, the first large scale multimodal dataset of viewer responses to Japanese manzai comedy, capturing facial videos and audio from 241 participants watching up to 10 professional performances in randomized order (94.6 percent watched >= 8; analyses focus on n=228). This addresses the Western centric bias in affective computing. Three key findings emerge: (1) k means clustering identified three distinct viewer types: High and Stable Appreciators (72.8 percent, n=166), Low and Variable Decliners (13.2 percent, n=30), and Variable Improvers (14.0 percent, n=32), with heterogeneity of variance (Brown Forsythe p < 0.001); (2) individual level analysis revealed a positive viewing order effect (mean slope = 0.488, t(227) = 5.42, p < 0.001, permutation p < 0.001), contradicting fatigue hypotheses; (3) automated humor classification (77 instances, 131 labels) plus viewer level response modeling found no type wise differences after FDR correction. The dataset enables culturally aware emotion AI development and personalized entertainment systems tailored to non Western contexts.

ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy

TL;DR

ManzaiSet addresses the Western-centric bias in affective computing by providing the first large-scale, multimodal dataset of viewer responses to culturally specific humor. The authors collected synchronized facial video and audio from 241 Japanese viewers as they watched 10 professional manzai performances, yielding 191.8 hours of data and enabling within-subject analyses (n=228). Three core findings emerge: (i) clear viewer typologies with trait-like stability and significant variance differences; (ii) a positive viewing-order effect, contradicting fatigue hypotheses; and (iii) no robust differences across nine humor types after multiple testing corrections. These results support culturally aware emotion AI and personalized entertainment systems that account for audience heterogeneity and temporal dynamics, with implications for cross-cultural emotion recognition and humor-enabled interaction design. The dataset and analyses offer a foundation for measuring moment-to-moment engagement, building audience-specific models, and exploring delivery and timing in non-Western comedic contexts.

Abstract

We present ManzaiSet, the first large scale multimodal dataset of viewer responses to Japanese manzai comedy, capturing facial videos and audio from 241 participants watching up to 10 professional performances in randomized order (94.6 percent watched >= 8; analyses focus on n=228). This addresses the Western centric bias in affective computing. Three key findings emerge: (1) k means clustering identified three distinct viewer types: High and Stable Appreciators (72.8 percent, n=166), Low and Variable Decliners (13.2 percent, n=30), and Variable Improvers (14.0 percent, n=32), with heterogeneity of variance (Brown Forsythe p < 0.001); (2) individual level analysis revealed a positive viewing order effect (mean slope = 0.488, t(227) = 5.42, p < 0.001, permutation p < 0.001), contradicting fatigue hypotheses; (3) automated humor classification (77 instances, 131 labels) plus viewer level response modeling found no type wise differences after FDR correction. The dataset enables culturally aware emotion AI development and personalized entertainment systems tailored to non Western contexts.
Paper Structure (41 sections, 2 equations, 5 figures, 4 tables)

This paper contains 41 sections, 2 equations, 5 figures, 4 tables.

Figures (5)

  • Figure 1: Data collection setup: participants watching manzai comedy videos at home via web browser.
  • Figure 2: Integrated temporal analysis combining individual temporal patterns with collective response dynamics. Four distinct temporal response patterns are identified: flat (48.7%), mixed (40.6%), volatile (10.0%), and ascending (0.8%). "Ascending" is defined conservatively as monotonic non-decreasing with a minimum slope threshold; many participants within the "mixed" class still exhibit net-positive linear slopes, consistent with the positive viewing-order effect reported in the main text. Percentages may not sum to 100 due to rounding.
  • Figure 3: Distribution of humor types across 77 instances in 10 manzai videos (multi-label classification). Unexpected twists (igai-sei) were most common (45.5% of instances), followed by repetition (tendon, 36.4%) and exaggeration (oogesa, 24.7%). Percentages sum to >100% as instances could have multiple labels. GPT-5-mini achieved 67.4% average confidence.
  • Figure 4: Effectiveness comparison of different humor types measured by instant and cumulative response rates with 95% bootstrap confidence intervals. Physical comedy (karada-gag) showed the highest cumulative response rate (26.2%, 95% CI: 11.3-41.0%), though with wide uncertainty due to small sample size (n=5). After FDR correction, no humor type achieved statistical significance over others (all adjusted p$>$0.05), with effect sizes remaining small (OR $<$ 1.5).
  • Figure 5: Comprehensive statistical validation results. (a) Pairwise comparisons with FDR correction showing no significant differences. (b) Permutation test p-value distribution confirming null findings. (c) Effect size matrix with OR values all below 1.5 (small effect threshold). (d) Bootstrap confidence interval overlap demonstrating minimal separation between humor types.