GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition

Zheng Lian; Licai Sun; Haiyang Sun; Kang Chen; Zhuofan Wen; Hao Gu; Bin Liu; Jianhua Tao

GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition

Zheng Lian, Licai Sun, Haiyang Sun, Kang Chen, Zhuofan Wen, Hao Gu, Bin Liu, Jianhua Tao

TL;DR

GPT-4V is evaluated on Generalized Emotion Recognition (GER) tasks across six task types and 21 datasets, establishing a zero-shot benchmark for multimodal emotion understanding. The study finds GPT-4V exhibits strong visual comprehension and temporal fusion capabilities but struggles with domain-specific micro-expressions and audio modalities. A GER-oriented calling strategy combining batch-wise, repeated, and recursive querying is proposed to respect API limits and minimize security-check failures. Overall, GPT-4V outperforms random baselines yet generally lags supervised systems, offering a practical zero-shot benchmark and analysis framework to guide future multimodal emotion research.

Abstract

Recently, GPT-4 with Vision (GPT-4V) has demonstrated remarkable visual capabilities across various tasks, but its performance in emotion recognition has not been fully evaluated. To bridge this gap, we present the quantitative evaluation results of GPT-4V on 21 benchmark datasets covering 6 tasks: visual sentiment analysis, tweet sentiment analysis, micro-expression recognition, facial emotion recognition, dynamic facial emotion recognition, and multimodal emotion recognition. This paper collectively refers to these tasks as ``Generalized Emotion Recognition (GER)''. Through experimental analysis, we observe that GPT-4V exhibits strong visual understanding capabilities in GER tasks. Meanwhile, GPT-4V shows the ability to integrate multimodal clues and exploit temporal information, which is also critical for emotion recognition. However, it's worth noting that GPT-4V is primarily designed for general domains and cannot recognize micro-expressions that require specialized knowledge. To the best of our knowledge, this paper provides the first quantitative assessment of GPT-4V for GER tasks. We have open-sourced the code and encourage subsequent researchers to broaden the evaluation scope by including more tasks and datasets. Our code and evaluation results are available at: https://github.com/zeroQiaoba/gpt4v-emotion.

GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition

TL;DR

Abstract

Paper Structure (21 sections, 13 figures, 10 tables, 1 algorithm)

This paper contains 21 sections, 13 figures, 10 tables, 1 algorithm.

Introduction
Related Works
Generalized Emotion Recognition
Unimodal Emotion Recognition
Multimodal Emotion Recognition
Multimodal Large Language Model
Evaluation on GPT-4V
Task Description
GPT-4V Calling Strategy
Results and Discussion
Main Results
Temporal Modeling Ability
Multimodal Fusion Ability
System Stability
Class-wise Performance Analysis
...and 6 more sections

Figures (13)

Figure 1: Performance of different methods on GER tasks. Here, "random" is a heuristic baseline that randomly selects labels from candidate categories.
Figure 2: Examples of different datasets. We pixelate faces due to the sensitive nature of human identity.
Figure 3: System stability test. We run GPT-4V 10 times, calculate the frequency with identical predictions, and report the test accuracy for different runs.
Figure 4: Confusion matrices for all evaluation datasets. Different colors represent distinct tasks.
Figure 5: Prediction consistency on RGB and grayscale images.
...and 8 more figures

GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition

TL;DR

Abstract

GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition

Authors

TL;DR

Abstract

Table of Contents

Figures (13)