Table of Contents
Fetching ...

A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection

Gaku Morio, Harri Rowlands, Dominik Stammbach, Christopher D. Manning, Peter Henderson

TL;DR

The paper tackles the problem of understanding how oil & gas advertising frames public narratives and potentially engages in greenwashing by leveraging a novel multimodal benchmark of Facebook and YouTube ads. It presents a cross-domain dataset with expert labels across 13 framing types, plus transcripts and broad entity/country coverage, designed to evaluate vision-language models on video framing detection. The authors benchmark multiple VLMs under zero-shot and 1-shot prompts, revealing that current models, while promising, struggle with certain labels and cross-cultural contexts, and that transcript information substantially aids inference. They also propose practical applications toward automated greenwashing detection and provide pilot analyses of temporal trends and company-level framing to illustrate the dataset’s utility for social science and policy research. The work enables reproducible, multimodal analysis of corporate messaging in the energy sector and highlights clear directions for improving robust, cross-domain framing detection in videos.

Abstract

Companies spend large amounts of money on public relations campaigns to project a positive brand image. However, sometimes there is a mismatch between what they say and what they do. Oil & gas companies, for example, are accused of "greenwashing" with imagery of climate-friendly initiatives. Understanding the framing, and changes in framing, at scale can help better understand the goals and nature of public relations campaigns. To address this, we introduce a benchmark dataset of expert-annotated video ads obtained from Facebook and YouTube. The dataset provides annotations for 13 framing types for more than 50 companies or advocacy groups across 20 countries. Our dataset is especially designed for the evaluation of vision-language models (VLMs), distinguishing it from past text-only framing datasets. Baseline experiments show some promising results, while leaving room for improvement for future work: GPT-4.1 can detect environmental messages with 79% F1 score, while our best model only achieves 46% F1 score on identifying framing around green innovation. We also identify challenges that VLMs must address, such as implicit framing, handling videos of various lengths, or implicit cultural backgrounds. Our dataset contributes to research in multimodal analysis of strategic communication in the energy sector.

A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection

TL;DR

The paper tackles the problem of understanding how oil & gas advertising frames public narratives and potentially engages in greenwashing by leveraging a novel multimodal benchmark of Facebook and YouTube ads. It presents a cross-domain dataset with expert labels across 13 framing types, plus transcripts and broad entity/country coverage, designed to evaluate vision-language models on video framing detection. The authors benchmark multiple VLMs under zero-shot and 1-shot prompts, revealing that current models, while promising, struggle with certain labels and cross-cultural contexts, and that transcript information substantially aids inference. They also propose practical applications toward automated greenwashing detection and provide pilot analyses of temporal trends and company-level framing to illustrate the dataset’s utility for social science and policy research. The work enables reproducible, multimodal analysis of corporate messaging in the energy sector and highlights clear directions for improving robust, cross-domain framing detection in videos.

Abstract

Companies spend large amounts of money on public relations campaigns to project a positive brand image. However, sometimes there is a mismatch between what they say and what they do. Oil & gas companies, for example, are accused of "greenwashing" with imagery of climate-friendly initiatives. Understanding the framing, and changes in framing, at scale can help better understand the goals and nature of public relations campaigns. To address this, we introduce a benchmark dataset of expert-annotated video ads obtained from Facebook and YouTube. The dataset provides annotations for 13 framing types for more than 50 companies or advocacy groups across 20 countries. Our dataset is especially designed for the evaluation of vision-language models (VLMs), distinguishing it from past text-only framing datasets. Baseline experiments show some promising results, while leaving room for improvement for future work: GPT-4.1 can detect environmental messages with 79% F1 score, while our best model only achieves 46% F1 score on identifying framing around green innovation. We also identify challenges that VLMs must address, such as implicit framing, handling videos of various lengths, or implicit cultural backgrounds. Our dataset contributes to research in multimodal analysis of strategic communication in the energy sector.
Paper Structure (87 sections, 13 figures, 7 tables)

This paper contains 87 sections, 13 figures, 7 tables.

Figures (13)

  • Figure 1: Overview of the task of our dataset.
  • Figure 2: The country distribution (based on headquarters location) of YouTube. Note that we primarily focus on English videos produced by multinational corporations.
  • Figure 3: The overview of the entity-aware 1-shot prompt construction.
  • Figure 4: The video length and F-score of GPT-4.1.
  • Figure 5: The region and F-score for YouTube. We exclude Middle East from Asia.
  • ...and 8 more figures