Table of Contents
Fetching ...

On the Interplay between Human Label Variation and Model Fairness

Kemal Kurniawan, Meladel Mistica, Timothy Baldwin, Jey Han Lau

TL;DR

This work investigates how human label variation (HLV) interacts with model fairness. It systematically compares training on majority-vote labels with four HLV methods across SBIC and TAG, using soft F1-based metrics and a group/class-aware fairness score $s_{kg}$. The study finds that HLV generally boosts performance and often preserves or improves fairness, with minority annotations driving fairness gains as shown by a temperature-scaling analysis. It highlights the importance of selecting fairness definitions and configurations tailored to the application and acknowledges dataset limitations, especially for the confidential TAG data.

Abstract

The impact of human label variation (HLV) on model fairness is an unexplored topic. This paper examines the interplay by comparing training on majority-vote labels with a range of HLV methods. Our experiments show that without explicit debiasing, HLV training methods have a positive impact on fairness.

On the Interplay between Human Label Variation and Model Fairness

TL;DR

This work investigates how human label variation (HLV) interacts with model fairness. It systematically compares training on majority-vote labels with four HLV methods across SBIC and TAG, using soft F1-based metrics and a group/class-aware fairness score . The study finds that HLV generally boosts performance and often preserves or improves fairness, with minority annotations driving fairness gains as shown by a temperature-scaling analysis. It highlights the importance of selecting fairness definitions and configurations tailored to the application and acknowledges dataset limitations, especially for the confidential TAG data.

Abstract

The impact of human label variation (HLV) on model fairness is an unexplored topic. This paper examines the interplay by comparing training on majority-vote labels with a range of HLV methods. Our experiments show that without explicit debiasing, HLV training methods have a positive impact on fairness.
Paper Structure (29 sections, 4 equations, 7 figures, 3 tables)

This paper contains 29 sections, 4 equations, 7 figures, 3 tables.

Figures (7)

  • Figure 1: Fraction of randomly sampled configurations where MV is not significantly fairer than each HLV training method on SBIC.
  • Figure 2: Fraction of randomly sampled configurations where MV is not significantly fairer than each HLV training method on SBIC for various levels of exponent $p$ in the group-wise aggregation.
  • Figure 3: Impact of minority annotations ($\tau$) to overall performance (left) and fairness (right) on TAG. Larger $\tau$ means minority annotations are weighted more.
  • Figure 4: Fraction of randomly sampled configurations where MV is not significantly fairer than each HLV training method on TAG. Whiskers indicate 95 bootstrap confidence intervals.
  • Figure 5: Fraction of randomly sampled configurations where each HLV training method is significantly fairer than MV. Whiskers indicate 95 bootstrap confidence intervals.
  • ...and 2 more figures