Semantic Properties of cosine based bias scores for word embeddings

Sarah Schröder; Alexander Schulz; Fabian Hinder; Barbara Hammer

Semantic Properties of cosine based bias scores for word embeddings

Sarah Schröder, Alexander Schulz, Fabian Hinder, Barbara Hammer

TL;DR

This work formalizes the semantic requirements for cosine-based bias scores and investigates two prominent measures, WEAT and Direct Bias, both theoretically and empirically. It shows that WEAT’s individual bias is not magnitude-comparable and its effect size can be unreliable for cross-model quantification, while Direct Bias is magnitude-comparable but not unbiased-trustworthy. Through experiments on multiple pretrained models, the paper demonstrates practical limitations in bias quantification and highlights the influence of attribute embeddings and biased directions on these scores. The results advocate for rigorous applicability checks, co-reporting of significance, and development of more robust bias quantification tools for fair evaluation in embedding spaces.

Abstract

Plenty of works have brought social biases in language models to attention and proposed methods to detect such biases. As a result, the literature contains a great deal of different bias tests and scores, each introduced with the premise to uncover yet more biases that other scores fail to detect. What severely lacks in the literature, however, are comparative studies that analyse such bias scores and help researchers to understand the benefits or limitations of the existing methods. In this work, we aim to close this gap for cosine based bias scores. By building on a geometric definition of bias, we propose requirements for bias scores to be considered meaningful for quantifying biases. Furthermore, we formally analyze cosine based scores from the literature with regard to these requirements. We underline these findings with experiments to show that the bias scores' limitations have an impact in the application case.

Semantic Properties of cosine based bias scores for word embeddings

TL;DR

Abstract

Paper Structure (18 sections, 7 theorems, 24 equations, 3 figures, 2 tables)

This paper contains 18 sections, 7 theorems, 24 equations, 3 figures, 2 tables.

Introduction
Related Work for Bias in World Embeddings
WEAT
Direct Bias
Terminology
Formal Requirements for Bias Scores
Formal Bias Definition and Notations
Requirements for Bias Metrics
Comparability
Trustworthiness
Analysis of Bias Scores
Analysis of WEAT
Analysis of the Direct Bias
Experiments
WEAT's effect size
...and 3 more sections

Key Result

Theorem 1

The bias score function $s(\mathbf{t},A,B)$ of WEAT is not magnitude-comparable.

Figures (3)

Figure 1: WEAT individual bias and effect size for distilBERT with different selections of target words. When selecting a smaller number of job titles (left), we observe that stereotypical male/female jobs are more distinct w.r.t. $s(\mathbf{t},A,B)$, while the effect size is lower.
Figure 2: WEAT's individual bias for job titles in GPT.
Figure 3: Correlation bias directions of individual word pairs (left: 0-24, right: 0-22) and the first principal component (last row) as selected for the Direct Bias. The lowest row in the heatmap shows the correlation of individual bias directions with the first principal component.

Theorems & Definitions (18)

Definition 3.1: individual bias
Definition 3.2: aggregated bias
Definition 3.3: magnitude-comparable
Definition 3.4: unbiased-trustworthy
Theorem 1
proof
Theorem 2
proof
Theorem 3
proof
...and 8 more

Semantic Properties of cosine based bias scores for word embeddings

TL;DR

Abstract

Semantic Properties of cosine based bias scores for word embeddings

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (3)

Theorems & Definitions (18)