Table of Contents
Fetching ...

Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations

Pingjun Hong, Beiduo Chen, Siyao Peng, Marie-Catherine de Marneffe, Benjamin Roth, Barbara Plank

TL;DR

This work reframes NLI annotation as a two-step process: reasoning (Step 1) and labeling (Step 2), analyzed through the LiTEx explanation taxonomy. By applying LiTEx to LiveNLI and VariErr, the authors jointly examine NLI labels, explanation categories, and explanation-text similarity to uncover how annotators diverge in both reasoning and label choice. They demonstrate that alignment in reasoning types better tracks explanation similarity than surface label agreement, and reveal robust annotator-level preferences and cross-dataset patterns. The findings underscore that explanations provide a richer view of human interpretive variation and caution against treating labels as ground truth without considering underlying reasoning paths.

Abstract

Natural Language Inference datasets often exhibit human label variation. To better understand these variations, explanation-based approaches analyze the underlying reasoning behind annotators' decisions. One such approach is the LiTEx taxonomy, which categorizes free-text explanations in English into reasoning types. However, previous work applying such taxonomies has focused on within-label variation: cases where annotators agree on the final NLI label but provide different explanations. In contrast, this paper broadens the scope by examining how annotators may diverge not only in the reasoning type but also in the labeling step. We use explanations as a lens to decompose the reasoning process underlying NLI annotation and to analyze individual differences. We apply LiTEx to two NLI English datasets and align annotation variation from multiple aspects: NLI label agreement, explanation similarity, and taxonomy agreement, with an additional compounding factor of annotators' selection bias. We observe instances where annotators disagree on the label but provide highly similar explanations, suggesting that surface-level disagreement may mask underlying agreement in interpretation. Moreover, our analysis reveals individual preferences in explanation strategies and label choices. These findings highlight that agreement in reasoning types better reflects the semantic similarity of free-text explanations than label agreement alone. Our findings underscore the richness of reasoning-based explanations and the need for caution in treating labels as ground truth.

Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations

TL;DR

This work reframes NLI annotation as a two-step process: reasoning (Step 1) and labeling (Step 2), analyzed through the LiTEx explanation taxonomy. By applying LiTEx to LiveNLI and VariErr, the authors jointly examine NLI labels, explanation categories, and explanation-text similarity to uncover how annotators diverge in both reasoning and label choice. They demonstrate that alignment in reasoning types better tracks explanation similarity than surface label agreement, and reveal robust annotator-level preferences and cross-dataset patterns. The findings underscore that explanations provide a richer view of human interpretive variation and caution against treating labels as ground truth without considering underlying reasoning paths.

Abstract

Natural Language Inference datasets often exhibit human label variation. To better understand these variations, explanation-based approaches analyze the underlying reasoning behind annotators' decisions. One such approach is the LiTEx taxonomy, which categorizes free-text explanations in English into reasoning types. However, previous work applying such taxonomies has focused on within-label variation: cases where annotators agree on the final NLI label but provide different explanations. In contrast, this paper broadens the scope by examining how annotators may diverge not only in the reasoning type but also in the labeling step. We use explanations as a lens to decompose the reasoning process underlying NLI annotation and to analyze individual differences. We apply LiTEx to two NLI English datasets and align annotation variation from multiple aspects: NLI label agreement, explanation similarity, and taxonomy agreement, with an additional compounding factor of annotators' selection bias. We observe instances where annotators disagree on the label but provide highly similar explanations, suggesting that surface-level disagreement may mask underlying agreement in interpretation. Moreover, our analysis reveals individual preferences in explanation strategies and label choices. These findings highlight that agreement in reasoning types better reflects the semantic similarity of free-text explanations than label agreement alone. Our findings underscore the richness of reasoning-based explanations and the need for caution in treating labels as ground truth.
Paper Structure (19 sections, 3 equations, 6 figures, 3 tables)

This paper contains 19 sections, 3 equations, 6 figures, 3 tables.

Figures (6)

  • Figure 1: Decomposing NLI annotations into a two-step decision-making framework: Step 1 (Reasoning) and Step 2 (Labeling). The two examples illustrate how annotators may diverge at one of the two steps, featuring the LiTEx categories Logical Conflict and Absence of Mention.
  • Figure 2: Co-occurrence of LiTEx explanation categories and NLI labels across three datasets (e-SNLI, LiveNLI, and VariErr).
  • Figure 3: Distribution of NLI labels (entailment, neutral, contradiction) across LiveNLI and VariErr annotators. The legend at the bottom specifies the color–label correspondence, while the area of each color segment represents the number of instances assigned to that label.
  • Figure 4: Distribution of explanation category per annotators in LiveNLI and VariErr. Colors correspond to different explanation categories.
  • Figure 5: Pairwise annotator agreement (conditional Cohen's $\kappa$) between taxonomy matches (T) and label matches (L).
  • ...and 1 more figures