Table of Contents
Fetching ...

From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene

Mojca Brglez, Špela Vintar

TL;DR

This work presents SloPragEval and SloPragMega, the first Slovene benchmarks for pragmatics understanding, addressing the gap in assessing context-driven meaning beyond syntax and semantics. By translating and adapting English pragmatics datasets with careful translation and localization procedures, the authors construct native Slovene data and validate it via a human baseline, then evaluate several open-source and proprietary LLMs using Slovene and English prompts. Results reveal substantial progress for top proprietary models in grasping nuanced pragmatic phenomena, yet persistent gaps for open-source models and notable difficulties with non-literal and culture-specific inferences, especially in Manner and Quantity tasks. The study highlights the importance of native-data benchmarks, human validation, and careful dataset design to avoid translation pitfalls and cultural misalignment, providing a resource for future benchmark development and cross-linguistic pragmatic evaluation.

Abstract

Large language models are demonstrating increasing capabilities, excelling at benchmarks once considered very difficult. As their capabilities grow, there is a need for more challenging evaluations that go beyond surface-level linguistic competence. Namely, language competence involves not only syntax and semantics but also pragmatics, i.e., understanding situational meaning as shaped by context as well as linguistic and cultural norms. To contribute to this line of research, we introduce SloPragEval and SloPragMega, the first pragmatics understanding benchmarks for Slovene that contain altogether 405 multiple-choice questions. We discuss the difficulties of translation, describe the campaign to establish a human baseline, and report pilot evaluations with LLMs. Our results indicate that current models have greatly improved in understanding nuanced language but may still fail to infer implied speaker meaning in non-literal utterances, especially those that are culture-specific. We also observe a significant gap between proprietary and open-source models. Finally, we argue that benchmarks targeting nuanced language understanding and knowledge of the target culture must be designed with care, preferably constructed from native data, and validated with human responses.

From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene

TL;DR

This work presents SloPragEval and SloPragMega, the first Slovene benchmarks for pragmatics understanding, addressing the gap in assessing context-driven meaning beyond syntax and semantics. By translating and adapting English pragmatics datasets with careful translation and localization procedures, the authors construct native Slovene data and validate it via a human baseline, then evaluate several open-source and proprietary LLMs using Slovene and English prompts. Results reveal substantial progress for top proprietary models in grasping nuanced pragmatic phenomena, yet persistent gaps for open-source models and notable difficulties with non-literal and culture-specific inferences, especially in Manner and Quantity tasks. The study highlights the importance of native-data benchmarks, human validation, and careful dataset design to avoid translation pitfalls and cultural misalignment, providing a resource for future benchmark development and cross-linguistic pragmatic evaluation.

Abstract

Large language models are demonstrating increasing capabilities, excelling at benchmarks once considered very difficult. As their capabilities grow, there is a need for more challenging evaluations that go beyond surface-level linguistic competence. Namely, language competence involves not only syntax and semantics but also pragmatics, i.e., understanding situational meaning as shaped by context as well as linguistic and cultural norms. To contribute to this line of research, we introduce SloPragEval and SloPragMega, the first pragmatics understanding benchmarks for Slovene that contain altogether 405 multiple-choice questions. We discuss the difficulties of translation, describe the campaign to establish a human baseline, and report pilot evaluations with LLMs. Our results indicate that current models have greatly improved in understanding nuanced language but may still fail to infer implied speaker meaning in non-literal utterances, especially those that are culture-specific. We also observe a significant gap between proprietary and open-source models. Finally, we argue that benchmarks targeting nuanced language understanding and knowledge of the target culture must be designed with care, preferably constructed from native data, and validated with human responses.
Paper Structure (26 sections, 3 figures, 5 tables)

This paper contains 26 sections, 3 figures, 5 tables.

Figures (3)

  • Figure 1: Example of a Quantity-flouting utterance from MultiPragEval (a) and SloPragEval (b). Culturally specific terms in yellow. Utterance in green, bolded.
  • Figure 2: Example from (Slo)PragMega: example from the Humor task. Original text on the left (a), Slovene example on the right (b). Problematic parts in red, adaptations in violet. Correct answer in bold.
  • Figure 3: Example from (Slo)PragMega: example from the Metaphor task. Original text on the left (a), Slovene example on the right (b). Problematic parts in red, adaptations in violet. Correct answer in bold.