Table of Contents
Fetching ...

Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs

Kyubyung Chae, Gihoon Kim, Gyuseong Lee, Taesup Kim, Jaejin Lee, Heejin Kim

TL;DR

The paper tackles whether sovereign LLMs truly align with local socio-cultural contexts and maintain safety, proposing a cross-national, multilingual evaluation framework and a new dataset. It combines quantitative accuracy-based testing with human QA and story-generation assessments, plus jailbreaking trials to probe safety. Across six languages and ten models, results show that domestic sovereignty does not guarantee superior socio-cultural understanding and that safety vulnerabilities persist, especially in smaller, locally tuned systems. The work highlights the need for broader, evidence-based evaluation that jointly considers linguistic coverage, cultural nuance, and robust safety to responsibly advance sovereign LLMs.

Abstract

Recent trends in LLMs development clearly show growing interest in the use and application of sovereign LLMs. The global debate over sovereign LLMs highlights the need for governments to develop their LLMs, tailored to their unique socio-cultural and historical contexts. However, there remains a shortage of frameworks and datasets to verify two critical questions: (1) how well these models align with users' socio-cultural backgrounds, and (2) whether they maintain safety and technical robustness without exposing users to potential harms and risks. To address this gap, we construct a new dataset and introduce an analytic framework for extracting and evaluating the socio-cultural elements of sovereign LLMs, alongside assessments of their technical robustness. Our experimental results demonstrate that while sovereign LLMs play a meaningful role in supporting low-resource languages, they do not always meet the popular claim that these models serve their target users well. We also show that pursuing this untested claim may lead to underestimating critical quality attributes such as safety. Our study suggests that advancing sovereign LLMs requires a more extensive evaluation that incorporates a broader range of well-grounded and practical criteria.

Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs

TL;DR

The paper tackles whether sovereign LLMs truly align with local socio-cultural contexts and maintain safety, proposing a cross-national, multilingual evaluation framework and a new dataset. It combines quantitative accuracy-based testing with human QA and story-generation assessments, plus jailbreaking trials to probe safety. Across six languages and ten models, results show that domestic sovereignty does not guarantee superior socio-cultural understanding and that safety vulnerabilities persist, especially in smaller, locally tuned systems. The work highlights the need for broader, evidence-based evaluation that jointly considers linguistic coverage, cultural nuance, and robust safety to responsibly advance sovereign LLMs.

Abstract

Recent trends in LLMs development clearly show growing interest in the use and application of sovereign LLMs. The global debate over sovereign LLMs highlights the need for governments to develop their LLMs, tailored to their unique socio-cultural and historical contexts. However, there remains a shortage of frameworks and datasets to verify two critical questions: (1) how well these models align with users' socio-cultural backgrounds, and (2) whether they maintain safety and technical robustness without exposing users to potential harms and risks. To address this gap, we construct a new dataset and introduce an analytic framework for extracting and evaluating the socio-cultural elements of sovereign LLMs, alongside assessments of their technical robustness. Our experimental results demonstrate that while sovereign LLMs play a meaningful role in supporting low-resource languages, they do not always meet the popular claim that these models serve their target users well. We also show that pursuing this untested claim may lead to underestimating critical quality attributes such as safety. Our study suggests that advancing sovereign LLMs requires a more extensive evaluation that incorporates a broader range of well-grounded and practical criteria.
Paper Structure (33 sections, 7 figures, 10 tables)

This paper contains 33 sections, 7 figures, 10 tables.

Figures (7)

  • Figure 1: Overview of our evaluation framework for socio-cultural understanding and safety in sovereign LLMs.
  • Figure 2: Examples of prompts used in the main experiments: The multiple-choice prompt (left) consists of a question, options, and the answer. The QA prompt (center) provides a scenario concerning specific socio-cultural aspects of a country. The story generation prompt (right) contains an overview of a novel, an original excerpt, and a request for story writing.
  • Figure 3: Attack Success Rate in Jailbreak Attempts. Despite undergoing continued pretraining from Llama3-8B, both Typhoon-Llama3-8B and Nordic-Llama3-8B showed attack success rates that were approximately two to three times higher than those of the original Llama3-8B.
  • Figure E.1: Example of a Survey for story generation tasks: Evaluators can view the given prompt for each language model in the the top-left. In the bottom-left, they can see anonymized models along with their responses. Then, as shown on the right, they can assign scores for each criterion and provide comments.
  • Figure E.2: (Please zoom in for a better view) This figure illustrates the input prompts used for QA assessments. These prompts serve as inputs to the target language models, and the evaluation is conducted based on their outputs. The assessment consists of a total of 30 questions, structured across five categories and six languages.
  • ...and 2 more figures