CLJun 9

Schützen: Evaluating LLM Safety in Bulgarian and German Contexts

arXiv:2606.11316v114.3h-index: 49Has Code
Predicted impact top 72% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the lack of safety evaluation datasets for non-English languages, particularly for German and Bulgarian, which is important for responsible LLM deployment in those regions.

The paper introduces Schützen, a German-Bulgarian safety dataset for evaluating LLM answerability under risk, revealing pronounced cross-language differences in safety behavior and highlighting the need for region-specific evaluation resources.

Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disrespectful content. Although substantial progress has been made in developing safety evaluation datasets, existing resources remain overwhelmingly English- and Chinese-centric. This limitation is particularly pronounced when evaluating languages that operate within shared sociocultural, legal, and ethical contexts. To address this gap, we introduce Schützen: a German--Bulgarian safety dataset designed to assess model answerability under risk, covering both a low-resource language (Bulgarian) and a high-resource language (German). Experiments with multilingual and language-specific LLMs reveal pronounced cross-language differences in safety behavior, highlighting the necessity of tailored, region-specific evaluation resources to support the responsible deployment of LLMs in Germany and Bulgaria. Datasets and code are available at https://github.com/xnlp-lab/Schutzen. Warning: this paper contains examples that may be offensive, harmful, or biased.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes