CLAINov 5, 2025

PLLuM: A Family of Polish Large Language Models

arXiv:2511.03823v16 citationsh-index: 24Has Code
Originality Synthesis-oriented
AI Analysis

This addresses the need for high-quality, culturally relevant language models for Polish speakers, though it is incremental as it applies existing methods to a new language.

The authors tackled the limited support for non-English languages in large language models by developing PLLuM, a family of open-source Polish language models, achieving the largest such models for Polish with a 140-billion-token corpus and demonstrating utility in public administration tasks.

Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for other languages. We present PLLuM (Polish Large Language Model), the largest open-source family of foundation models tailored specifically for the Polish language. Developed by a consortium of major Polish research institutions, PLLuM addresses the need for high-quality, transparent, and culturally relevant language models beyond the English-centric commercial landscape. We describe the development process, including the construction of a new 140-billion-token Polish text corpus for pre-training, a 77k custom instructions dataset, and a 100k preference optimization dataset. A key component is a Responsible AI framework that incorporates strict data governance and a hybrid module for output correction and safety filtering. We detail the models' architecture, training procedures, and alignment techniques for both base and instruction-tuned variants, and demonstrate their utility in a downstream task within public administration. By releasing these models publicly, PLLuM aims to foster open research and strengthen sovereign AI technologies in Poland.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes