CLLGDec 8, 2025

Do Generalisation Results Generalise?

arXiv:2512.07832v1h-index: 4
Originality Synthesis-oriented
AI Analysis

This work addresses the problem of accurately evaluating LLM generalization for deployment, but it is incremental as it builds on prior single-dataset assessments.

The study investigated whether out-of-distribution (OOD) generalization results for large language models (LLMs) are consistent across different testsets, finding no overarching trend and that correlations depend strongly on the specific model analyzed.

A large language model's (LLM's) out-of-distribution (OOD) generalisation ability is crucial to its deployment. Previous work assessing LLMs' generalisation performance, however, typically focuses on a single out-of-distribution dataset. This approach may fail to precisely evaluate the capabilities of the model, as the data shifts encountered once a model is deployed are much more diverse. In this work, we investigate whether OOD generalisation results generalise. More specifically, we evaluate a model's performance across multiple OOD testsets throughout a finetuning run; we then evaluate the partial correlation of performances across these testsets, regressing out in-domain performance. This allows us to assess how correlated are generalisation performances once in-domain performance is controlled for. Analysing OLMo2 and OPT, we observe no overarching trend in generalisation results: the existence of a positive or negative correlation between any two OOD testsets depends strongly on the specific choice of model analysed.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes