The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs
For practitioners and regulators relying on trust benchmarks for open-source LLMs, this work reveals that static scores are unreliable across model checkpoints, necessitating longitudinal reporting.
This paper audits four open-source LLM release lines across twelve checkpoints and finds that trustworthiness benchmark scores drift significantly between successive checkpoints, with mean absolute drift well above a no-drift null. The authors conclude that trust scores should be reported as checkpoint-bound, dated artefacts rather than carried forward across releases.
Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpoints of one release line as if the model behind it had not shifted. We test whether it has shifted by auditing four open-source release lines, Yi, Qwen, Mistral, and Gemma, at three successive generations each, on a fixed basket of trust benchmarks under multiple prompt templates. Mean absolute adjacent-generation drift lands well above an independence-based no-drift reference null, and the gap persists when we drop a benchmark, drop a release line, or switch to strict scoring. We therefore conclude that a trust score attached to a release line should not be carried forward to the next checkpoint without remeasurement; it should instead be reported as a checkpoint-bound, dated artefact, which we package as a longitudinal model card. Closed APIs, larger models, canonical benchmark protocols, and fixed month-cadence rules lie outside the audited scope and require their own evaluation.