CLAIJul 9

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

arXiv:2607.0804619.6h-index: 7
Predicted impact top 26% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners deploying LLM forecasters, this work provides a practical method to calibrate and audit model behavior, revealing that reasoning traces can be unreliable while internal representations offer a more faithful window.

The paper shows that probing internal representations of LLM forecasters yields better calibration than chain-of-thought reasoning, and that these probes can detect when reasoning is unfaithful, predicting direction of change in 84% of cases. Pre-reasoning representations already encode the final forecast, enabling token savings of 30-47% without accuracy loss.

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a result that also holds for GLM-4.7-Flash and GLM-4.5-Air. We then assess CoT faithfulness through evidence ablation and diversionary injection: removing an influential source in the prompt often changes the model's forecast while leaving the reasoning trace untouched. The same probes function as lie detectors: their activations track behavioral shifts far better than the reasoning trace does, and they also predict the direction of change in 84% of cases, including when the CoT conceals the perturbation's influence. Finally, forced answering reveals that forecasts are largely fixed before reasoning begins: a single pre-reasoning pass recovers the committed answer and confidence, and routing questions by the spread of this pre-set answer distribution saves 30-47% of generated tokens, with no loss of accuracy. Together, these results establish probing internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes