CLJun 14

FinBalance: A Multi-Document Accounting Reconciliation Benchmark

arXiv:2606.1594919.2
Predicted impact top 44% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For NLP and finance researchers, this benchmark exposes that current LLMs fail at multi-document accounting reconciliation, a critical real-world task, despite plausible numerical outputs.

FinBalance is a multi-document accounting reconciliation benchmark that tests LLMs on reconciling source documents into journal entries and balance sheets, finding that even the best model achieves only 46% exact final-balance-sheet accuracy, with a 26-41 percentage point gap between reported and ledger-replayed balance sheets.

Existing financial-NLP benchmarks mostly evaluate prepared artifacts such as filings, tables, or extracted values. Real accounting begins earlier: source documents must be reconciled into cited journal entries, aggregated into a balance sheet, and checked for contradictions. We introduce FinBalance, a multi-document accounting reconciliation benchmark built from source-document bundles across eight industries, three period types, and five difficulty levels. Human-authored business scenarios, accounting policies, tax/FX treatments, document schemas, distractors, and inconsistency templates are composed by a deterministic generator whose ledger produces journal entries,balance sheets, and 23 inconsistency-code labels. On a 710-record evaluation split, six contemporary LLMs reach at most 46% exact final-balance-sheet accuracy. Four models show a 26-41 pp gap between BS_exact, the model's reported balance sheet, and BS_recon, the balance sheet obtained by replaying its entries through our ledger. Models often recover numerically plausible entries but fail to bind them to supporting documents and aggregate them consistently. Citation-pressure prompting barely changes document-linking errors, while ledger-feedback ablations substantially improve reported balance sheets and expose inconsistency-detection trade-offs. Expert finance reviewers validate the benchmark design and labels.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes