Grounded receipts (verbatim critique quotes) + benchmark-dominance edges (verbatim table cells): every quote is checked to appear in the citing paper's prose and every number to appear in its tables with the win direction recomputed — a grounding check, not a measured precision rate. Verify magnitudes against the cited paper. beaten_by shows ONE row per winning paper, chosen as its WIDEST relative margin, so each row is that paper's most favourable reported condition rather than a typical one — compare evidence_caveats.beaten_by_shown against beaten_by_pairs_total to see how much is off-screen, and expect unshown rows to be narrower. Each row carries margin_pct and effect: 'tie' means the reported win is under 1% relative, i.e. within table-rounding noise; 'unknown' means the margin was not computable (zero or non-numeric baseline) and is NOT evidence of a large win; metric_scope is a keyword guess at whether the column is a benchmark aggregate or a single subtask, not a verified classification. Receipts carry scope='class' when the paper's subject was a group of methods or an unresolved reference ('these methods', 'it') rather than this one by name; some unresolved subjects are still served as scope='method'. Coverage is arXiv benchmark tables only: 'superseded' means beaten in a published comparison, not that a method is dead or unusable, and production frameworks appear only as baselines. A literature signal, not a deployment recommendation.
← All of Speculative decoding