AIJul 20

PEARL: Auditable Repair for Scientific Reasoning Graph Extraction

arXiv:2607.1791717.1Has Code
Predicted impact top 23% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers needing inspectable reasoning traces in scientific AI workflows, PEARL provides a reliability layer by repairing LLM graph outputs, but the method is incremental as it builds on existing extraction and repair techniques.

PEARL is a training-free framework that repairs noisy LLM-generated scientific reasoning graphs into auditable, semantically valid structures, raising strict gate passes from 0/350 to 300/350 and average REA from 0.339 to 0.906 on the ARCHE benchmark.

Scientific Reasoning Graph Extraction (SRGE) aims to recover explicit links among observations, evidence, intermediate claims, and paper-level conclusions. LLMs can produce graph-like scientific explanations, but their outputs often mix malformed syntax, drifting edge labels, incorrectly oriented roots, and weak source anchors. We propose PEARL (Peircean Extraction via Abstraction and Repair Layer), a training-free framework that turns noisy LLM graph responses into auditable reasoning graphs and repairs them toward strict semantic validity. PEARL first materializes explicit graph content under a closed Peircean schema, then uses matched evidence-grounded judge feedback to repair rejected edge types, local inference steps, and terminal roots while preserving an audit trail. On five 70-paper model archives from ARCHE, a benchmark for latent reasoning-chain extraction, PEARL raises strict gate passes from 0/350 for the LLM baseline to 300/350, with average REA improving from 0.339 to 0.906. The graphs provide a reliability layer for research-agent and AI scientist workflows that need inspectable reasoning traces rather than unconstrained graph regeneration. Code and audit artifacts are available at https://github.com/BohanSu/auditable-repair-reasoning-graphs/tree/300-350_workshop .

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes