LGAIJun 12

CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation

arXiv:2606.14581v110.2
Predicted impact top 39% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For scientists using LLMs in costly experiments, CARE provides a safe and auditable way to leverage LLM creativity without unsafe exploration.

CARE introduces an auditable controller that allows LLMs to revise challenger ranking policies for high-throughput experimentation optimization while keeping a non-LLM incumbent as default, achieving final-best improvements from 80.0 to 88.5 on Minerva/Olympus and from 83.9 to 92.1 on ChemLex benchmarks.

Granting LLMs direct control over costly, irreversible scientific experiments leads to unsafe exploration and unstable performance, but discarding LLM creativity entirely sacrifices significant optimization potential. We introduce CARE (Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation), an auditable controller for high-throughput experimentation (HTE) optimization that keeps a non-LLM incumbent optimizer as the default action path while using LLMs to revise challenger ranking policies. Before each outcome is revealed, a public-evidence intervention gate compares the challenger with the incumbent. It authorizes the challenger's selection only when the evidence available before selection supports the change, with the decision recorded in the audit log. CARE outperforms all other evaluated methods on Minerva/Olympus and ChemLex benchmarks, with final-best improving from 80.0 to 88.5 on Minerva/Olympus and from 83.9 to 92.1 on ChemLex, relative to the public incumbent. Our experiments indicate that LLM self-evolution is more reliable when it expands the proposal space under an auditable controller, rather than directly choosing experiments.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes