AIJun 11

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

arXiv:2606.13602v18.2
Predicted impact top 77% in AI · last 90 daysOriginality Incremental advance
AI Analysis

This benchmark provides a rigorous, domain-specific evaluation for AI agents in epigenomics, revealing that current systems lack the necessary scientific reasoning for complex analysis tasks.

EpiBench is a verifiable benchmark for short-horizon epigenomics analysis that evaluates AI agents on realistic workflow decisions. Across 5,088 trajectories from 16 model-harness pairs, the best system (GPT-5.5 / Pi) passed only 45.0% of attempts, with most failures due to lack of deep scientific judgment.

We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic workflow states and return deterministically gradable answers. The benchmark includes 106 evaluations across CUT\&Tag/CUT\&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows. Across 5,088 valid trajectories from 16 model-harness pairs, no system passed a majority of attempts: GPT-5.5 / Pi led at 45.0\% (143/318 attempts; 95\% confidence interval (CI), 36.3--53.7), followed by GPT-5.5 / OpenAI Codex at 39.9\% (127/318 attempts; 95\% CI, 31.6--48.3). Claude Opus 4.8 Max / Pi and GPT-5.4 / Pi each passed 39.0\% (124/318 attempts; 95\% CI, 30.2--47.8 and 31.0--47.0, respectively). Performance varies across assay types, and many failed runs still contain parts of the correct answer. Agents often found the right files and computed useful intermediate results, but failed when the task required deeper, assay-specific scientific judgment.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes