Stress-testing large language model agents in a robotic chemistry laboratory
For researchers deploying AI in physical labs, this work provides a measurable benchmark revealing that current LLM agents are far from ready for autonomous scientific experimentation.
The study tested LLM agents in a robotic chemistry lab, finding that only 3.3% of 4,608 trials produced executable workflows, with the best system achieving 28.1%. Agents struggled with long-horizon planning and failed to replan workflows based on experimental feedback.
AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.