22.5CLMar 29
PRBench: End-to-end Paper Reproduction in Physics ResearchShi Qiu, Junyi Deng, Yiwei Deng et al.
This benchmark provides a rigorous, expert-validated test for evaluating AI agents' capability to autonomously reproduce scientific research, revealing systematic failures that must be addressed for progress in AI-driven science.
24.3CHEM-PHMay 18Code
Harnessing AtomisticSkills for Agentic Atomistic ResearchBowen Deng, Bohan Li, Matthew Cox et al.
For computational materials science and chemistry researchers, AtomisticSkills provides a modular infrastructure to automate complex atomistic workflows, addressing the challenge of scaling monolithic agents in fragmented software ecosystems.
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMsYadi Cao, Sicheng Lai, Jiahe Huang et al.
This addresses the gap in cost-aware evaluation for physics simulations, providing a practical benchmark for researchers and practitioners, though it is incremental in extending existing benchmarking approaches.
El Agente Forjador: Task-Driven Agent Generation for Quantum SimulationZijian Zhang, Aiwei Yin, Amaan Baweja et al.
This work addresses the bottleneck of static, hand-curated toolsets in scientific agentic systems, enabling autonomous tool creation and reuse for quantum chemistry and dynamics tasks.
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational ScienceNithin Somasekharan, Youssef Hassan, Shiyao Lin et al.
For researchers building scientific AI assistants, this benchmark addresses the overlooked problem of refining ill-posed user requests through dialogue, but the results show current models are far from reliable.
11.0LGMar 29
Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learningYuanzhao Zhang, William Gilpin
For researchers developing time-series foundation models, this work reveals a critical performance gap and failure mode, providing a simple baseline that should be used to validate whether models are truly learning beyond trivial copying.
11.9COMP-PHApr 1
Grading the Unspoken: Evaluating Tacit Reasoning in Quantum Field Theory and String Theory with LLMsXingyang Yu, Yinghuan Zhang, Yufei Zhang et al.
It provides a sensitive evaluation lens for epistemic limits of LLMs in highly abstract theoretical physics, but the dataset is small and the findings are incremental.
11.2CEMar 28
Sparse Autoencoders as a Steering Basis for Phase Synchronization in Graph-Based CFD SurrogatesYeping Hu, Ruben Glatt, Shusen Liu
For users of graph-based CFD surrogates in digital twins and closed-loop control, this work provides a method to correct phase drift post hoc, extending latent-space steering to time-dependent physical systems.
11.6LGMay 25
Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE SolversShyam Sankaran, Hanwen Wang, Paris Perdikaris
For researchers building neural PDE solvers, this shows that small models with strong priors can rival huge foundation models, challenging the scaling trend in this domain.
11.3DIS-NNMay 12
The critical slowing down in diffusion modelsLuca Maria Del Bono, Giulio Biroli, Patrick Charbonneau et al.
Provides theoretical insight into the limitations of diffusion models near criticality and demonstrates how architectural design can overcome these bottlenecks, relevant for statistical physics and generative modeling.
9.4LGApr 10
EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention TransformersYi-Lun Liao, Alexander J. Hoffman, Sabrina C. Shen et al.
This work addresses the need for scalable and physically consistent models in computational chemistry and materials science, representing an incremental advancement over previous versions.
8.2LGJun 1
Speculative Sampling For Faster Molecular DynamicsArthur Kosmala, Stephan Günnemann, Meng Gao et al.
For computational chemists and physicists, LSD addresses the serial bottleneck in molecular dynamics to increase single-system throughput without introducing error.
9.3LGMar 20
Neural Uncertainty Principle: A Unified View of Adversarial Fragility and LLM HallucinationDong-Xiao Zhang, Hu Lou, Jun-Jie Zhang et al.
It addresses reliability issues in AI systems by providing a unified framework for diagnosing and mitigating anomalies across perception and generation tasks, though it is incremental in applying existing uncertainty concepts to new contexts.
8.8LGMar 11
On the Value of Tokeniser Pretraining in Physics Foundation ModelsHadi Sotoudeh, Payel Mukhopadhyay, Ruben Ohana et al.
This provides practical guidance for training efficient physics emulators in data-limited settings, though it is incremental as it focuses on optimizing an existing approach.
6.8COMP-PHJun 2
An efficient and energy stable framework for phase field simulations of grain growth in additive manufacturingChaoqian Yuan, Chinnapat Panwisawas, Ye Lu
For researchers simulating microstructure evolution in additive manufacturing, this framework offers a way to drastically reduce computational cost without sacrificing stability.
8.5LGMar 15
Excited Pfaffians: Generalized Neural Wave Functions Across Structure and StateNicholas Gao, Till Grutschus, Frank Noé et al.
This work addresses the scalability challenge for quantum chemistry simulations, offering a more efficient method for modeling excited states across molecules, though it is incremental in improving existing neural-network wave function approaches.
Mesh Based Simulations with Spatial and Temporal awarenessPaul Garnier, Vincent Lannelongue, Elie Hachem
For researchers in geometric deep learning and computational physics, this work addresses a critical bottleneck in training physics surrogates by incorporating numerical analysis principles, leading to more accurate and stable simulations.
7.5LGApr 8
Bayesian Optimization for Mixed-Variable Problems in the Natural SciencesYuhao Zhang, Ti John, Matthias Stosiek et al.
This work provides a practical BO framework for mixed-variable optimization problems in the natural sciences, particularly useful in autonomous laboratory settings with noise and limited data, but it is incremental as it builds on existing PR methods.
Building Trust in PINNs: Error Estimation through Finite Difference MethodsAleksander Krasowski, René P. Klausen, Aycan Celik et al.
This addresses the trust issue for researchers and practitioners using PINNs in scientific computing by providing interpretable error explanations, though it is incremental as it builds on existing PINN frameworks.
13.6MLMay 29
Free energy Estimation on Any State SpaceJiajun He, Zijing Ou, Francisco Vargas et al.
This work provides a more general and efficient method for free energy estimation, which is a fundamental problem across physics and statistics.