22.5CLMar 29
PRBench: End-to-end Paper Reproduction in Physics ResearchShi Qiu, Junyi Deng, Yiwei Deng et al.
This benchmark provides a rigorous, expert-validated test for evaluating AI agents' capability to autonomously reproduce scientific research, revealing systematic failures that must be addressed for progress in AI-driven science.
Fine-Tuning Small Reasoning Models for Quantum Field TheoryNathaniel S. Woodward, Zhiqi Gao, Yurii Kvasiuk et al.
This work provides a foundation for developing domain-specific reasoning capabilities in small LLMs for theoretical physics, addressing the scarcity of verifiable training data.
10.3LGMay 13
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis ReproductionDarius A. Faroughy, Sofia Palacios Schweitzer, Ian Pang et al.
This benchmark addresses the need for realistic, domain-specific evaluation of AI agents in scientific research, particularly for complex, long-horizon tasks.
11.8AIMay 7
When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic ReasoningVasilis Niarchos, Constantinos Papageorgakis, Alexander G. Stapleton et al.
This work provides a controlled testbed for evaluating interaction structures in AI-driven scientific discovery, offering practical insights for researchers using LLMs in theoretical physics.
Efficient AI-Inspired Reduction of Feynman Integrals via Tube SeedingJustin Berman, Francois Charton, Andres Luna et al.
This work addresses a key bottleneck in high-precision calculations for particle and gravitational-wave physics, offering a practical improvement for multi-loop integral reduction.
9.7HEP-PHApr 2
Generative models on phase spaceZachary Bogorad, Ibrahim Elsharkawy, Yonatan Kahn et al.
For high-energy physicists, this provides interpretable and reliable generative models that respect physical conservation laws exactly, addressing a key limitation of approximate methods.
20.8SEJul 7
Articulating Assumptions in AI-Generated Scientific Analyses through Task DecompositionAhmed Hammad, Mihoko Nojiri
For researchers using LLMs in scientific computing, the framework enhances reproducibility and understanding of generated analyses.
9.7HEP-PHMay 27
Neural Scaling Laws for Jet GenerationOz Amram, Darius A. Faroughy, Tjarko Gerdes et al.
For researchers training large generative models for collider physics, this work provides the first empirical evidence that scaling laws for jet generation differ from language models, highlighting fundamental limits in data and compute scaling.
18.5HEP-PHJul 9
Revisiting One-Zero and Two-Zero Neutrino Mass Textures in Light of Recent Oscillation and Cosmological DataHaruto Kitagawa, Coh Miyao, Satsuki Nishimura et al.
This work refines the classification of neutrino mass textures for particle physicists studying neutrino mass models and experimental searches.
9.6AIMay 25
Experiments in Agentic AI for ScienceJudy Fox, Geoffrey Fox
For scientists and researchers, this work provides practical agentic AI systems that automate data curation and report generation, though the approach is incremental.
Descending into the Modular BootstrapNathan Benjamin, A. Liam Fitzpatrick, Wei Li et al.
This work addresses the challenge of identifying unknown CFTs in theoretical physics, particularly in a parameter range lacking known examples, though it is incremental as it builds on existing modular bootstrap methods with technical improvements.
16.6HEP-LATJul 16
LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics ResearchHaofei Gao, Tingjia Miao, Wenkai Jin et al.
For lattice QCD researchers, this tool lowers the barrier to complex computing workflows, enabling rapid exploration and verification of non-standard scientific ideas.
16.4HEP-PHAug 14
Pairton: Iterative Reconstruction of Short-Lived ParticlesAndreas Hermansen, Chris Scheulen, Tobias Golling
For high-energy physics researchers, Pairton offers a new general paradigm for particle reconstruction, though its demonstrated gains are specific to one decay topology.
16.4HEP-PHAug 18
VERaiPHY -- Validation & Evaluation for Robust AI in PHYsicsGaia Grosso, Ramon Winterhalder, Lydia Brenner et al.
This initiative aims to improve the rigor and reliability of AI applications in physics by addressing the lack of systematic statistical validation, uncertainty quantification, and robustness assessment for researchers in fundamental physics.
Cross-Domain Transfer with Particle Physics Foundation Models: From Jets to Neutrino InteractionsGregor Krzmanc, Vinicius Mikuni, Benjamin Nachman et al.
This work demonstrates that particle physics foundation models can generalize across vastly different energy scales and detector technologies, enabling detector-agnostic inference for the particle physics community.
16.2HEP-THJun 8
Calling the Brane Next Door: The Kaluza-Klein Tower as a Gravitational Information ChannelKarim Benakli
For theoretical physicists exploring extra dimensions, this work provides a novel information-theoretic perspective on Kaluza-Klein towers as communication carriers, but remains speculative and incremental.
16.2HEP-PHAug 15
Uncovering Hidden Leptonic Correlations with Flow Matching and AutoencodersHaruto Kitagawa, Satsuki Nishimura, Hajime Otsuka
This work provides a new tool for exploring parameter spaces in particle physics, potentially aiding in understanding lepton flavor structure, though the immediate impact is domain-specific.
6.4HEP-PHMay 11
Dissecting Jet-Tagger Through Mechanistic InterpretabilitySaurabh Rai, Sanmay Ganguly
For jet physics practitioners, this work demonstrates that mechanistic interpretability methods from NLP can uncover physically meaningful circuits in jet taggers, providing a new tool for understanding and validating deep learning models in high-energy physics.
6.3HEP-PHMay 28
Generative Models and Statistical ValidationSascha Diefenbacher, Sofia Palacios Schweitzer, Gregor Kasieczka
This work addresses the problem of validating generative models for physicists using them as fast surrogates and density estimators.
15.4HEP-PHJun 12
Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training ObjectivesIbrahim Elsharkawy, Joschka Birk, Vinicius Mikuni et al.
For researchers building foundation models in simulation-based sciences, this study provides a systematic comparison of pre-training objectives, revealing task-specific optimal strategies and the need for multi-objective pre-training for transfer across classification and generation.