An Luo, Jin Du, Xun Xian et al.
This addresses the problem of assessing AI capabilities versus human expertise in domain-specific data science for researchers and practitioners, showing incremental insights by benchmarking current limitations.
Statistical methodology, experimental design
An Luo, Jin Du, Xun Xian et al.
This addresses the problem of assessing AI capabilities versus human expertise in domain-specific data science for researchers and practitioners, showing incremental insights by benchmarking current limitations.
Ander Artola Velasco, Stratis Tsirtsis, Manuel Gomez-Rodriguez
For regulators and product developers in high-stakes domains like drug development, this work addresses the inefficiency of costly trials that may deter socially valuable 'moonshot' products.
Jonas Arruda, Niels Bracher, Ullrich Köthe et al.
It provides a comprehensive guide for researchers and practitioners applying diffusion models to simulation-based inference, but is a review rather than novel research.
Tobias Holtdirk, Georg Ahnert, Joseph W Sakshaug et al.
For survey researchers and data analysts, this provides a more accurate imputation method for public opinion data, especially under non-random missingness.
Takashi Ishida, Thanawat Lodkaew, Ikko Yamane
For LLM benchmark creators and evaluators, CapBencher provides a built-in alarm to detect test-set overfitting and evaluation gaming, addressing a critical issue in reliable model assessment.
Johnny Tian-Zheng Wei, Jerry Li, Ameya Godbole et al.
This work addresses the underexplored problem of correcting test set contamination for machine learning practitioners, offering a practical method to obtain more reliable evaluation scores.
Victoria Lin, Taedong Yun, Maja Matarić et al.
For researchers using LLMs as human simulators in causal inference, this paper identifies and provides diagnostics for a critical confounding bias, though the mitigation approach is incremental.
Bin Zhu, Yanghui Rao
For practitioners deploying LLM judge panels, the paper provides a practical regime map to decide calibration strategy under limited human labels, showing that the key question is whether the next judge's information is estimable.
Xun Huan, Jayanth Jagalur, Youssef Marzouk
For researchers and practitioners in modeling and prediction across sciences and engineering, this survey provides a comprehensive overview of OED methods and identifies key open problems.
Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori et al.
For researchers evaluating coding agents, this provides a principled method to detect and prevent deceptive performance, addressing a growing failure mode in agent evaluation.
Nicholas Brawand, Nima Leclerc, Anhthy Ngo et al.
For healthcare AI monitoring, AM-PPI provides a statistically valid, label-efficient method that leverages multiple predictors of varying cost and accuracy, outperforming existing single-predictor approaches.
Elynn Chen, Xi Chen, Yi Zhang
This work addresses a domain-specific problem in e-commerce and retail for sellers needing efficient pricing strategies across diverse markets, representing an incremental advance by adapting existing transfer learning and bandit methods to a structured utility model.
Bohan Wu, Julius von Kügelgen, David M. Blei
This work addresses the estimation challenge in causal representation learning for researchers in machine learning, though it is incremental as it builds on known identifiability results with a novel method.
Geping Chen, Chunlin Li, Tianzhong Yang et al.
For practitioners and researchers in causal inference, TabCF provides an easy-to-use, tuning-light method for distributional causal effect estimation, serving as a strong baseline.
Mengchu Li, Jin Zhu, Jinglai Li et al.
For researchers and practitioners needing to localize LLM-generated segments in mixed text, this provides a principled segmentation approach, though it is an incremental adaptation of existing methods.
Alexander G. Reisach, Antoine Chambaz, Gilles Blanchard et al.
For researchers evaluating causal discovery algorithms on synthetic data, this work highlights a structural artifact that can inflate performance, suggesting the need for more realistic benchmarks.
Qi Qin, Jiajie Zhu, Dali Chen et al.
This work provides a method to significantly reduce the computational overhead of Tabular Foundation Models, making them more practical for large-scale deployment for practitioners in machine learning.
Peiman Mohseni, Nick Duffield, Raymond K. W. Wong
This work provides a more interpretable and efficient framework for translation-equivariant neural processes, benefiting applications in scientific and engineering domains requiring modeling of irregularly sampled functions.
Grégoire Martinon, Ibrahim Merad, Mohammed Raki
This work addresses the problem of costly and biased evaluation of agentic systems for researchers and practitioners by providing a unified, open-source library for Prediction-Powered Inference.
Zhanyu Wang, Arin Chang, Jordan Awan
This addresses the challenge of accurate statistical inference for researchers and practitioners using privacy-preserving data, offering a solution to biases that affect confidence intervals and hypothesis tests, though it is incremental by building on existing parametric bootstrap methods.