Antonio Valerio Miceli Barone, Poon Tsz Nok
This work addresses the problem of improving LLM reasoning for code equivalence, particularly for Haskell, with potential benefits for program verification tasks.
Language design, compilers, type systems
Antonio Valerio Miceli Barone, Poon Tsz Nok
This work addresses the problem of improving LLM reasoning for code equivalence, particularly for Haskell, with potential benefits for program verification tasks.
Shubham Agarwal, Alexander Krentsel, Shu Liu et al.
For developers of safety-critical distributed systems, IDS dramatically reduces the effort and cost of formal verification, which previously required months to years of expert work.
Simon Yu, Derek Chong, Ananjan Nandi et al.
Provides an efficient infrastructure for programming meta-agents, enabling runtime intervention, counterfactual optimization, and tree-RL training.
Marcus J. Min, Mike He, Zhaoyu Li et al.
For researchers in formal verification and AI, this position paper highlights a conceptual gap but offers no empirical evidence or new method.
Poorva Garg, Renato Lui Geh, Daniel Israel et al.
This work addresses the computational bottleneck of sampling many programs from LLMs for code generation and reasoning tasks, offering a more efficient test-time method.
Prashanth Vijayaraghavan, Apoorva Nitsure, Luyao Shi et al.
This work addresses the problem of improving RTL synthesis and summarization for hardware designers by incorporating symbolic planning, though it appears incremental as it builds on existing prompting and RAG methods.
Yaoxiang Wang, Qi Shi, ShangZhan Li et al.
This addresses the need for high-quality hardware design automation, though it is incremental by integrating existing tools into a multi-agent framework.
Jonas Bayer, Stefan Zetzsche, Olivier Bouissou et al.
For developers and researchers using LLMs for program verification, this work addresses the critical gap in violation detection, showing that training on formal verification traces can significantly improve performance.
Le Chen, Nuo Xu, Winson Chen et al.
For practitioners needing code translation in low-resource programming languages or frameworks, this work provides a method to generate high-quality training data, significantly boosting functional correctness.
Zhensu Sun, Zhihao Lin, Zhi Chen et al.
This addresses latency inefficiencies in coding agents for developers, though it is an incremental improvement over existing serial workflows.
Derek Egolf, Yuhao Zhou, Stavros Tripakis
This addresses the problem of evaluating LLMs' capabilities in program synthesis for AI and software engineering, showing they are currently incremental compared to specialized tools.
Pouya Pezeshkpour, Estevam Hruschka
For practitioners needing reliable, interpretable verification of LLM outputs across diverse tasks, this work provides an automated method to build compact executable verifiers that outperform both LLM-based and hand-crafted verifiers.
Timothy Zhou, Loris D'Antoni, Nadia Polikarpova
For developers of agentic applications, LBAC provides a principled way to enforce security policies uniformly across both agent-generated and developer-written code.
Jing Xiong, Qi Han, Chenchen Ding et al.
For hardware designers using LLM-based code generation, CktFormalizer provides a correctness firewall that prevents subtle defects from causing silent failures in synthesis and routing.
Size Zheng, Xuegui Zheng, Hanshi Sun et al.
For developers of large language models, DITRON provides a flexible, high-performance distributed programming solution that overcomes the rigidity of existing libraries and compilers.
Anh T. V. Dau, Shin Hwei Tan, Jinqiu Yang et al.
This addresses the challenge of maintaining legacy mainframe systems for enterprises and developers, though it is incremental as it adapts existing LLM methods to a specific domain.
Lize Shao, Michael Cardei, Zichen Xie et al.
For developers and security engineers, CDC provides a method to enforce program-level constraints during code generation without retraining, addressing a key limitation of autoregressive models.
Cormac Guerin, Frank Guerin
For developers of LLM-based autonomous agents, KAIJU offers a system-level abstraction that improves security and execution efficiency over the ReAct paradigm.
Yuming Feng, Frederick Pu, One An et al.
This benchmark addresses the need for scalable, coherent auto-formalization of interdependent mathematical theories, a critical bottleneck for formal verification.
Jacek Karwowski, Younesse Kaddar, Zihuiwen Ye et al.
This addresses a critical failure mode in automated Bayesian model discovery for researchers and practitioners using probabilistic programming, though it is incremental as it builds on existing language frameworks.