AIJun 30

AxDafny: Agentic Verified Code Generation in Dafny

arXiv:2606.3200710.7
Predicted impact top 54% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For developers using Dafny for verified programming, AxDafny provides a practical agentic framework that significantly boosts automated proof generation.

AxDafny improves verification success for Dafny code generation by 6.5 percentage points over the previous best baseline, achieving 92.7% on DafnyBench, and introduces a new benchmark LCB-Pro-Dafny.

We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iteratively generates implementations, invariants, assertions, and termination arguments. We also introduce LiveCodeBench-Pro-Dafny (LCB-Pro-Dafny), a benchmark of 250 competition-style programming problems translated into Dafny with formal specifications and a verifier-based evaluation harness. On LCB-Pro-Dafny, AxDafny substantially improves verification success over baseline GPT-5.5 performance. On DafnyBench, AxDafny achieves 92.7\% verification success, outperforming the strongest previously reported proof-hint baseline by 6.5 percentage points. Lastly, we show that verification success and runtime test performance measure different aspects of generated code.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes