From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
For AI4Math researchers, this paper provides a critical analysis and strategic vision for advancing AI from solving predefined problems to engaging in open-ended mathematical research.
This position paper argues that current LLM-driven theorem provers are limited to well-defined problems and cannot tackle frontier research mathematics, which is open-ended and abstract. The authors review the field and identify core limitations, proposing a roadmap for developing AI research agents capable of formal mathematical reasoning at the research frontier.
Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures, which are often open-ended, under-specified, and involve multiple layers of abstraction. We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges with rigorous formal mathematical reasoning. In this position paper, we provide a systematic review of the field, covering datasets, auto-formalization, and proof synthesis. More importantly, we identify core limitations of existing systems in serving as mathematical research agents, examining issues across datasets, relational structure, mathematical exploration, tool ecosystem, and human-AI collaboration, outlining a strategic road-map for the future of AI4Math.