CLJul 30

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

arXiv:2607.2856834.23 citationsHas Code
Predicted impact top 1% in CL · last 90 daysOriginality Highly original
AI Analysis

This work addresses the challenge of building AI systems that can improve the process of building other AI systems, which is crucial for advancing recursive self-improvement research in machine learning engineering.

The paper introduces Frontis-MA1, a 35B parameter AI4AI model designed for recursive self-improvement in machine learning engineering. It improves the Medal Average from 39.39% to 60.61% over its base model on MLE-Bench Lite, and reaches 71.21% with an enhanced search strategy, outperforming GPT-5.5 + Codex and nearing GPT-5.6 Sol and Kimi K3. On NatureBench Lite, the model improves Match-SOTA from 50% to 70%, and the evolutionary framework improves it from 20% to 50%.

Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes