LGAICLJun 15

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

arXiv:2606.1649721.7
Predicted impact top 6% in LG · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the challenge of automated GPU kernel optimization for performance engineers, offering a significant improvement over existing RL-based methods on a standard benchmark.

daVinci-kernel introduces a reinforcement learning framework for GPU kernel optimization that co-evolves skill selection, summarization, and utilization. It achieves 37.2%, 70.6%, and 32.2% on KernelBench Level 1, 2, and 3 under the Fast$_1$ threshold, outperforming the prior best RL model Dr.Kernel-14B.

GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present daVinci-kernel, a reinforcement learning framework that couples skill discovery with skill exploitation through a dynamically evolving skill library. daVinci-kernel jointly trains three agents sharing one LLM backbone: a Skill Selection Agent that retrieves relevant techniques via BM25 and LLM reranking, a Policy Agent that generates multi-turn CUDA/Triton kernels conditioned on selected skills, and a Skill Summary Agent that distills successful rollouts into reusable skills. Candidate skills are added only after execution-based verification confirms reproducible speedups. All three agents share a single LLM backbone, are initialized via a structured SFT cold start on diversity-filtered data, and are then jointly optimized end-to-end with multi-turn REINFORCE and per-agent advantage estimation. On KernelBench, daVinci-kernel-14B achieves 37.2%, 70.6%, and 32.2% on Level 1, Level 2, and Level 3 under the Fast$_1$ threshold, outperforming the strongest prior RL-trained model, Dr.Kernel-14B.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes