SEJul 1

Rise From The Ashes: LLM-based Static Analysis for Deep Learning Framework Bugs

arXiv:2607.0055512.6
Predicted impact top 30% in SE · last 90 daysOriginality Highly original
AI Analysis

For DL framework developers, Phoenix provides a practical static analysis complement to dynamic testing, addressing the high cost of fuzzing and the lack of effective static tools for multilingual tensor bugs.

Phoenix is the first LLM-based static analysis technique for deep learning frameworks, using a structured semantic bridge intermediate representation (SBIR) to model cross-language tensor flows. It found 31 real new bugs in PyTorch across heterogeneous hardware backends, with 20 patches merged.

Deep learning (DL) frameworks are critical AI infrastructures that often hide bugs with serious security implications. While dynamic approaches such as fuzzing are effective in uncovering these bugs, they require real test execution and incur high computational costs. Static analysis is a natural complement because it can detect bugs without runtime execution, offering fast and scalable testing. Unfortunately, there is still limited work targeting static analysis for DL frameworks due to their multilingual architectures and tensor-related program state. We present Phoenix, the first LLM-based static analysis technique for DL frameworks. Our key insight is that cross-language tensor flows in DL frameworks can be modeled, together with concrete code context, as a structured semantic bridge intermediate representation (SBIR) that LLMs can analyze for potential bugs in tensor semantic propagation. We implement this insight through a multi-agent workflow. A summarization agent first distills bug summaries from historical bug-fix patches and CWE rules. Guided by each summary, an extraction agent identifies bug-relevant repository symbols for code retrieval, and a generation agent synthesizes grounded SBIRs from the retrieved context. Finally, an analysis agent is leveraged to check SBIRs and report potential bugs. Our evaluation shows that Phoenix is a practical complement to dynamic DL framework testing for bug finding. To date, Phoenix has found 31 real new bugs in PyTorch for different heterogeneous hardware backends (Intel CPU, NVIDIA CUDA, and Apple MPS). Among them, 20 submitted bug-fixing patches have been merged into upstream.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes