AISEJul 28

TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

arXiv:2607.263077.5h-index: 3
Predicted impact top 75% in AI · last 90 daysOriginality Synthesis-oriented
AI Analysis

For developers and auditors needing transparency in LLM-based code generation, TraceCoder provides explainability and auditability, though it is incremental in combining existing ideas.

TraceCoder introduces a code generation system with snippet versioning and provenance tracking, enabling auditable and explainable code repair. On 30 algorithmic tasks, it achieves 30% mean Chg% and 30% of snippets have traceable repair events, compared to 21% with Gemini 2.0 Flash.

Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible. We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that records, per repair event, the benchmark reference, round number, failure text, and LLM explanation, enabling full provenance queries; (ii) a browser-based visualisation tool that renders this history as heat-mapped, hover-annotated source code; and (iii) a competitive fractional position-key indexing scheme with tree-node delimiters that assigns stable, lexicographically-ordered identifiers to each code snippet, enabling fine-grained tracking without disrupting surrounding lines. We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two provider configurations. Of these, 10 exhaust the 6-iteration budget on tasks with subtle edge-case behaviour. Mean Chg% reaches 30%, three in ten code snippets carry a traceable repair-event row, compared to 21% when using Gemini 2.0 Flash as sole provider on a 20-task subset. Three detailed case studies demonstrate how the system explains which specific benchmark failures shaped each line of the final program. The proposed mechanism makes the internal "narrative" of automated code generation auditable and replayable, a property essential for trust and accountability in production deployments.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes