LGAIJul 9

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

arXiv:2607.080547.4h-index: 5
Predicted impact top 44% in LG · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the blind spot of self-validation in LLM-assisted safety analysis tools, providing a method for deriving governance principles that can be enforced, though the approach is incremental and the results are preliminary.

The authors identify that LLM-assisted safety analysis tools are themselves safety-relevant systems that lack self-analysis, and propose Constitutional Meta-STPA, a closed-loop tool that applies STPA to itself to derive a governance constitution of 21 Tool Principles and 8 Meta-Safety Principles. They show that a frontier LLM ensemble recovers 18/21 canonical and 8/8 governance principles from the tool's own design, while a weaker pair recovers only 12/21 and 3/8, indicating the meta layer is model-limited.

Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, inside rigorous processes such as Systems-Theoretic Process Analysis (STPA). Yet a blind spot runs through this fast-growing literature: every system gets analysed except the LLM-assisted tool doing the analysing, which is itself a safety-relevant system that can hallucinate standards, emit unverifiable constraints, and leave no audit trail from prompt to artifact. We take seriously the question the field has skipped -- {who analyses the analyser?} and answer it by turning STPA on the tool itself. We present \{Constitutional Meta-STPA}, an LLM-assisted STPA tool built around a closed loop: the tool runs a {meta-STPA} of the class of AI-assisted safety tools and {derives} rather than asserts, its governance constitution from the resulting loss$\to$hazard$\to$UCA$\to$constraint chain, yielding a published constitution of $21$ Tool Principles and $8$ Meta-Safety Principles, each bound to a code enforcement point. We formalise the measured object as a constitution-marginal coverage operator over a principle set $P$ ($|P|{=}29$) with a soundness lemma that isolates coverage from model and scanner, and report four findings. {(i)~Self-derivation:} a frontier ensemble ({claude-opus-4.8}${+}${claude-sonnet-4}) recovers $18/21$ canonical and all $8/8$ governance principles from the tool's own design, while a weaker pair recovers $12/21$ and $3/8$, so the meta layer is model-limited, not constitution-limited, and the same $8/8$ re-emerge from a second, independently authored tool.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes