AICLCRJul 22

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

arXiv:2607.199139.7
Predicted impact top 30% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For developers of autonomous agents, Janus offers a foresight-oriented safety framework that reduces operational failures without sacrificing task performance.

Janus trains guards to anticipate delayed risks from partial trajectories in long-horizon agent tasks, improving average protection by 15.9 percentage points and benign task completion by 5.1 percentage points over baseline guards.

Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution. Across four agent-safety benchmarks, Vanguard improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes