CRAIAug 12

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

arXiv:2608.1227312.8
Predicted impact top 7% in CR · last 90 daysOriginality Highly original
AI Analysis

This work identifies a new vulnerability in skill-based LLM agents, demonstrating that correct task completion does not guarantee cost safety or trajectory integrity, which is a problem for users and operators of such agents.

This paper introduces Convergent Detour Hijacking (CDH), a text-only attack that manipulates LLM agents by coupling skill selection and planning stages. It causes agents to take unnecessarily costly trajectories while preserving task completion, leading to a 66.91% increase in token consumption and 92.45% increase in execution time on DeepSeek-V4-Pro, with 80.02% of tasks selecting the attacker-controlled coordinator.

LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unnecessarily costly trajectory. Prior work studies selection manipulation, malicious skill instructions, and tool-chain resource amplification largely separately, leaving their end-to-end composition unclear. We introduce Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack that couples these stages. Under shared semantic cover, a description establishes relevance during selection, while an aligned body reuses that rationale to fabricate plausible dependencies during planning. CDH attracts an attacker-controlled coordinator alongside legitimate skills, recruits unnecessary benign skills into a bounded detour, and then re-enters the original route to preserve task completion. We evaluate it across multiple LLM backends and 491 held-out tasks under single-task and multi-turn conditions. On DeepSeek-V4-Pro, the matched coordinator is selected in 80.02% of tasks; among coordinator-hit runs that complete tasks, token consumption and end-to-end execution time increase by 66.91% and 92.45%, respectively, while aggregate task completion remains comparable. Thus, correct outcomes do not guarantee trajectory integrity or cost safety.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes