LGAIROJul 10

Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

arXiv:2607.0933610.4h-index: 4
Predicted impact top 25% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners of offline reinforcement learning, STP offers a more efficient generative planning method that reduces both training and inference costs without sacrificing performance.

STP introduces a single-stage shortcut trajectory model for offline RL that enables adjustable one-step or few-step inference, achieving strong performance across D4RL benchmarks while simplifying training and reducing inference cost compared to diffusion-based planners.

Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher--student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based reinforcement learning framework that incorporates shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, supports adjustable one-step and few-step inference through step-size conditioning, and selects candidate plans using a critic augmented with feasibility-aware correction. Across standard D4RL benchmarks, including locomotion, navigation, manipulation, and dexterous control tasks, STP achieves strong performance while simplifying the training pipeline for fast generative planning.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes