ROAIJun 18

Temporal Self-Imitation Learning

arXiv:2606.197524.8
Predicted impact top 76% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For robot manipulation, TSIL provides a self-supervisory signal from temporal efficiency, reducing reliance on reward shaping.

TSIL mines temporally efficient successful trajectories during RL training and uses them as self-supervision, improving learning efficiency, task-completion efficiency, and robustness across 15 long-horizon manipulation tasks.

Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while rare efficient behaviors may be forgotten during training. We argue that temporal efficiency itself provides a powerful and underutilized source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors through efficiency-weighted self-imitation learning. Across 15 distinct long-horizon manipulation tasks, TSIL consistently improves learning efficiency, task-completion efficiency, revisitation of fast successful behaviors, and robustness to unstable training conditions. More broadly, our results suggest that the temporal structure of successful behavior itself provides a scalable self-supervisory signal for reinforcement learning beyond manually engineered reward shaping alone.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes