ROJun 24

Visual-Language-Guided Task Planning for Horticultural Robots

arXiv:2601.1190613.7h-index: 10
Predicted impact top 23% in RO · last 90 daysOriginality Incremental advance
AI Analysis

It provides a deployable framework and benchmark for VLM-guided agricultural robotics, highlighting limitations in long-horizon reasoning and noisy context grounding.

The paper introduces a modular framework using a Vision Language Model (VLM) for robotic task planning in horticulture, achieving 87% success on short-horizon tasks but under 10% on complex long-horizon tasks, with task completion above 76% under noiseless conditions.

Crop monitoring is essential for precision agriculture, but current systems lack high-level reasoning. We introduce a novel, modular framework that uses a Vision Language Model (VLM) to guide robotic task planning by actively querying heterogeneous data sources, including enriched RGB camera feeds and 2D semantic occupancy maps, interleaved with robotic action primitives. We contribute a comprehensive benchmark for short- and long-horizon crop monitoring tasks in monoculture and polyculture environments. Our results show that while zero-shot VLMs perform robustly for short-horizon tasks (achieving 87% success, comparable to human experts), success drops significantly to under 10% for complex long-horizon, multi-target tasks. Despite this decline, task completion rates remain above 76% under noiseless conditions. Critically, the system degrades when relying on noisy semantic maps, demonstrating a key limitation in current VLM context grounding for sustained robotic operations. This work offers a deployable framework and critical insights into VLM capabilities and shortcomings for complex agricultural robotics.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes