Yao Zou

h-index29
2papers
2,852citations

2 Papers

1.2SYNov 9, 2017
Coordinated trajectory tracking of multiple vertical take-off and landing UAVs

Yao Zou, Ziyang Meng

This paper investigates the coordinated trajectory tracking problem of multiple vertical takeooff and landing (VTOL) unmanned aerial vehicles (UAVs). The case of unidirectional information flow is considered and the objective is to drive all the follower VTOL UAVs to accurately track the trajectory of the leader. Firstly, a novel distributed estimator is developed for each VTOL UAV to obtain the leader's desired information asymptotically. With the outputs of the estimators, the solution to the coordinated trajectory tracking problem of multiple VTOL UAVs is transformed to individually solving the tracking problem of each VTOL UAV. Due to the under-actuated nature of the VTOL UAV, a hierarchical framework is introduced for each VTOL UAV such that a command force and an applied torque are exploited in sequence, then the position tracking to the estimated desired position and the attitude tracking to the command attitude are achieved. Moreover, an auxiliary system with proper parameters is implemented to guarantee the singularity-free command attitude extraction and to obviate the use of the unavailable desired information. The stability analysis and simulations effectively validate the achievement of the coordinated trajectory tracking of multiple VTOL UAVs with the proposed control approach.

9.4LGSep 2, 2025
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo et al.

Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundation Models (VFMs) and Multimodal Large Language Models (MLLMs), enhance semantic understanding but suffer from high computational overhead and limited cross-domain transferability when fine-tuned. To address these limitations, we propose a novel zero-shot framework employing Test-Time Scaling (TTS) that leverages MLLMs' vision-language alignment capabilities through Guidance-based methods for direct similarity scoring. Our approach eliminates two-stage processing by employing structured prompts that generate length-controllable JSON outputs. The TTS framework with Uncertainty-Aware Self-Consistency (UASC) enables real-time adaptation without additional training costs, achieving superior generalization across diverse environments. Experimental results demonstrate significant improvements in cross-domain VPR performance with up to 210$\times$ computational efficiency gains.