IRJun 26
R$^2$-Searcher: Calibrating Retrieval and Reasoning Boundaries for Agentic SearchSheng Zhang, Junyi Li, Wenlin Zhang et al.
Recent search agents for multi-hop reasoning often fail by either retrieving incomplete evidence or reasoning over irrelevant portions of the retrieved content, leading to a retrieval-reasoning boundary shift. We propose R$^2$-Searcher, a novel framework that explicitly explores and calibrates the retrieval and reasoning boundaries via fine-grained, query-token-guided evidence modeling and post-retrieval reflection. Specifically, R$^2$-Searcher: (1) constructs fine-grained reasoning contexts by extracting precise facts from retrieved content based on query token semantics (e.g., subjects, actions, temporal markers, and degree modifiers), thereby guiding the attention of search agent; (2) introduces a retrieval reflection mechanism that evaluates and corrects boundary deviations after each retrieval step, guiding the generation of improved queries grounded in the extracted reasoning contexts; and (3) employs an end-to-end reasoning-reflection-guided reinforcement learning algorithm, R$^2$PO, which jointly optimizes both boundaries through a tree-based exploration of reasoning regions and reflections. Our method significantly enhances the quality of both retrieval and reasoning, establishing an iterative loop where retrieval and reasoning mutually enhance each other. Extensive experiments on seven complex multi-hop QA benchmarks demonstrate that R$^2$-Searcher significantly outperforms state-of-the-art agentic search methods in answer accuracy and retrieval-reasoning quality. Ablation studies further confirm the critical role of retrieval-reasoning boundary calibration.
SYJun 26
Resilient Control Lyapunov Function-based Quadratic Program for Quadrotors Under CyberattacksYichao Wang, Sameeha Tasneem, Mohamadamin Rajabinezhad et al.
Ensuring the operational safety of quadrotors under partial actuator failures, lumped external disturbances, and malicious cyberattacks is a critical challenge due to the system's underactuated and highly nonlinear nature. Building on the existing result of a fault-tolerant control approach for a quadrotor experiencing a complete loss of two opposing rotors \cite{chen2024quadrotor}, this letter further addresses the additional challenge of malicious cyberattacks, which could be unknown and unbounded. While the baseline control law, rooted in proportional-derivative (PD) feedback and observer-based decoupling, effectively handles mismatched disturbances, it remains vulnerable to maliciously injected cyberattacks on the pseudo-control channels. To address this, a Resilient Control Lyapunov Function-based Quadratic Program (RCLF-QP) is developed, where a resilient compensational term with real-time online adaptation is designed in the conventional CLF to compensate for the maliciously injected unknown and unbounded attacks. Compared with the PD feedback control, the proposed QP-based constrained optimization control framework provides a systematic and extensible framework that allows new control objectives and constraints to be seamlessly integrated without altering the underlying stability guarantees. The overall proposed controller integrates a model-based extended state observer with the proposed RCLF-QP mechanism to mitigate both lumped disturbances caused by aerodynamics and strong wind, and adversarial cyberattacks injected by malicious adversaries. Simulations in a high-fidelity environment demonstrate that the proposed RCLF-QP control architecture prevents trajectory divergence and system instability in scenarios where the baseline controller fails in maintaining the stability of Quadrotors under malicious attacks.