Jiaxin Wang

2papers

2 Papers

18.4ROJun 29
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models

Siyao Chen, Jiakang Yuan, Jiaxin Wang et al.

Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning. However, existing RL methods typically require external environmental feedback, relying on predefined success signals to guide policy updates. In this work, we show that VLA models possess useful internal evaluative capabilities: in discrete-action VLAs, trajectories with higher generation confidence are significantly more likely to succeed. Based on this observation, we introduce T^2VLA (Test-time VLA), an architecture-agnostic test-time RL framework that enables VLA models to achieve self-bootstrapping policy improvement. Instead of relying on external rewards, T^2VLA leverages trajectory-level similarity to high-confidence expert demonstrations as an intrinsic reward signal. In addition, we propose a Confidence-Driven Dual Expert Bootstrapping mechanism, which dynamically balances a Local Pseudo-Expert for exploration and a Global Expert Pool for training stability. Extensive experiments on the LIBERO and RoboTwin benchmarks show that T^2VLA consistently outperforms supervised baselines and approaches oracle RL performance with ground-truth rewards, achieving effective improvement without external reward feedback. Furthermore, T^2VLA adapts to distinct VLA paradigms, including both OpenVLA-OFT and the pi series.

7.4ITJun 29
New families of asymptotically optimal codebooks from vectorial dual-bent functions

Yadi Wei, Jiaxin Wang, Fang-Wei Fu et al.

Codebooks with small maximum cross-correlation amplitudes play an important role in many applications, such as code division multiple access (CDMA) communication systems, multiple-input multiple-output (MIMO) communications, compressed sensing, and coding theory. In this paper, by using vectorial dual-bent functions, we construct several families of codebooks that asymptotically achieve the Welch bound. The maximum cross-correlation amplitudes and the distributions of the cross-correlation amplitudes of the constructed codebooks are explicitly determined. Furthermore, these codebooks have new parameters, and some of them have very small alphabet sizes.