30.0CVSep 30, 2024
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuningHaotian Zhang, Mingfei Gao, Zhe Gan et al.
We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development.
9.9SYApr 4
Hybrid Voltage-Current Control of Grid-Forming and Grid-Following InvertersZirui Wang, Yitong Li, Quanchi Wu et al.
Grid-connected inverters are required to operate stably under a wide range of grid conditions. However, conventional grid-following (GFL) control may suffer from instability under weak-grid conditions, while grid-forming (GFM) control may exhibit unstable oscillations under strong-grid conditions. To address these issues, a hybrid voltage-current control method is proposed in this article. A voltage control is introduced on the d-axis, while a current control is adopted on the q-axis, enabling the inverter to exhibit voltage-source characteristics on the d-axis and current-source characteristics on the q-axis. In this way, the proposed control integrates the characteristics of both conventional GFL and GFM control. A full-order model is established to analyze the port characteristics and small-signal stability of the systems. Finally, the effectiveness of the proposed control strategy is validated through simulations and experiments on a 1.5 kW inverter experimental platform. The results show that the proposed control maintains stable operation under different grid conditions with varying short-circuit ratios (SCRs).
11.8SYApr 5
Ideally-Smooth Transition between Grid-Forming and Grid-Following Inverters based on State Mapping MethodZhenshuai Liu, Yitong Li, Zirui Wang et al.
There has been widespread global increasing use of renewable energy sources, which are usually connected to the electricity grids via power electronic inverters. Traditionally, these inverter-based resources operate in either grid-forming (GFM) or grid-following (GFL) mode. But more recently, the need of switching between these two modes are glowingly required because of the complex operation scenarios of systems such as source-side limitations, grid-side services, fault disturbances, etc. However, due to the differences between GFM and GFL modes, a direct switching between them would lead to large oscillations or even instability of inverters. Therefore, in this paper, a method called state mapping method for analyzing the switching transient and designing the switching control is proposed. Based on this method, an ideally-smooth transition between GFM and GFL can be achieved. The effectiveness of the proposed method is verified by both the theoretical analysis and experiment tests.
9.4LGSep 30, 2025
Learning to Reason as Action Abstractions with Scalable Mid-Training RLShenao Zhang, Donghan Yu, Yihao Feng et al.
Large language models excel with reinforcement learning (RL), but fully unlocking this potential requires a mid-training stage. An effective mid-training phase should identify a compact set of useful actions and enable fast selection among them through online RL. We formalize this intuition by presenting the first theoretical result on how mid-training shapes post-training: it characterizes an action subspace that minimizes both the value approximation error from pruning and the RL error during subsequent planning. Our analysis reveals two key determinants of mid-training effectiveness: pruning efficiency, which shapes the prior of the initial RL policy, and its impact on RL convergence, which governs the extent to which that policy can be improved via online interactions. These results suggest that mid-training is most effective when the decision space is compact and the effective horizon is short, highlighting the importance of operating in the space of action abstractions rather than primitive actions. Building on these insights, we propose Reasoning as Action Abstractions (RA3), a scalable mid-training algorithm. Specifically, we derive a sequential variational lower bound and optimize it by iteratively discovering temporally-consistent latent structures via RL, followed by fine-tuning on the bootstrapped data. Experiments on code generation tasks demonstrate the effectiveness of our approach. Across multiple base models, RA3 improves the average performance on HumanEval and MBPP by 8 and 4 points over the base model and the next-token prediction baseline. Furthermore, RA3 achieves faster convergence and higher asymptotic performance in RLVR on HumanEval+, MBPP+, LiveCodeBench, and Codeforces.