AICLJul 24

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

arXiv:2607.2208332.91 citations
Predicted impact top 1% in AI · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a compact model with strong agentic capabilities, enabling efficient local deployment for personal assistants.

Nanbeige4.2-3B is a 3B-parameter agentic model that outperforms larger models like Qwen3.5-9B and Gemma4-12B on agentic benchmarks while maintaining competitive reasoning abilities.

We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks while remaining competitive on reasoning and alignment tasks. Performance with OpenClaw further supports its use as a compact local personal assistant.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes