15.1LGJun 22
DiT-Reward: Generative Representations for Text-to-Image Reward ModelingYuanming Yang, Guoqing Ma, Bo Wang et al.
Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative representation learning. To this end, we introduce DiT-Reward, which converts a pretrained text-to-image Diffusion Transformer into a reward model by processing near-clean image latents and aggregating text-conditioned image representations across transformer layers. Under the same training data mixture as HPSv3, DiT-Reward outperforms HPSv3 on all four evaluated preference benchmarks, reaching 85.6% on HPDv2 and 77.6% on HPDv3. When the generative backbone is frozen, a lightweight learned head can still extract meaningful preference predictions from its representations. Probing across depth further reveals that downstream reward performance is strongest in the middle-to-late layers and benefits from combining representations across different stages. We also observe consistent positive scaling with generative backbone capacity. Finally, when used to optimize Stable Diffusion 3.5 Large with Flow-GRPO, DiT-Reward outperforms HPSv3 along the matched training trajectory, with particularly clear gains in realism. Direct latent scoring also achieves a 1.65x inference speedup over HPSv3 with comparable peak memory. These results show that pretrained generative DiTs provide transferable representations for reward modeling and policy optimization.
2.7SYJun 19
Nonholonomic directional pursuit and evasion: Global feedbacksBo Wang, Miroslav Krstic
In a recent paper by the second coauthor, directional pursuit-evasion for strictly forward-moving nonholonomic vehicles was solved "half-globally", namely under favorable initial line-of-sight conditions. In this paper, we develop feedback designs that achieve the directional pursuit and evasion objectives from arbitrary initial relative configurations. We achieve globality with completely different approaches to both the design and the analysis. Our designs are less aggressive in both the forward-speed and steering laws, allowing transient overshoot of the pursuer-evader range during global reorientation. The only price that we pay for globality is that our feedback laws require a priori knowledge of the opponent's maximal turning rate, whereas in the half-global work, no known bound of the opponent's turning rate was assumed. For the pursuit problem, the feedback law guarantees finite-time capture with prescribed directional alignment. For the evasion problem, the feedback law guarantees capture avoidance with a prescribed safety margin and achieves spinaway under a decay condition on the pursuer's turning rate. We illustrate the global capture and spinaway with simulations. The analysis is based on an integral input-to-state stability type mechanism induced by an endogenous time dilation, together with finite-time coextinction and safety-margin persistence lemmas for singularly coupled scalar inequalities.