CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
Provides a high-fidelity, controllable generative model for 3D medical images, enabling synthesis of diverse chest CT volumes with specified clinical findings.
CONFLUX introduces a latent diffusion model for 3D chest CT synthesis conditioned on clinical attributes, achieving a tri-planar FID of 32.3 vs. 74.6 for MAISI. Post-training with RL reduces the reliability shortfall relative to real scans by 47%.
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning. We present CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space. Generation is conditioned on structured radiological metadata (18 abnormality findings, sex, age, and reconstruction kernel) through adaptive layer normalization. The model leads strong volumetric baselines on tri-planar Frechet distance (FID 32.3 vs. 74.6 for MAISI) while exposing direct control over clinical attributes. To strengthen that control we add an online reinforcement-learning post-training stage (group-relative policy optimization) that rewards how reliably a classifier recovers the requested findings from each generated volume. Judged by a separate, independent classifier, post-training removes 47% of the shortfall relative to real-scan reliability. We release the model and a ~200k synthetic chest-CT dataset with conditioning metadata spanning a wide variety of clinical findings.