CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
AuthorsMax Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
Resources
This paper trains a 3D latent generative model to synthesize realistic chest CT scans with controllable clinical findings, then uses reinforcement learning to make the requested attributes appear more reliably.
Key results
Compared with 74.6 for MAISI and 145.4 for GenerateCT.
One-scan-per-patient CT-RATE subset used to train the flow model.
Independent image-space judge score, up from 0.330 before RL.
Independent judge score, up from 0.684 before RL.
Fraction of the pre-RL-to-real-scan gap recovered after GRPO.
Approximate number of metadata-linked synthetic chest-CT volumes.
What the paper found
Researchers at Stanford University and Ghent University present CONFLUX, a controllable, natively 3D chest-CT generator that combines a 3D variational autoencoder with a rectified-flow transformer operating in latent space. The model compresses CT volumes by a factor of 8 into 16-channel latents and uses adaptive layer normalization to condition generation on 18 abnormality findings, sex, age range, and reconstruction kernel. Its design relates to FLUX from Black Forest Labs, while its group-relative policy optimization, or GRPO, adapts reinforcement-learning ideas popularized by DeepSeekMath. Trained on 18,417 CT-RATE scans, CONFLUX achieves a tri-planar FID of 32.3, compared with 74.6 for MAISI and 145.4 for GenerateCT. The paper’s main novelty is RL post-training: a frozen classifier rewards whether generated volumes realize their requested findings, and an independent image-space judge shows macro average precision increasing from 0.330 to 0.344 and macro AUROC from 0.684 to 0.699. This recovers 47% of the reliability gap between pre-RL synthesis and real scans, without improving unrelated sex, age, or kernel conditioning. The authors release the model and approximately 200K metadata-linked synthetic chest-CT volumes spanning diverse clinical findings.
Original abstract
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning. We present CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space. Generation is conditioned on structured radiological metadata (18 abnormality findings, sex, age, and reconstruction kernel) through adaptive layer normalization. The model leads strong volumetric baselines on tri-planar Frechet distance (FID 32.3 vs. 74.6 for MAISI) while exposing direct control over clinical attributes. To strengthen that control we add an online reinforcement-learning post-training stage (group-relative policy optimization) that rewards how reliably a classifier recovers the requested findings from each generated volume. Judged by a separate, independent classifier, post-training removes 47% of the shortfall relative to real-scan reliability. We release the model and a ~200k synthetic chest-CT dataset with conditioning metadata spanning a wide variety of clinical findings.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.