AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow
AuthorsSalah Eddine Bekhouche, Abdellah Zakaria Sellam, Fadi Dornaika, Abdenour Hadid
Resources
AffectFlow-DINO uses conditional rectified flow to represent ambiguity in facial emotion predictions while improving valence-arousal, expression, and Action Unit estimation.
Key results
Dimensional vector combining valence-arousal, expressions, and Action Units.
Improvement in CCC-V from rectified-flow decoding over deterministic decoding.
Fear F1 increased from 3.8% to 33.1% through post-hoc expression calibration.
Final validation score after backbone fine-tuning, flow retuning, and both calibrations.
Challenge baseline exceeded by the final AffectFlow-DINO configuration.
What the paper found
Researchers at the University of the Basque Country, the University of Salento and CNR, and Universiti Malaysia Kelantan introduce AffectFlow-DINO for the ABAW multi-task challenge on s-Aff-Wild2. The model uses a frozen DINOv3 ViT-S/16 backbone and a conditional rectified-flow head to learn a distribution rather than a single estimate over a 22-dimensional affect vector combining valence-arousal, eight facial expressions, and twelve Action Units. Masked supervision handles incomplete annotations, while Monte Carlo flow sampling captures ambiguity in in-the-wild facial behavior. Relative to deterministic decoding, rectified-flow decoding improves valence CCC-V by 0.058, although backbone fine-tuning becomes the largest performance lever. Post-hoc per-AU and per-expression calibration addresses severe class imbalance without retraining, raising Fear F1 from 3.8% to 33.1% and improving the final composite PMTL to 1.177, compared with the official baseline of 0.45. The study also finds that flow and deterministic objectives are complementary, inference saturates rapidly with few Euler steps, and flow retuning is necessary after backbone adaptation; nevertheless, deterministic decoding remains stronger than flow decoding in the fully fine-tuned variants.
Original abstract
We present \textbf{AffectFlow-DINO}, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior. Instead of predicting a single affect estimate, the model learns a conditional generative distribution, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling. The system jointly estimates continuous valence-arousal, classifies eight facial expressions, and detects twelve Action Units from static face images. Built on a frozen DINOv3 ViT-S/16 backbone, extensive ablation studies show that rectified-flow decoding consistently improves deterministic prediction, particularly for valence-arousal estimation (CCC-V $+0.058$). We further show that post-hoc threshold calibration effectively recovers performance on severely imbalanced rare classes (e.g., Fear: $3.8\% \rightarrow 33.1\%$) without retraining. Combined with backbone fine-tuning and flow retuning, the final model achieves $\mathbf{P_{MTL}=1.177}$, substantially outperforming the official challenge baseline of $P_{MTL}=0.45$.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.