NTH

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

AuthorsSalah Eddine Bekhouche, Abdellah Zakaria Sellam, Fadi Dornaika, Abdenour Hadid

July 18, 2026 2 min read
Watch on YouTube
The one-line take

AffectFlow-DINO uses conditional rectified flow to represent ambiguity in facial emotion predictions while improving valence-arousal, expression, and Action Unit estimation.

Key results

22
Joint affect representation

Dimensional vector combining valence-arousal, expressions, and Action Units.

0.058
Valence CCC gain

Improvement in CCC-V from rectified-flow decoding over deterministic decoding.

33.1%
Fear F1 after calibration

Fear F1 increased from 3.8% to 33.1% through post-hoc expression calibration.

1.177
Best PMTL

Final validation score after backbone fine-tuning, flow retuning, and both calibrations.

0.45
Official baseline PMTL

Challenge baseline exceeded by the final AffectFlow-DINO configuration.

What the paper found

Researchers at the University of the Basque Country, the University of Salento and CNR, and Universiti Malaysia Kelantan introduce AffectFlow-DINO for the ABAW multi-task challenge on s-Aff-Wild2. The model uses a frozen DINOv3 ViT-S/16 backbone and a conditional rectified-flow head to learn a distribution rather than a single estimate over a 22-dimensional affect vector combining valence-arousal, eight facial expressions, and twelve Action Units. Masked supervision handles incomplete annotations, while Monte Carlo flow sampling captures ambiguity in in-the-wild facial behavior. Relative to deterministic decoding, rectified-flow decoding improves valence CCC-V by 0.058, although backbone fine-tuning becomes the largest performance lever. Post-hoc per-AU and per-expression calibration addresses severe class imbalance without retraining, raising Fear F1 from 3.8% to 33.1% and improving the final composite PMTL to 1.177, compared with the official baseline of 0.45. The study also finds that flow and deterministic objectives are complementary, inference saturates rapidly with few Euler steps, and flow retuning is necessary after backbone adaptation; nevertheless, deterministic decoding remains stronger than flow decoding in the fully fine-tuned variants.

Original abstract

We present \textbf{AffectFlow-DINO}, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior. Instead of predicting a single affect estimate, the model learns a conditional generative distribution, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling. The system jointly estimates continuous valence-arousal, classifies eight facial expressions, and detects twelve Action Units from static face images. Built on a frozen DINOv3 ViT-S/16 backbone, extensive ablation studies show that rectified-flow decoding consistently improves deterministic prediction, particularly for valence-arousal estimation (CCC-V $+0.058$). We further show that post-hoc threshold calibration effectively recovers performance on severely imbalanced rare classes (e.g., Fear: $3.8\% \rightarrow 33.1\%$) without retraining. Combined with backbone fine-tuning and flow retuning, the final model achieves $\mathbf{P_{MTL}=1.177}$, substantially outperforming the official challenge baseline of $P_{MTL}=0.45$.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis