ResearchPod Summary
Micro-expression recognition (MER) is notoriously difficult due to the brevity and subtlety of facial movements. Researchers often rely on optical flow (to capture motion) and motion magnification (to amplify appearance changes). However, these modalities frequently suffer from asymmetric failure: one may be noisy or distorted while the other remains informative. The authors investigate how to effectively fuse these heterogeneous modalities while accounting for their spatially varying reliability and the complex, non-linear relationship between facial muscle movements (Action Units) and emotion categories.
To address these challenges, the authors propose SAC2-Net, which follows an "align-first, fuse-later" strategy.
SAC2-Net achieves state-of-the-art or highly competitive performance across five standard MER benchmarks. The authors demonstrate that by explicitly modeling the asymmetric failure patterns of optical flow and motion magnification, the model can effectively suppress modality-specific noise. The use of AU-based soft labels is shown to be superior to traditional hard contrastive learning, as it better captures the nuanced semantic structure of micro-expressions.
This work provides a robust framework for multimodal fusion in scenarios where input modalities are inherently unreliable or heterogeneous. By leveraging semantic knowledge (via text prompts) to guide the alignment of visual features, the approach offers a scalable way to handle limited training data in affective computing, where the relationship between physical signals and subjective emotions is often ambiguous.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.