ResearchPod Summary
Multimodal Large Language Models (Omni-MLLMs) are increasingly used for Multimodal Emotion Reasoning (MER), where they must generate emotion predictions grounded in visual and acoustic signals. However, existing models often suffer from two critical failures: underutilization of multimodal cues and unfaithfulness, where models hallucinate modality-specific details (e.g., describing a facial expression) based on cues from a different modality (e.g., audio tone) rather than the actual visual input.
The authors propose Omni-Perception Policy Optimization (OPPO), a reinforcement learning framework designed to enforce grounded reasoning. OPPO introduces two primary mechanisms:
To evaluate these capabilities, the authors introduce MEP-Bench, a diagnostic benchmark that quantifies both the recall of multimodal cues (utilization) and the model's ability to maintain faithfulness under unimodal masking.
Experiments demonstrate that OPPO achieves state-of-the-art performance on standard benchmarks like MER-UniBench and MME-Emotion. More importantly, diagnostic results on MEP-Bench show that OPPO significantly outperforms baseline models in both utilization (higher recall of multimodal cues) and faithfulness (higher accuracy when probed with masked inputs). This confirms that explicitly optimizing for perception-grounded reasoning is essential for reliable multimodal emotion analysis.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.