Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He, Ziyang Ma, Qize Yang, Yunfei Chu, Jin Xu, Junyang Lin, Chi-Wing Fu, Pheng-Ann Heng
4 min
Abstract
Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although interactive frameworks have emerged, they often rely on global pre-scanning, and their context cost still scales with video length. We propose OmniAgent, the first native omni-modal agent that formulates video understanding as a POMDP-based iterative Observation-Thought-Action cycle. OmniAgent executes on-demand actions to selectively distill audio-visual cues into a persistent textual memory, effectively decoupling reasoning complexity from raw video duration. To operationalize this, we introduce (1) Agentic Supervised Fine-Tuning to bootstrap native active perception via best-of-N trajectory synthesis with dual-stage quality control, and (2) Agentic Reinforcement Learning with TAURA (Turn-aware Adaptive Uncertainty Rescaled Advantage), which leverages turn-level entropy to steer credit assignment toward pivotal discovery turns. Crucially, OmniAgent exhibits positive test-time scaling, where performance improves as the number of reasoning turns increases, validating the efficacy of active perception. Empirical results across ten benchmarks (e.g., VideoMME, LVBench) demonstrate that OmniAgent achieves state-of-the-art performance among open-source models. Notably, on LVBench, our 7B agent outperforms the 10$\times$ larger Qwen2.5-VL-72B (50.5% vs. 47.3%).
Sam: So it's trained on the whole messy process, not just the clean final answer.
Alex: Right. And on top of that, they developed a method called TAURA. The name is dense, but the idea is straightforward. Not all moments in a search are equally important. Finding a key piece of evidence is more significant than confirming something you already knew. TAURA teaches the model to weight those "discovery" moments more heavily during training, so it gets better at recognising when it's actually found something that matters.
Sam: Like learning to pay attention at the right moments, not just any moment.
Alex: That's a fair way to put it. There's also a step called a Rationality Audit, where a separate model checks whether the reasoning chain actually holds together logically. This prevents the agent from stumbling onto correct answers through lucky guesses rather than sound thinking — which would make it unreliable in practice.
Sam: Does all of this actually translate into better performance?
Alex: The study suggests it does. On a standard long-video benchmark called LVBench, a version of OmniAgent with around seven billion parameters — that's a rough measure of model size and complexity — outperformed a passive model roughly ten times larger. The smaller, more deliberate system beat the bigger, brute-force one.
Sam: That's a notable result. A much smaller model winning because of how it thinks, not how big it is.
Alex: Though the researchers are clear about the trade-off. This approach depends entirely on the quality of the model's reasoning at each step. If it draws a wrong conclusion early on — misreads a clue, so to speak — it may go looking in entirely the wrong place and never recover. The system is only as reliable as its own logic.
Sam: So it's more efficient, but also more fragile in a specific way.
Alex: That's a fair characterisation. It trades raw computational brute force for something more like directed intelligence — which is more powerful when it works, but requires the reasoning to hold up throughout. The paper frames this as a meaningful step toward AI systems that can handle long-form content by thinking carefully about where to look, rather than simply processing everything in sight. Thanks for listening to ResearchPod.