ResearchPod Summary
Deep learning models for video action recognition often suffer from significant performance degradation when deployed in real-world environments due to distribution shifts (e.g., varying lighting, sensor noise, or compression). While Test-Time Adaptation (TTA) aims to mitigate this by adapting models to unlabeled target data during inference, existing methods often struggle with the temporal dynamics of video and the risk of catastrophic forgetting. This paper introduces Test-time Adaptation via Dual Distillation (TADD) to address these challenges in an online, source-free setting.
TADD utilizes a frozen CLIP backbone to provide a robust, domain-agnostic foundation. The core of the framework is a lightweight projection adapter that is the only component updated during inference. To ensure stable adaptation without losing the model's original discriminative power, the authors propose two complementary distillation losses:
This dual-loss objective allows the model to adapt to new, unlabeled video streams in real-time while maintaining a balance between general visual-language alignment and specific action-recognition accuracy.
The authors evaluated TADD on three challenging video action recognition benchmarks: UCF-HMDB, Daily-DA, and Sports-DA. The results demonstrate that TADD consistently outperforms state-of-the-art TTA baselines, achieving accuracy improvements of up to +3.81% on UCF-HMDB, +2.63% on Daily-DA, and +3.03% on Sports-DA. The framework shows particular robustness in closed-set scenarios and maintains high performance even when compared against offline Source-Free Domain Adaptation (SFDA) methods that have access to the entire target dataset.
This work provides a practical, efficient solution for deploying video recognition models in the wild. By avoiding the need for source data and operating in a strictly online, test-time fashion, TADD offers a scalable approach to maintaining model reliability in dynamic environments where data distributions are constantly shifting.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.