ResearchPod Summary
Multimodal automated fact-checking (MAFC) systems are designed to verify claims by retrieving and reasoning over external evidence. A major challenge in evaluating these systems is data contamination, where LLMs can verify claims using their internal, pre-trained knowledge rather than performing the intended retrieval and reasoning tasks. While researchers have shifted toward dynamic benchmarks—which use claims published after an LLM's knowledge cut-off date—the effectiveness of this approach remains under-examined. This study investigates the extent of contamination in both static (AVeriTeC) and dynamic (ClaimReview2025Q4) benchmarks and its impact on performance evaluation.
The authors find that dynamic evaluation is not a panacea for contamination. Even for claims published after an LLM's knowledge cut-off, a substantial portion (17.09%–29.30%) remains potentially contaminated. This occurs because many new claims can be verified by synthesizing public knowledge that was already available before the cut-off. The study demonstrates that this contamination is not merely a theoretical concern; it significantly inflates Macro-F1 scores and distorts the relative rankings of different MAFC systems. When the authors re-evaluated models on a strictly contamination-controlled subset, they found that all tested SOTA models performed below 56% Macro-F1, revealing that current systems are less capable than previously reported.
This research highlights a critical vulnerability in how we measure progress in AI-driven fact-checking. By showing that even "dynamic" benchmarks are susceptible to déjà vu effects, the authors provide a cautionary tale for researchers relying on these datasets to claim state-of-the-art performance. The study offers practical guidelines for building more robust evaluation pipelines, emphasizing the need for rigorous, contamination-controlled testing to ensure that MAFC systems are truly capable of handling novel, real-world misinformation.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.