ResearchPod Summary
As vision-language models (VLMs) are increasingly deployed in safety-critical applications, their vulnerability to adversarial prompting has become a significant concern. While many benchmarks exist, there is a lack of large-scale, ready-to-use adversarial datasets that cover diverse harmful intents and multimodal attack strategies. This paper addresses this gap by introducing PHANTOM, a resource designed to lower the barrier for researchers to evaluate VLM robustness and develop defensive guardrails.
The authors constructed PHANTOM by consolidating harmful intents from multiple existing benchmarks and introducing a new category for Child Safety. The final dataset includes 7,826 distinct harmful intents categorized into 10 high-level domains and 55 subcategories. Using five state-of-the-art attack strategies—BAP, IDEATOR, MML, FC ATTACK, and CSDJ—the researchers generated 47,524 adversarial image-text pairs. These samples were tested against a variety of open-source VLMs and evaluated for transferability to proprietary, black-box models using the Abel-24-HarmClassifier to determine attack success rates (ASR).
PHANTOM provides a comprehensive, structured resource for VLM safety research. By consolidating and cleaning existing benchmarks, the authors created a diverse set of adversarial samples that span a wide range of potential risks, from cybersecurity threats to ethical and social issues. The inclusion of a dedicated Child Safety category addresses a critical, previously underrepresented area of concern. The dataset is designed to be practical for researchers, offering pre-generated samples that eliminate the high computational costs typically associated with generating large-scale multimodal adversarial attacks, thereby enabling more reproducible and comparable safety evaluations.
This dataset is a significant contribution to the field of AI safety, as it provides a standardized, large-scale benchmark for testing the robustness of multimodal models. By making these adversarial samples open-source, the authors enable resource-constrained research groups to perform rigorous safety testing that was previously impractical. This work supports the development of more resilient AI systems and helps practitioners better understand the cross-modal vulnerabilities that current alignment techniques may fail to address.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.