False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.
Key points
- Image-generation models can now produce text-rich, natural-looking visual artifacts that are hard to distinguish from real-world evidence, such as news reports and textbook pages.
- Yet, the same capability introduces a new risk: these models can just as easily fabricate visual misinformation.
- Curiously, we find that these models can recognize a claim as false when asked, yet still render that very claim as credible visual evidence.
- We further introduce EpiReal-Attack, a skill-guided black-box optimization framework that uses Pareto-based selection and multimodal feedback to identify commands that bypass alignment safeguards while preserving visual realism, textual legibility, and semantic fidelity.
Sources (1)
- [1]False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image GeneratorsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:35 AM
To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats.
Image-generation models can now produce text-rich, natural-looking visual artifacts that are hard to distinguish from real-world evidence, such as news reports and textbook pages.
Extractive summary: sentences quoted from the sources.