ResearchResearch paperInterpretability · Computer Vision · Multimodal Models1 source · Oct 7, 2026

MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection

In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues.

Key points

  • Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions.
  • For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information.
  • To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map.
  • We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks.

Sources (1)

  • [1]MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:30 PM
    In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues.
    Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
  2. Oct 7, 2026Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
  3. Oct 6, 2026FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
  4. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  5. Oct 1, 2026nvidia/PixelUMM
  6. Oct 1, 2026RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

Related