ResearchResearch paperRobotics & Embodied AI · Multimodal Models · Computer Vision1 source · Oct 6, 2026

A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification.

Key points

  • Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature 0, across eight pipeline configurations on one backbone, 16--37% of clips change their predicted label between repeated runs, and the cause lies in the serving stack.
  • On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC $0.851 \pm 0.012$, 86.0% accuracy, and F$1$ 71.8 over three runs.
  • It labels 74.4% of clips correctly in every run (60.5% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot).
  • An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.

Sources (1)

  • [1]A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:25 PM
    Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification.
    Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature 0, across eight pipeline configurations on one backbone, 16--37% of clips change their predicted label between repeated runs, and the cause lies in the serving stack.

Extractive summary: sentences quoted from the sources.

Related