A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video
Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification.
ProofPaper ↗
Key points
- Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature 0, across eight pipeline configurations on one backbone, 16--37% of clips change their predicted label between repeated runs, and the cause lies in the serving stack.
- On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC $0.851 \pm 0.012$, 86.0% accuracy, and F$1$ 71.8 over three runs.
- It labels 74.4% of clips correctly in every run (60.5% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot).
- An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.
Sources (1)
- [1]A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home VideoarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:25 PM
Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification.
Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature 0, across eight pipeline configurations on one backbone, 16--37% of clips change their predicted label between repeated runs, and the cause lies in the serving stack.
Extractive summary: sentences quoted from the sources.