Identity-Duplication Auditing in National-Scale Neuroimaging Repositories
In this work, we present HAPPEN, a human-in-the-loop pipeline for auditing identity duplication in T1-weighted brain MRI repositories.
ProofPaper ↗
Key points
- National-scale magnetic resonance imaging (MRI) repositories increasingly integrate data from different studies and institutions.
- However, subject identifiers that are valid only within individual datasets are no longer guaranteed to remain globally unique after aggregation, making it possible for the same subject to be assigned multiple identifiers, which we define as identity duplication.
- It combines SHA-256 fingerprinting for exact-duplicate detection with supervised contrastive retrieval of non-identical scans that may originate from the same person.
- We deployed the workflow in a 95,129-scan aggregated repository and assessed end-to-end recovery using 54 genetic-reference pairs.
Sources (1)
- [1]Identity-Duplication Auditing in National-Scale Neuroimaging RepositoriesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:57 AM
In this work, we present HAPPEN, a human-in-the-loop pipeline for auditing identity duplication in T1-weighted brain MRI repositories.
National-scale magnetic resonance imaging (MRI) repositories increasingly integrate data from different studies and institutions.
Extractive summary: sentences quoted from the sources.