MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction
Large-scale construction of medical vision-language model (VLM) benchmarks is increasingly feasible with richly annotated imaging datasets and large language models (LLMs), yet existing automation largely focuses on generating evaluation items within predefined benchmark specifications.
ProofPaper ↗
Key points
- We study the broader problem of automatically deriving the specification itself: what to evaluate, which annotations support each task, and how to translate this evidence into reliable evaluation items.
- We formulate benchmark construction as constrained compilation, in which the benchmark specification is progressively derived from evaluation requirements, heterogeneous annotations, and medical knowledge.
- Based on this formulation, we introduce MedBenchAgent, a multi-agent framework with a Benchmark Intermediate Representation (BIR) that encodes task definitions, evidence mappings, evaluation protocols, and item specifications across construction stages.
- These results establish constrained compilation as a scalable and auditable framework for medical VLM benchmark construction beyond question generation.
Sources (1)
- [1]MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark ConstructionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:17 AM
Large-scale construction of medical vision-language model (VLM) benchmarks is increasingly feasible with richly annotated imaging datasets and large language models (LLMs), yet existing automation largely focuses on generating evaluation items within predefined benchmark specifications.
We study the broader problem of automatically deriving the specification itself: what to evaluate, which annotations support each task, and how to translate this evidence into reliable evaluation items.
Extractive summary: sentences quoted from the sources.