Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction
Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language.
Key points
- In next-visit diagnosis prediction, multiple diagnoses can be simultaneously valid, but independently rewarding one diagnosis per trajectory does not distinguish repeated hits from coverage of different diagnoses.
- To address both challenges, we propose CARing, a framework that represents diagnoses with compositional Semantic IDs (SIDs) and optimizes reasoning trajectories for multi-label coverage.
- Concretely, we first encode ontology-enriched disease semantics into compact SIDs through residual quantization, and ground the resulting SID tokens in natural language and longitudinal EHR contexts through multi-task alignment and reasoning-enriched training to unlock transferable LLM reasoning.
- CARing further improves unordered multi-label prediction through a coverage reward for reinforcement learning and multi-positive supervision.
Sources (1)
- [1]Coverage-Aware Reasoning with Medical Tokens for Diagnosis PredictionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:39 PM
Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language.
In next-visit diagnosis prediction, multiple diagnoses can be simultaneously valid, but independently rewarding one diagnosis per trajectory does not distinguish repeated hits from coverage of different diagnoses.
Extractive summary: sentences quoted from the sources.