CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology
In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction.
ProofPaper ↗
Key points
- In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models.
- If the evaluation reference changes while the intended target remains the original full-bag prediction, however, not only can the measured fidelity of the same compact evidence change, but comparisons between competing candidate strategies can also change.
- To make this dependence explicit, we introduce CHARTER, a reference-aware evaluation charter that asks researchers to DECLARE the intended target and reference, QUANTIFY candidate-induced prediction shift, and AUDIT the stability of comparative conclusions.
- Across the 15 comparisons in our main five-seed Random-K audit, 4 showed determinate reversals; in a matched native-ranking stress test, the ACMIL comparison changed from REVERSED to PRESERVED.
Sources (1)
- [1]CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational PathologyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:45 AM
In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction.
In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models.
Extractive summary: sentences quoted from the sources.