TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites.
Key points
- System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface.
- We assess probability quality through selective prediction, ordinal proper scoring rules, and realised cost under asymmetric cost matrices.
- We evaluate a hosted model, an open encoder, and a family of open decoders spanning 0.8B-9B parameters on identical items.
- Overall, typed decision models must be evaluated jointly on policy adherence, wording robustness, probability quality, and induced decision outcomes.
Sources (1)
- [1]TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:22 AM
We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites.
System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface.
Extractive summary: sentences quoted from the sources.