MMLU
Also known as: MMLU-Pro
4stories this week
4last 30 days
4all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceSpecialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual TasksWe study how such a specialized decision model compares with general-purpose large language models (LLMs).
- Oct 8, 2026 · Research paper · 1 sourceDeflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion TransformersIn diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
- Oct 7, 2026 · Research paper · 1 sourcePatchBench: Measuring Collateral Damage in Activation PatchingTo address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.
- Oct 6, 2026 · Research paper · 1 sourcePenalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid ChoicesMultiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy.