AION
Benchmark

MMLU

Also known as: MMLU-Pro

4stories this week
4last 30 days
4all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks
    We study how such a specialized decision model compares with general-purpose large language models (LLMs).
  2. Oct 8, 2026 · Research paper · 1 source
    Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
    In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
  3. Oct 7, 2026 · Research paper · 1 source
    PatchBench: Measuring Collateral Damage in Activation Patching
    To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers.
  4. Oct 6, 2026 · Research paper · 1 source
    Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices
    Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy.

Often appears with