AION
Research paperSafety & Alignment · Evaluation & Benchmarks1 source · Oct 6, 2026

Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair

We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes.

Key points

  • Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost.
  • We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects.
  • Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs.
  • These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.

Sources (1)

  • [1]Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:27 AM
    We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes.
    Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost.

Extractive summary: sentences quoted from the sources.