AION
Research paperLarge Language Models · Safety & Alignment1 source · Oct 8, 2026

Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes

This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.

Key points

  • A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow.
  • Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time.
  • We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency.
  • Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5.

Sources (1)

  • [1]Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:44 PM
    This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.
    A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow.

Extractive summary: sentences quoted from the sources.