Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.
Key points
- A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow.
- Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time.
- We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency.
- Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5.
Sources (1)
- [1]Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug FixesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:44 PM
This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment.
A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow.
Extractive summary: sentences quoted from the sources.