Closed-loop evaluation of LLM agents for embedded software development
We present a benchmark for closed-loop evaluation of embedded coding agents.
Key points
- Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively.
- Embedded firmware is a demanding target because correctness depends on closed-loop behavior under sensing, timing, and safety constraints, not only on static source quality.
- The implementation targets simulated ESP32 firmware for reproducibility.
- We evaluate seven GPT-family and Qwen-family configurations across five tasks and four scenarios, with three repetitions per condition for 420 runs. gpt-5.4 has the highest pass rate among evaluated configurations but does not saturate the benchmark; qwen3.5-27B is the strongest observed local model; and smaller local models degrade sharply in pass rate and search efficiency.
Sources (1)
- [1]Closed-loop evaluation of LLM agents for embedded software developmentarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:02 AM
We present a benchmark for closed-loop evaluation of embedded coding agents.
Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively.
Extractive summary: sentences quoted from the sources.