ResearchResearch paperLarge Language Models · Efficiency & Inference · Safety & Alignment1 source · Oct 6, 2026

Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs

Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated.

Key points

  • We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts.
  • The 8-bit-4-bit recovery comparison changes direction across prompts and evaluation targets.
  • Full-pipeline point estimates favor 8-bit Llama under all five prompts, whereas the Qwen comparison changes direction across prompts.
  • These findings show that one prompt, one screened task set, and one scoring policy do not establish a stable conclusion about quantized-agent robustness.

Sources (1)

  • [1]Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:17 AM
    Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated.
    We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts.

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 29, 2026NVIDIA/TensorRT-LLM v1.3.0rc29
  2. Sep 28, 2026Holo4: powering generalist computer-use agents
  3. Sep 22, 2026vllm-project/vllm v0.30.0
  4. Aug 10, 2026vllm-project/vllm v0.27.0
  5. Jun 29, 2026vllm-project/vllm v0.24.0
  6. Jun 15, 2026vllm-project/vllm v0.23.0

Related