AnalysisOpinion / analysisEfficiency & Inference · MLOps, Tooling & Infrastructure · Large Language Models1 source · Oct 10, 2026

Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work

Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).

Proof1 community thread

Key points

  • Today I decided to do a few tests on things that I'd consider analogous to "real" work that I do, trying out some different combinations of models and harnesses.
  • Harnesses: Codex CLI vs OpenCode CLI
  • Types of Work: "Tell me about this code" and "Let's make something new"
  • I thought of an application that's concisely scoped, doesn't rely on a legacy codebase or significant external dependencies, and would probably take me ~1-2 days to hammer out by hand. ~8 paragraphs of prompt covering general concept, user experience in a couple different roles, design constraints and future-proofing needs.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 9, 2026Microsoft's Decision-1 model enters the fast-growing AI decision model race
  2. Oct 9, 2026A new feature for my blog, built using my voice
  3. Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
  4. Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
  5. Oct 8, 2026A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
  6. Oct 7, 2026[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing

Related