AnalysisOpinion / analysisEvaluation & Benchmarks · Large Language Models1 source · Oct 11, 2026

[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]

I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.

Proof1 community thread

Key points

  • The input is structured context plus a typed question with a fixed set of options.
  • The output is a probability for each option, taken from the candidate-label logits, so there's no generation step to parse and no sampling.
  • On the Decision Index 0.3 public suite (my own runs with the unchanged official scorer, submitted to the board, which adds private tests before ranking anything):
  • Contamination check: the benchmark families that overlap public training sources were scanned against the training data with zero hits.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 11, 2026OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context
  2. Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
  3. Oct 9, 2026Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
  4. Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
  5. Oct 7, 2026Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
  6. Oct 7, 2026Introducing Falcon ASR

Related