AION
Opinion / analysisMultimodal Models1 source · Oct 7, 2026

Multimodal open d1 decision models for the edge

Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).

Key points

  • Multimodal: d1-3B supports text and images, while d1-omni-600M supports text and images or text and audio
  • Fast: d1-3B answers a question in 16 ms on an NVIDIA Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50ms on a Jetson Orin Nano
  • We benchmarked d1-3B and d1-omni-600M on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding. d1-3B achieves a mean score of 82.9, the highest in the table and above Decider 4B. d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with only a quarter of the parameters.
  • Reach for d1 decision models when you need fast, structured decisions, including multimodal inputs. d1-3B delivers the highest decision quality at its size, while d1-omni-600M fits where footprint matters.

Sources (1)

  • [1]Multimodal open d1 decision models for the edge
    Hugging Face Blog · Oct 7, 04:54 PM
    - Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
    - Multimodal: d1-3B supports text and images, while d1-omni-600M supports text and images or text and audio

Extractive summary: sentences quoted from the sources.