AION
Open-source releaseLarge Language Models1 source · Oct 6, 2026

huggingface/transformers v5.19.0: Release v5.19.0

EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.

Key points

  • All MoE models whose routers compute logits now return them when outputrouterlogits=True, following the Qwen3-MoE pattern (a routerlogits recorder on the base model, MoeModelOutputWithPast from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.
  • Owlv2ForObjectDetection.embedimagequery now selects the query box with the highest objectness score, as in the original OWLv2 notebook, instead of the OWL-ViT heuristic, so image-guided query embeddings and detections may differ from earlier releases.
  • The regular flash and SDPA attention functions (flashattention.py, sdpaattention.py) now support continuous batching directly, and "paged|..." implementations for these are redirected to them, while eager still requires the "paged|eager" prefix.
  • 🚨 [CB] 🚨 Fuse update for index and block table path (#49088) by @remi-or

Sources (1)

  • [1]huggingface/transformers v5.19.0: Release v5.19.0
    GitHub: huggingface/transformers · Oct 6, 04:39 PM
    EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
    All MoE models whose routers compute logits now return them when `output_router_logits=True`, following the Qwen3-MoE pattern (a `router_logits` recorder on the base model, `MoeModelOutputWithPast` from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.

Extractive summary: sentences quoted from the sources.