AION
Open-source releaseLarge Language Models · Speech & Audio · Multimodal Models1 source · Sep 30, 2026

huggingface/transformers v5.18.0: Release 5.18.0

Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.

Key points

  • A single checkpoint supports configurable latency profiles, from an 80 ms input buffer to a 30.4 s offline-style buffer, and configurable output frame resolution in multiples of 10 ms.
  • With chunked inference, the maximum audio duration is not limited.
  • embedding space at the <image> / <video> context-token positions; audio clips are projected in the same way at
  • HyperCLOVAX Vision V2 is a multimodal vision-language model developed by NAVER.

Sources (1)

  • [1]huggingface/transformers v5.18.0: Release 5.18.0
    GitHub: huggingface/transformers · Sep 30, 04:46 PM
    Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio.
    A single checkpoint supports configurable latency profiles, from an 80 ms input buffer to a 30.4 s offline-style buffer, and configurable output frame resolution in multiples of 10 ms.

Extractive summary: sentences quoted from the sources.