AION
Research paperSpeech & Audio1 source · Oct 7, 2026

BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics

In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.

Key points

  • Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models.
  • Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications.
  • To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient.
  • We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology.

Sources (1)

  • [1]BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:41 PM
    In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.
    Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Leaner Transformers Can Easily Learn to Cluster
  2. Oct 6, 2026Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
  3. Oct 6, 2026Towards In-Parameter Memory Augmentation for Large Language Models
  4. Oct 6, 2026Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
  5. Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
  6. Sep 29, 2026In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

Related