BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.
Key points
- Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models.
- Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications.
- To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient.
- We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology.
Sources (1)
- [1]BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for BioacousticsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:41 PM
In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.
Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Leaner Transformers Can Easily Learn to Cluster
- Oct 6, 2026Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
- Oct 6, 2026Towards In-Parameter Memory Augmentation for Large Language Models
- Oct 6, 2026Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
- Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
- Sep 29, 2026In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks