MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface
We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder.
Key points
- System One models output constrained decisions and probability distributions rather than free-form text generation.
- While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface.
- To scale effectively across both small closed-set (< 256) and massive open-set (millions) candidate spaces, MetaEncoder employs a bi-encoder architecture trained via unidirectional contrastive learning for request-candidate alignment.
- We conduct extensive evaluations across 11 benchmark suites and 190 tasks spanning multimodal decision-making, understanding (closed-set) and retrieval (open-set), highlighting where MetaEncoder beats SOTA multimodal encoders, as well as its current limits on reasoning-intensive tasks.
Sources (1)
- [1]MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language InterfacearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:22 AM
We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder.
System One models output constrained decisions and probability distributions rather than free-form text generation.
Extractive summary: sentences quoted from the sources.