AION
Research paperLarge Language Models1 source · Oct 8, 2026

SteerablePlex: Can We Steer Full-Duplex Models?

We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order.

Key points

  • Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows.
  • The benchmark reveals that current open-source full-duplex models struggle to follow such constraints.
  • We then introduce a Group Reward-Decoupled Normalization Policy Optimization (GDPO)-based training recipe that enables a full-duplex model to follow textual instructions during an ongoing conversation while maintaining its turn-taking ability.
  • By connecting the resulting SteerablePlex to an asynchronous backend language model that monitors the conversation and provides instructions when needed, we build a more controllable full-duplex user simulator that follows multi-stage constraints more reliably than existing open-source models and GPT-Realtime.

Sources (1)

  • [1]SteerablePlex: Can We Steer Full-Duplex Models?
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:56 PM
    We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order.
    Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows.

Extractive summary: sentences quoted from the sources.