Model-initiated spoken interruptions during user speech based on context rather than turn boundaries.
Verbal interjections
Verbal interjections are model-initiated spoken contributions that occur during user speech, triggered by real-time assessment of the ongoing context rather than a completed turn. In a conventional half-duplex system, the model waits for the user to finish speaking before generating any response. A verbal interjection system allows the model to jump in when the accumulated context warrants it. If the user says something factually incorrect, the model might interject with a correction. If the user asks a question and then immediately answers it themselves, the model might stay quiet. If the user seems confused, the model might offer clarification unprompted.
The capability depends on the same architectural foundation as full-duplex interaction and micro-turn processing. The model must perceive and reason about the ongoing input stream while simultaneously generating output, without being blocked by either. Turn-based architectures prohibit this: the model cannot generate output until the entire input turn has been received and processed. Verbal interjections emerge from the time-aligned micro-turn design and do not require a separate interjection-detection module.
Several practical applications matter. A model that can interject on incorrect mispronunciations in real time (the CueSpeak benchmark) differs from one that can only correct errors after the fact. A model that can say "hold on, I see what you mean" in response to a visual cue (visual interjections) while the user is still speaking enables a collaborative style that mirrors how humans naturally work together. Live translation, generating speech in language B while the user is still speaking in language A, is another core use case that requires simultaneous speech and the ability to interject proactively rather than reactively.
The risks of verbal interjection are significant. A model that interjects incorrectly or excessively would be experienced as rude, distracting, or unhelpful. Calibrating interjection behavior, deciding when the model's confidence that an interjection is warranted is high enough to override the user's current turn, is an open problem. The model's decisions here have real social consequences and represent a safety-critical dimension of interaction model behavior that does not exist in turn-based systems.
Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Verbal Interjections already loaded, so edit it or scan as is.