Internal benchmark testing simultaneous speech responses triggered by specific user cues.
CueSpeak is an internal benchmark evaluating whether interaction models can produce simultaneous speech responses, meaning they speak at the same time as the user with an appropriate, semantically correct response to a specific conversational cue. The benchmark presents scenarios like: "every time I code-switch and use another language, give me the correct word in the original language." A full score requires both that the model's response has the correct semantic content (the right word in the right language) and that it is delivered at the same time as the user's code-switch, rather than after the user has finished their sentence.
The core capability being tested is cue-triggered simultaneous speech: the model must recognize that a specific event has occurred in the user's speech and produce a response at the same moment, not after. CueSpeak differs from TimeSpeak, which tests time-triggered initiation, and from standard verbal interjections, which test context-triggered corrections. CueSpeak specifically tests the intersection of semantic recognition and temporal concurrency. The question is whether the model can detect a meaningful event in the ongoing speech stream and respond concurrently.
The evaluation uses an LLM judge that assesses semantic correctness of the response and a separate timing analysis that determines whether the response overlapped with the relevant segment of user speech. Responses that are semantically correct but delivered after the user finishes speaking receive partial or no credit. Both criteria must be met simultaneously, so models cannot compensate for slow timing with higher accuracy.
Current frontier models perform near-zero on CueSpeak, consistent with the finding that simultaneous speech capability is a genuine frontier. Thinking Machines Lab designed the benchmark to quantify capabilities that existing benchmarks do not cover, specifically the ability to engage in real-time collaborative correction, translation, and co-reading that characterizes expert human collaboration. Practical applications include real-time language tutoring, accessibility tools for users who code-switch, and collaborative writing or coding assistance where the model provides concurrent feedback.
Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with CueSpeak Benchmark already loaded, so edit it or scan as is.