Benchmark suite measuring AI model behavior across interruption, backchannel, and overlapping speech.
FD-Bench (Full-Duplex Bench) is a benchmark suite designed to measure the quality of an AI model's interactive behavior in full-duplex settings — specifically, how well the model handles interruption, backchanneling, overlapping speech, and responsiveness in multi-party or continuous-speech scenarios. Unlike standard benchmarks that evaluate a model's response to a complete input turn, FD-Bench evaluates model behavior at specific moments during ongoing audio and video streams, testing whether the model responds correctly at the right time rather than only when explicitly prompted.
The benchmark exists in multiple versions. FD-Bench V1 focuses on turn-taking latency — measuring the delay between a natural turn transition point in an audio recording and the model's response. FD-Bench V1.5 tests average performance across several interaction scenarios: user interruption (can the model handle being cut off mid-sentence?), user backchannel (does the model appropriately acknowledge without dominating?), talking to others (multi-party interaction), and background speech (can the model stay on task when other people are speaking?). FD-Bench V3 adds response quality metrics and tests with tools enabled.
FD-Bench V3 introduced a Pass@1 metric that measures whether the model's first response at a given interaction moment is correct — a stringent criterion that rewards both timely and accurate responses. Interaction models score substantially higher than turn-based systems on FD-Bench, but still well below human performance on naturalness metrics, suggesting that the interactive gap between AI and human conversation remains significant even when AI responses are technically correct.
The FD-Bench suite is notable because it was specifically designed for the evaluation gap in interactive AI — most existing benchmarks measure intelligence or instruction-following but not the temporal and conversational dynamics of interactive behavior. Its development reflects the growing recognition that interaction quality cannot be inferred from turn-based benchmark performance, and that metrics like latency and appropriateness of interruption are qualitatively different from accuracy metrics.
Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with FD-Bench already loaded, so edit it or scan as is.