Most voice assistants take turns: they wait for you to stop, think, then speak. A full-duplex model listens and speaks at the same time, so it can acknowledge you mid-sentence and stop the moment you interrupt. On October 1, 2026, Tavus previewed Griffin-Lite, a full-duplex model for video calls. Five models now make that claim in their own documentation, and they differ sharply in access, license and what their makers have measured.
Built from each maker’s own announcement, model card or repository, read October 3, 2026. Every latency and benchmark figure here is the maker’s own; none was reproduced by Digital Applied.
- 01Five models qualifyOpenAI’s GPT-Live-1, Tavus’s Griffin-Lite and three open-weight models from NVIDIA and Kyutai.
- 02Only one is a paid APIGPT-Live-1 costs $0.05 a minute for the voice layer. Griffin-Lite is a closed preview.
- 03Latency claims do not line upEach maker measures something different, on different hardware. Test on your own calls.
- 04Open models are runnableVoiceChat, PersonaPlex and Moshi can run on your own GPUs, with licenses to check first.
01 — DefinitionWhat full-duplex means
The term comes from telephony: a full-duplex line carries sound both ways at once. A traditional voice agent, as OpenAI describes it, chains three systems: speech recognition, a language model, then speech synthesis. That chain has to decide when you have finished before it can start, which is how pauses get cut off and interruptions land late.
A full-duplex model handles both audio streams in one model. Kyutai’s Moshi repository describes the design plainly: one stream for the model’s own speech, one for the user’s, modelled together. OpenAI makes the business case in its GPT-Live-1 announcement: one model reasoning over incoming and outgoing audio avoids the delays and fragile handoffs of a chained system. It cites the language-learning company Speak, which found interruptions fell by almost 80% against earlier turn-based systems, a customer-reported figure.
02 — The newsWhat changed this week
Tavus announced Griffin on October 1 as a model that sees, hears and talks at once in a video call, and released Griffin-Lite as a research preview to a small group of testers. It decides what to do at sub-second intervals rather than once per turn, and generates the video face as well as the voice. Tavus reports that 26 of 54 people (48%) in its own study thought they had spoken to a real person after a one-minute call; that is Tavus’s study, with Tavus’s method.
On September 30, Inworld announced it had acquired Ultravox, a platform for building real-time voice agents. Inworld says the deal advances its work on speech-to-speech experiences and that more will follow in the coming weeks. Its announcement does not describe Ultravox as full-duplex, so Ultravox is not in the tables below.
03 — The modelsThe models that do it
Each row is a model whose maker states that it listens and speaks at the same time. Two are commercial; three publish their weights.
| Model | Access | Price or license | What it adds |
|---|---|---|---|
| GPT-Live-1 (OpenAI) | Closed API; first introduced in ChatGPT | $0.05 a minute for the voice layer, plus the backend model | Delegates reasoning and tool calls to a separate model; phone calls supported |
| Griffin-Lite (Tavus) | Research preview for selected early testers | No public price | Video as well as voice: sees the user and generates a talking face |
| NemotronLabs VoiceChat 11B (NVIDIA) | Open weights on Hugging Face; NVIDIA container for live use | OpenMDW 1.1 license | Tool calls during the conversation; English |
| PersonaPlex 7B (NVIDIA) | Gated open weights; accept the license on Hugging Face | Weights under NVIDIA’s Open Model License; code MIT | Set the role with a text prompt and the voice with an audio prompt |
| Moshi (Kyutai) | Open weights; code and setup on GitHub | Weights CC-BY 4.0; code MIT and Apache | The open base PersonaPlex is built on |
The two NVIDIA models answer different needs. The VoiceChat model card says it is the first open full-duplex model to call tools, and lets a developer set a holding phrase the agent speaks while a tool runs. The PersonaPlex repository focuses on persona: a text prompt sets the role, such as a customer service agent, and an audio prompt sets the voice. GPT-Live-1 takes a third route and hands hard questions and tool calls to a separate text model, which our GPT-Live-1 API guide covers in detail.
04 — The latencyThe latency each maker reports
Latency is the number everyone quotes and nobody measures the same way. The table gives each figure with what it measures and where. Read across a row, never down a column.
| Model | Reported figure | What and where |
|---|---|---|
| GPT-Live-1 | 30 percentage points better on Full Duplex Bench than GPT-Realtime-2.1 | OpenAI’s own evaluation; no millisecond figure in the announcement text |
| Griffin-Lite | 0.43 s from audio arriving to its effect on screen | Tavus, measured on H100s, for the video generator |
| NemotronLabs VoiceChat | 448 ms turn-taking; 480 ms to yield when interrupted | NVIDIA, Full-Duplex-Bench 1.0, on H100 |
| Moshi | 160 ms theoretical; as low as 200 ms in practice | Kyutai, on an L4 GPU |
| PersonaPlex | No figure in the public README | NVIDIA points to FullDuplexBench prompts for evaluation |
Moshi’s figure is its overall latency on an L4 GPU. NVIDIA’s is time to reply after a user finishes, on a data-center GPU. Tavus’s covers video, not speech. OpenAI gives a benchmark gain, not a time. Putting them in one chart would rank four different measurements. Our voice agent latency reference separates the delays worth timing.
05 — MethodHow to test one yourself
The public benchmark most of these makers cite, Full-Duplex-Bench, scores four behaviors, and they make a good test plan for your own calls. Record 20 real conversations from your use case and score each model on the same four.
Pause handling
Pause mid-thought for two seconds. A good model waits; a bad one answers half a question.
Backchannels
Talk for 30 seconds. Listen for short acknowledgements that do not take the turn.
Turn-taking
Time from the end of your sentence to its first word, on your network and hardware.
Interruption
Cut in mid-answer. Time how long it keeps talking, and check it kept the context.
Add one test the benchmark does not cover: background noise. OpenAI lists noise handling as a strength of GPT-Live-1, and it is where phone-line deployments fail first. Benchmark rankings and listener preference also disagree, as our review of voice benchmarks against listener votes found, so include a few real listeners in the scoring.
06 — Practical implicationsWhich one fits
Full duplex fixes the rhythm of a call, not its substance. The agent still needs the right tools, data and limits behind it, and a plan for when it is wrong. Our AI transformation work covers voice agents from model choice to call-flow testing.
07 — The claimsReading a maker’s latency claim
The latency table asks you to read across a row and never down a column, and the reason is worth spelling out, because a single number in a product page is designed to be read down the column. Every figure in that table answers a different question: what is being timed, from which event to which, on what hardware, in what kind of call. A number without those four answers is not comparable to anything, including the same maker’s number from a different page.
So the habit when a vendor quotes latency is to ask the four questions before writing the number down. Is it end-to-end or one stage of the chain? Does the clock start when the user stops speaking, when their last word is detected, or when audio arrives at the server? Is it a data-center GPU or a consumer one, a lab network or a phone line? Is it speech only, or speech plus a generated video face? The table’s right-hand column records what each maker discloses on those questions, and it is the column to read first.
The only figure that resolves all four questions for your product is the one you measure in the test plan above, on your calls, your network and your hardware. A maker’s number tells you the model can be fast somewhere. Yours tells you whether it is fast where you need it.
Keep the maker’s figure in the record anyway, with its four answers beside it. When your own measurement comes out very different, the gap is informative: a model that is fast on a data-center GPU and slow on yours is telling you about your hardware, and a model whose reply time is good but whose interruption handling is poor is telling you that reply time was the wrong thing to quote.
08 — The failuresFour ways a voice pilot misleads
The illustrative cases below each produce a confident pilot result that does not survive the first week of real calls. They are what happens when the test plan is run with a step missing.
The ranked latency: two models are compared on their makers’ quoted figures, one measured overall on an L4 and the other as reply time on a data-center GPU, and the comparison picks a winner on numbers that were never measuring the same thing. The quiet room: every test call is made from a desk with a good headset, the model handles it beautifully, and the first customer on a speakerphone in a car exposes the noise handling nobody tried. The polite tester: the people running the pilot wait for the model to finish, as people do with software, so interruption and barge-in are never exercised and the first impatient caller finds out what happens. The benchmark proxy: the model with the better published Full-Duplex-Bench score is chosen without listening, and real listeners prefer the other one, the kind of disagreement our review of benchmarks against listener votes found.
The test plan catches all four if it is run as written: the four benchmark behaviors plus background noise, on 20 recorded conversations from your own use case, scored by a few real listeners as well as by the team.
A pilot scored on calls recorded before the model was chosen is much harder to tilt toward one model. The same 20 conversations run through two candidates, with the same listeners rating both, is the comparison in this post that produces a number you can rank.
Score two models on 20 of your own calls
Pick one paid and one open model from the tables, run the four tests plus background noise on 20 recorded conversations, and let a few real listeners rate the results. That answers more than any number a maker has published.