AI DevelopmentDecision Matrix7 min readPublished October 2, 2026

Voice models that can say mm-hm while you are still talking

Full-Duplex Voice Models Compared: Listen and Talk at Once

Tavus's Griffin-Lite joins OpenAI's GPT-Live-1 and three open models that listen while they speak. Latency, price, license and how to test each one yourself.

DA
Digital Applied Team
Research and practical guidance
CoverageOctober 2, 2026

Most voice assistants take turns: they wait for you to stop, think, then speak. A full-duplex model listens and speaks at the same time, so it can acknowledge you mid-sentence and stop the moment you interrupt. On October 1, 2026, Tavus previewed Griffin-Lite, a full-duplex model for video calls. Five models now make that claim in their own documentation, and they differ sharply in access, license and what their makers have measured.

Built from each maker’s own announcement, model card or repository, read October 3, 2026. Every latency and benchmark figure here is the maker’s own; none was reproduced by Digital Applied.

Key takeaways
  1. 01
    Five models qualifyOpenAI’s GPT-Live-1, Tavus’s Griffin-Lite and three open-weight models from NVIDIA and Kyutai.
  2. 02
    Only one is a paid APIGPT-Live-1 costs $0.05 a minute for the voice layer. Griffin-Lite is a closed preview.
  3. 03
    Latency claims do not line upEach maker measures something different, on different hardware. Test on your own calls.
  4. 04
    Open models are runnableVoiceChat, PersonaPlex and Moshi can run on your own GPUs, with licenses to check first.

01 — DefinitionWhat full-duplex means

The term comes from telephony: a full-duplex line carries sound both ways at once. A traditional voice agent, as OpenAI describes it, chains three systems: speech recognition, a language model, then speech synthesis. That chain has to decide when you have finished before it can start, which is how pauses get cut off and interruptions land late.

A full-duplex model handles both audio streams in one model. Kyutai’s Moshi repository describes the design plainly: one stream for the model’s own speech, one for the user’s, modelled together. OpenAI makes the business case in its GPT-Live-1 announcement: one model reasoning over incoming and outgoing audio avoids the delays and fragile handoffs of a chained system. It cites the language-learning company Speak, which found interruptions fell by almost 80% against earlier turn-based systems, a customer-reported figure.

02 — The newsWhat changed this week

Tavus announced Griffin on October 1 as a model that sees, hears and talks at once in a video call, and released Griffin-Lite as a research preview to a small group of testers. It decides what to do at sub-second intervals rather than once per turn, and generates the video face as well as the voice. Tavus reports that 26 of 54 people (48%) in its own study thought they had spoken to a real person after a one-minute call; that is Tavus’s study, with Tavus’s method.

On September 30, Inworld announced it had acquired Ultravox, a platform for building real-time voice agents. Inworld says the deal advances its work on speech-to-speech experiences and that more will follow in the coming weeks. Its announcement does not describe Ultravox as full-duplex, so Ultravox is not in the tables below.

03 — The modelsThe models that do it

Each row is a model whose maker states that it listens and speaks at the same time. Two are commercial; three publish their weights.

Sources: OpenAI announcement; Tavus research post; NVIDIA model card and repository; Kyutai repository. Read October 3, 2026.
ModelAccessPrice or licenseWhat it adds
GPT-Live-1 (OpenAI)Closed API; first introduced in ChatGPT$0.05 a minute for the voice layer, plus the backend modelDelegates reasoning and tool calls to a separate model; phone calls supported
Griffin-Lite (Tavus)Research preview for selected early testersNo public priceVideo as well as voice: sees the user and generates a talking face
NemotronLabs VoiceChat 11B (NVIDIA)Open weights on Hugging Face; NVIDIA container for live useOpenMDW 1.1 licenseTool calls during the conversation; English
PersonaPlex 7B (NVIDIA)Gated open weights; accept the license on Hugging FaceWeights under NVIDIA’s Open Model License; code MITSet the role with a text prompt and the voice with an audio prompt
Moshi (Kyutai)Open weights; code and setup on GitHubWeights CC-BY 4.0; code MIT and ApacheThe open base PersonaPlex is built on

The two NVIDIA models answer different needs. The VoiceChat model card says it is the first open full-duplex model to call tools, and lets a developer set a holding phrase the agent speaks while a tool runs. The PersonaPlex repository focuses on persona: a text prompt sets the role, such as a customer service agent, and an audio prompt sets the voice. GPT-Live-1 takes a third route and hands hard questions and tool calls to a separate text model, which our GPT-Live-1 API guide covers in detail.

04 — The latencyThe latency each maker reports

Latency is the number everyone quotes and nobody measures the same way. The table gives each figure with what it measures and where. Read across a row, never down a column.

Sources: as in the table above. All figures are reported by the model’s maker.
ModelReported figureWhat and where
GPT-Live-130 percentage points better on Full Duplex Bench than GPT-Realtime-2.1OpenAI’s own evaluation; no millisecond figure in the announcement text
Griffin-Lite0.43 s from audio arriving to its effect on screenTavus, measured on H100s, for the video generator
NemotronLabs VoiceChat448 ms turn-taking; 480 ms to yield when interruptedNVIDIA, Full-Duplex-Bench 1.0, on H100
Moshi160 ms theoretical; as low as 200 ms in practiceKyutai, on an L4 GPU
PersonaPlexNo figure in the public READMENVIDIA points to FullDuplexBench prompts for evaluation
Why these cannot be ranked

Moshi’s figure is its overall latency on an L4 GPU. NVIDIA’s is time to reply after a user finishes, on a data-center GPU. Tavus’s covers video, not speech. OpenAI gives a benchmark gain, not a time. Putting them in one chart would rank four different measurements. Our voice agent latency reference separates the delays worth timing.

05 — MethodHow to test one yourself

The public benchmark most of these makers cite, Full-Duplex-Bench, scores four behaviors, and they make a good test plan for your own calls. Record 20 real conversations from your use case and score each model on the same four.

1
Pause handling
Does it wait?

Pause mid-thought for two seconds. A good model waits; a bad one answers half a question.

Patience
2
Backchannels
Does it acknowledge?

Talk for 30 seconds. Listen for short acknowledgements that do not take the turn.

Presence
3
Turn-taking
How fast is the reply?

Time from the end of your sentence to its first word, on your network and hardware.

Speed
4
Interruption
Does it stop?

Cut in mid-answer. Time how long it keeps talking, and check it kept the context.

Control

Add one test the benchmark does not cover: background noise. OpenAI lists noise handling as a strength of GPT-Live-1, and it is where phone-line deployments fail first. Benchmark rankings and listener preference also disagree, as our review of voice benchmarks against listener votes found, so include a few real listeners in the scoring.

06 — Practical implicationsWhich one fits

Phone or app agent in production now
GPT-Live-1 with a backend model for tools
Paid API
Self-hosted agent that must call tools
NemotronLabs VoiceChat on your own GPUs
Open
A fixed persona and voice for a service line
PersonaPlex, after the license review
Open
Video calls with a face on screen
Ask Tavus for preview access; plan for change
Preview

Full duplex fixes the rhythm of a call, not its substance. The agent still needs the right tools, data and limits behind it, and a plan for when it is wrong. Our AI transformation work covers voice agents from model choice to call-flow testing.

07 — The claimsReading a maker’s latency claim

The latency table asks you to read across a row and never down a column, and the reason is worth spelling out, because a single number in a product page is designed to be read down the column. Every figure in that table answers a different question: what is being timed, from which event to which, on what hardware, in what kind of call. A number without those four answers is not comparable to anything, including the same maker’s number from a different page.

So the habit when a vendor quotes latency is to ask the four questions before writing the number down. Is it end-to-end or one stage of the chain? Does the clock start when the user stops speaking, when their last word is detected, or when audio arrives at the server? Is it a data-center GPU or a consumer one, a lab network or a phone line? Is it speech only, or speech plus a generated video face? The table’s right-hand column records what each maker discloses on those questions, and it is the column to read first.

The only figure that resolves all four questions for your product is the one you measure in the test plan above, on your calls, your network and your hardware. A maker’s number tells you the model can be fast somewhere. Yours tells you whether it is fast where you need it.

Keep the maker’s figure in the record anyway, with its four answers beside it. When your own measurement comes out very different, the gap is informative: a model that is fast on a data-center GPU and slow on yours is telling you about your hardware, and a model whose reply time is good but whose interruption handling is poor is telling you that reply time was the wrong thing to quote.

08 — The failuresFour ways a voice pilot misleads

The illustrative cases below each produce a confident pilot result that does not survive the first week of real calls. They are what happens when the test plan is run with a step missing.

The ranked latency: two models are compared on their makers’ quoted figures, one measured overall on an L4 and the other as reply time on a data-center GPU, and the comparison picks a winner on numbers that were never measuring the same thing. The quiet room: every test call is made from a desk with a good headset, the model handles it beautifully, and the first customer on a speakerphone in a car exposes the noise handling nobody tried. The polite tester: the people running the pilot wait for the model to finish, as people do with software, so interruption and barge-in are never exercised and the first impatient caller finds out what happens. The benchmark proxy: the model with the better published Full-Duplex-Bench score is chosen without listening, and real listeners prefer the other one, the kind of disagreement our review of benchmarks against listener votes found.

The test plan catches all four if it is run as written: the four benchmark behaviors plus background noise, on 20 recorded conversations from your own use case, scored by a few real listeners as well as by the team.

The reason to record the 20 calls first

A pilot scored on calls recorded before the model was chosen is much harder to tilt toward one model. The same 20 conversations run through two candidates, with the same listeners rating both, is the comparison in this post that produces a number you can rank.

Next step

Score two models on 20 of your own calls

Pick one paid and one open model from the tables, run the four tests plus background noise on 20 recorded conversations, and let a few real listeners rate the results. That answers more than any number a maker has published.

Voice AI implementation

Voice agents that know when to stop talking

Digital Applied tests voice models on your real calls, builds the agent behind them and measures what callers actually experience.

Model testingCall-flow designLatency measurement
The test plan

Five checks per model

  • →Pause handling
  • →Backchannels
  • →Turn-taking speed
  • →Interruption
  • →Background noise
Questions and answers

Practical questions

A model that processes your speech and its own speech at the same time, in one model, instead of waiting for you to finish before replying. It can acknowledge you while you talk and stop when you interrupt.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Voice AI Benchmarks and Listener Votes Disagree: 14 Models

Artificial Analysis scores realtime voice systems two ways. The arena system with the best reasoning and turn-taking scores ranks last on listener preference.

September 25, 2026 · 6 minRead
AI Development

Anthropic: Open GLM-5.3 Nearly Matches Mythos at Exploits

In Anthropic's tests the open GLM-5.3 model succeeded at 50 of 410 exploit attempts, close to Claude Mythos Preview at 56. What the report shows and omits.

September 29, 2026 · 7 minRead
AI Development

Eleven v4 and v4 Turbo: What Changed for Voice Agents

ElevenLabs' Eleven v4 tops Artificial Analysis' voice arena and v4 Turbo claims ~100ms. The two latency figures, the vendor tests and the v4 price.

September 28, 2026 · 6 minRead
AI Development

Best Text-to-Speech Models, September 2026: Ranked, Priced

92 text-to-speech models ranked by blind listener vote, with the price per million characters, open-weight licences and the job each leader fits.

September 24, 2026 · 9 minRead
AI Development

State of AI Agents 2026: 200+ Data Points Compiled

The definitive State of AI Agents 2026 — 247 data points across adoption, ROI, autonomy, and governance, sourced from McKinsey, Stanford HAI, and Gartner.

May 22, 2026 · 16 minRead
AI Development

AI Video Generation 2026: Omni vs Sora vs Veo 3 Compared

Gemini Omni, OpenAI Sora 2, and Google Veo 3.1 compared for video — quality, per-second cost spread of 17x, and the September 24 Sora API sunset clock.

May 22, 2026 · 15 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source