LAUNCH

Google’s Gemma 4 31B hits 354ms audio latency on LiveKit

Signals Inbox·July 16, 2026·Voice AI

Google just put Gemma 4 31B on LiveKit with 354ms time to first audio and 192ms time to first token. The real tension is that a fairly large open model is getting close to human conversational timing while Google says it also beats GPT-4.1 on a tool-use benchmark.

The Signal, Explained in 3 Minutes

Q1What actually launched?

Google officially announced Gemma 4 31B on LiveKit Inference, a hosted version tuned for real-time voice agents. Google reports 192ms to the first text token, 354ms to the first audio, and a 76.9% score on tau2-bench.

Q2Why does 354ms matter?

Because voice agents feel broken when every reply arrives after an awkward pause. Human turn-taking often happens within a few hundred milliseconds, while many voice pipelines stack speech recognition, model reasoning, text-to-speech, and network delay. A 354ms first-audio result puts the model layer near conversational territory, though it is not the same as total end-to-end call latency.

Q3Is it faster than the market?

It is competitive, not untouchable. Specialized voice systems such as Deepgram have reported roughly 200 to 300ms response times. The difference is that Gemma 4 31B is a much broader language model built to reason and call tools, not only process speech. Google is trying to combine speed with deeper agent behavior.

Q4What does beating GPT-4.1 mean?

Google says Gemma 4 31B reaches 76.9% on tau2-bench, ahead of GPT-4.1 in its setup. Tau2-bench tests conversational agents that must guide a user and use tools inside a shared telecom task. That makes the result more relevant to support calls than a normal chatbot benchmark, but it is still one benchmark and the claim needs independent replication.

Q5Why is the 31B size important?

A 31B-parameter model is large enough to handle serious reasoning, but small enough to serve more cheaply than the biggest frontier models. If LiveKit can keep it fast under real traffic, developers may get capable voice agents without paying frontier-model prices for every second of every call.

Q6What is the real shift?

Voice AI is moving from scripted phone bots toward agents that can listen, reason, use software, and answer before the silence feels awkward. The key test now is production performance: total response delay, interruptions, accuracy under noise, tool failures, and cost across thousands of simultaneous calls.

← Back to the signals