For voice agents, Gemma 4 31B claims lower latency than GPT-4.1.
Google's Gemma 2 9B model delivers its first audio in 354 ms on LiveKit. This speed reduces voice latency by a factor of five compared to GPT-4o.
Gemma 4 31B, Google's open-weights model, joins LiveKit Inference as the recommended default LLM for voice agents. The offering is integrated into LiveKit Cloud, and the model runs on the platform's GPU infrastructure, built for low latency.
The argument lies in the response times. According to LiveKit, the model starts its generation in about 192 ms (time to first token) and produces its first sound in 354 ms (time to first audio). Compared to GPT-4o, the provider claims about five times lower latency and nearly six times lower cost, at a comparable response quality, all measured on its own reference agent, a hotel receptionist. Google also puts forward a score of 76.9% on tau2bench in agentic tool use, above GPT-4o.
The service highlights an optimization choice: where other inference platforms prioritize throughput and tolerate more latency, LiveKit says it does the opposite, even if it means reducing throughput, on the grounds that voice cannot wait. On its benchmark, the co-located GPU route delivers 158 tokens per second and starts a sentence much faster than a generalist router like OpenRouter, measured at 33 tokens per second and 1,876 ms.
All of this remains quantified by the provider, based on a single test agent. LiveKit Inference operates with zero data retention by default and opens access to other recognition, generation, and speech synthesis models, from Deepgram to Cartesia. Connecting the model takes just one line in a LiveKit Agents session.