Skip to main content
Back to News
news/AI Applications

ElevenLabs v4 Pushes Voice Agents Toward Real Conversation

ElevenLabs v4 adds faster streaming, richer expression controls, and 90 languages, raising the bar for voice agents in customer-facing workflows.

Stefan Trbojevic

Stefan Trbojevic

28 September 20262 min read
LinkedIn
Abstract audio waveform and routing nodes for a voice AI model launch

The takeaway

For voice-agent builders, latency and conversational control are becoming core infrastructure concerns, not cosmetic features.

Why it matters for builders

Voice agents now need streaming, expressive turn-taking, and multilingual consistency to work inside real customer workflows.

ElevenLabs v4 Pushes Voice Agents Toward Real Conversation

ElevenLabs is raising the ceiling for voice agents. The company launched ElevenLabs v4 and v4 Turbo on September 28, adding richer expression controls, lower latency, and support for more than 90 languages. The release targets a problem that has become increasingly visible in production: a voice agent can be technically correct and still feel painfully slow, flat, or unable to handle a difficult conversation.

What changed

According to TechCrunch, the new models can begin generating audio while the underlying language model is still producing its answer. That streaming path should reduce dead air in calls and make interruptions feel more natural. ElevenLabs also says v4 can clone a voice from 10 seconds of audio, preserve identity over longer passages, and follow stacked inline expression controls. The company has expanded language support from 70 to more than 90, with notable quality gains in Japanese, Brazilian Portuguese, Mandarin, and Cantonese.

Abstract routing topology for low-latency voice agents

Why it matters for builders

The practical shift is not just better narration. Voice agents increasingly need to manage holds, escalations, objections, and handoffs without sounding like they have fallen out of the workflow. Lower latency gives orchestration layers more room to stream partial responses, while richer prosody can signal whether an agent is clarifying, apologizing, or escalating. For teams connecting telephony, CRM, and automation tools, that means voice quality becomes part of system design rather than a final presentation layer.

ElevenLabs says more than 55% of its business now comes from large companies, while competition is intensifying from Cartesia, Deepgram, Google, OpenAI, and other speech startups. The next differentiator will be operational: consistent identity, safe cloning, reliable turn-taking, and measurable outcomes across real calls.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

28 September 2026

Updated

28 September 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.