Skip to main content
Back to News
news/AI Applications

Synthesia's Interactive AI Avatars Turn Scripts to Conversations

Synthesia's interactive AI avatars now listen and answer back, not just read scripts, as the $4B startup pushes enterprise video toward agentic conversation.

Stefan Trbojevic

Stefan Trbojevic

26 September 20262 min read
LinkedIn
Abstract glowing audio waveforms and data streams around a central routing hub

The takeaway

Synthesia's interactive avatar runs a four-model pipeline where the language model is a replaceable component, pointing toward avatars as a thin orchestration layer over interchangeable AI models.

Why it matters for builders

Avatars are becoming an orchestration layer over interchangeable reasoning, voice, and video models. Synthesia treats the language model as swappable, mirroring how agentic automation assembles best-of-breed models rather than locking customers into one vendor.

Synthesia's Interactive AI Avatars Turn Scripts to Conversations

Synthesia, the British AI avatar startup valued at $4 billion, has pushed its technology beyond one-way videos. A hands-on demo now shows an interactive digital twin that listens, understands, and answers back.

What happened

In a piece for TechCrunch, journalist Dominic-Madori Davis became the first person outside the company to receive a personal interactive avatar from Synthesia. The company built it from a short studio session: a set of photos plus a two-minute voice recording. It was then trained on a single article so it could field questions about that specific story.

The avatar runs on a four-model pipeline. A voice-to-text model transcribes what someone says, an agentic language model makes sense of it and can take actions, a text-to-voice model turns the response into audio, and Synthesia's own video model animates the avatar as it speaks.

Synthesia builds its video and voice models in-house but lets customers swap in alternatives from Cartesia, ElevenLabs, Google, or OpenAI, and host avatars on the cloud of their choice. The company, which says it crossed $100 million in annual recurring revenue last year, now splits its product into three tiers: a classic video-creation platform, an agentic Sessions product for interactive roleplay, and an API platform for combining its models with other services.

Why it matters

The demo signals a wider shift in the enterprise avatar market, where rivals such as D-ID, HeyGen, and Colossyan are also moving from scripted, one-way videos toward avatars that hold conversations. Synthesia's Roleplay Sessions, launched in July, already lets employees rehearse sales pitches and performance reviews against an avatar that pushes back and scores their answers.

Synthesia interactive avatar model pipeline

For AI builders, the architecture is the real story. Avatars are becoming a thin orchestration layer over interchangeable reasoning, voice, and video models. Synthesia treats the language model as a replaceable component, echoing how agentic automation is assembling best-of-breed models rather than locking customers into one vendor.

That points toward a near future where interactive avatars become a standard interface for training, onboarding, and customer-facing automation, powered by the same agentic model plumbing that runs the rest of the AI stack.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

26 September 2026

Updated

26 September 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.