The takeaway
Synthesia's interactive avatar runs a four-model pipeline where the language model is a replaceable component, pointing toward avatars as a thin orchestration layer over interchangeable AI models.
Why it matters for builders
Avatars are becoming an orchestration layer over interchangeable reasoning, voice, and video models. Synthesia treats the language model as swappable, mirroring how agentic automation assembles best-of-breed models rather than locking customers into one vendor.
Synthesia's Interactive AI Avatars Turn Scripts to Conversations
Synthesia, the British AI avatar startup valued at $4 billion, has pushed its technology beyond one-way videos. A hands-on demo now shows an interactive digital twin that listens, understands, and answers back.
What happened
In a piece for TechCrunch, journalist Dominic-Madori Davis became the first person outside the company to receive a personal interactive avatar from Synthesia. The company built it from a short studio session: a set of photos plus a two-minute voice recording. It was then trained on a single article so it could field questions about that specific story.
The avatar runs on a four-model pipeline. A voice-to-text model transcribes what someone says, an agentic language model makes sense of it and can take actions, a text-to-voice model turns the response into audio, and Synthesia's own video model animates the avatar as it speaks.
Synthesia builds its video and voice models in-house but lets customers swap in alternatives from Cartesia, ElevenLabs, Google, or OpenAI, and host avatars on the cloud of their choice. The company, which says it crossed $100 million in annual recurring revenue last year, now splits its product into three tiers: a classic video-creation platform, an agentic Sessions product for interactive roleplay, and an API platform for combining its models with other services.
Why it matters
The demo signals a wider shift in the enterprise avatar market, where rivals such as D-ID, HeyGen, and Colossyan are also moving from scripted, one-way videos toward avatars that hold conversations. Synthesia's Roleplay Sessions, launched in July, already lets employees rehearse sales pitches and performance reviews against an avatar that pushes back and scores their answers.
![]()
For AI builders, the architecture is the real story. Avatars are becoming a thin orchestration layer over interchangeable reasoning, voice, and video models. Synthesia treats the language model as a replaceable component, echoing how agentic automation is assembling best-of-breed models rather than locking customers into one vendor.
That points toward a near future where interactive avatars become a standard interface for training, onboarding, and customer-facing automation, powered by the same agentic model plumbing that runs the rest of the AI stack.
The Automation Brief
Read 5 AI stories instead of 50.
The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.
No noise. Unsubscribe anytime.
Editorial notes
Stefan Trbojevic
n8n Lab Editorial
26 September 2026
26 September 2026
AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.



