Inworld TTS generates natural, expressive speech for narration. Generate AI narrated videos with VideoGen's script to video workflow.
By Inworld AI · TTS 1.5 released May 2026 · Artificial Analysis Speech Arena Elo ~1,236 (July 2026)
Real prompts, real outputs, generated with Inworld TTS in VideoGen.
“Malcolm: Smooth, authoritative narrator with a deep, resonant timbre.”
Generated with Inworld TTS in VideoGen
“Celeste: Polished, refined female voice for high-end product showcases.”
Generated with Inworld TTS in VideoGen
“Rupert: Crisp, sophisticated British voice for corporate promos.”
Generated with Inworld TTS in VideoGen
“Deborah: Bright, articulate female voice for explainers and instructional videos.”
Generated with Inworld TTS in VideoGen
“Hades: Commanding, deep character voice for dramatic narration.”
Generated with Inworld TTS in VideoGen
“Hélène: Warm French female voice with a smooth, melodic cadence.”
Generated with Inworld TTS in VideoGen
“Asuka: Bright, energetic Japanese female voice for social media content.”
Generated with Inworld TTS in VideoGen
| Languages | 15, including English, Spanish, Japanese, and Arabic |
|---|---|
| Models | TTS 1.5 Max, Standard, and Mini |
| Voice cloning | Instant cloning from a short sample |
| Latency | Under 250ms to first audio (1.5 Max) |
| Output | Natural, expressive speech with emotional range |
The fastest and most affordable tier, built for drafts and quick iteration.
The balanced default: good quality voiceovers at an everyday cost.
Premium quality voiceovers for the content you publish.
MAX always integrates the state of the art. Inworld TTS is part of that set.
Every quality tier routes each request to the best model based on our internal evals, compliance requirements, and your intent. Requests stay fully compliant at every tier: sensitive data is never routed to models hosted by foreign adversaries such as China.
Sign up for VideoGen to open the AI playground and editor.
Paste or write your script and VideoGen generates a fully narrated video from it.
Choose a narrator from the voice catalog, or clone your own voice.
VideoGen builds the finished video with narration, visuals, music, and captions, ready to export.
Developers can generate voiceovers through the VideoGen API with the same quality tier routing as the app, so requests reach models like Inworld TTS without managing provider keys.
curl -X POST "https://api.videogen.io/v1/tools/text-to-speech" \
-H "Authorization: Bearer $VIDEOGEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{"ttsText": "Welcome to the future of video.", "voiceId": "<voice id>"}'
Turn any script into natural, expressive narration with the top-ranked text-to-speech model.
“VideoGen is awesome for scaling video production and improving turnaround time... The complex process of video editing, which could take days or months, now takes minutes!”


“VideoGen solves the biggest pain points of video production—complexity, cost, and time. In just a few clicks, anyone can create professional, copyright-free videos.”

“VideoGen is the most underrated tool for content creators who want to put out the highest quality content in the shortest amount of time.”

Every VideoGen workflow uses models like Inworld TTS automatically to build complete videos with visuals, narration, music, and captions.
Sign in to generate images, video clips, and more with the models you choose.