Narration & Subtitles: Adding Voice to AI Videos
Talking-head mode ships with built-in TTS voice and lip sync; plain text-to-video needs you to add narration and captions yourself.
A common question when starting out: do AI-generated videos come with a voice-over? Do I need to add subtitles myself? The answer depends on which generation mode you use. The talking-head presenter mode generates the voice-over and lip sync internally. Regular text-to-video, however, produces only visuals β no narration, no captions. This guide explains both cases and where to add subtitles.
Talking-head mode: one reference image is enough for a narrated presenter video
Talking-head mode needs only a single presenter reference image to produce a video of one person speaking. The narration is generated automatically by the built-in TTS, and lip sync is handled inside the model β no extra voice recording or manual alignment required. Note that multi-character scenes and advanced motion libraries live on external platforms such as HeyGen, Synthesia or D-ID, and are not part of the presenter mode here.
Plain text-to-video: no narration by default
If you simply turn a text prompt into a video and no one is speaking in the frame, the result comes with no auto-generated voice and no captions. You get clean visuals only, and you'll handle the narration and subtitles yourself in your editing tool. Also note the generation limits: up to 20 seconds per clip, 720p resolution, and a rate limit of 16 requests per minute.
How do I add subtitles?
Plain text-to-video clips have no captions: burn them in after export using an editor such as CapCut, Premiere or similar. If you want captions built in, switch to talking-head mode instead β videos from that mode come with subtitles synced to the narration.
Give it a try in the demo
Both the online Demo and Studio offer a talking-head entry point, and /api-docs exposes the full API. Generate a narrated presenter video from a single image right now.
Try the DemoVoiceover & Subtitles References
The TTS voiceover, lip-sync, and subtitle-burning approaches on this page reference the following industry products and open-source tech.
- HeyGen β ζ°εδΊΊθ§ι’εδ½εΉ³ε°HeyGen Β· heygen.com (2025)
Digital-human platforms like HeyGen mentioned here are described per their official product docs.
- ElevenLabs β AI θ―ι³δΈ TTSElevenLabs Β· elevenlabs.io (2025)
For DIY voiceovers, ElevenLabs is a mainstream AI TTS option; its voices and language support follow its official site.
- Whisper β θͺε¨θ―ι³θ―ε«δΈεεΉθ½¬ε½OpenAI Β· GitHub (openai/whisper) (2023)
For auto subtitle transcription, OpenAI's open-source Whisper model is a reference implementation.
- CapCut β θ§ι’εͺθΎδΈεεΉε·₯ε
·ByteDance Β· capcut.com (2025)
The subtitle-burning features of editors recommended here (Jianying/CapCut, etc.) follow their official docs.
Why you can trust this
Written by a practicing developer and cross-checked against authoritative primary sources such as arXiv papers, vendor technical reports, and the Stanford AI Index.
SandGrid@lcy362
Author of Agnes Video Generator Β· Full-stack Developer
Independent developer and author of Agnes Video Generator, an open-source (MIT) AI video generation tool built on Agnes AI's free video models. Has helped 1,000+ creators produce AI videos at zero cost. Focused on making video generation models accessible and production-ready; content is based on hands-on practice and primary research.
View on GitHubLast updatedοΌ2026-08-21
Ready to Start Creating?
"Making world-class AI belong to everyone." β Bruce Yang. It's completely free, no credit card, and you won't need a high-end GPU. Your first AI video starts at zero cost. Want to use Agnes AI's free video models? This is the easiest way in.
Clone the GitHub repo and launch in 2 minutes