Text to Speech Avatar: Turn a Script into a Talking Photo
A text to speech avatar turns written words into speech and animates a presenter to deliver them. In the DIY AI avatar generator, start with one authorised portrait and a short script. This tutorial covers the text-input workflow, including checks for pronunciation, timing, cost, and the final animation.
The workflow: prepare your portrait and script, select Text & voice, approve the narration, add the portrait, review the avatar quote and create the clip. Keep the narration within the 30-second limit. Finalise what the viewer will hear before paying to animate it.
A talking photo needs a face and a voice, approved separately
The portrait supplies the presenter’s appearance; text-to-speech supplies the narration. The animation then makes the person appear to speak. A photograph alone cannot reproduce its subject’s voice: Microsoft’s text to speech avatar overview describes avatar appearance and voice as separate choices.
A preset voice is enough for this walkthrough. A recognisable personal voice requires a separate, authorised process, such as cloning your own voice. Do not assume that uploading a headshot performs both jobs.
1. Prepare a portrait with an unobstructed mouth
Choose a clear, front-facing photograph of one person, with the whole face visible and some space around the head. Avoid group photographs, extreme angles, heavy shadows and objects covering the mouth. A simple head-and-shoulders composition gives you fewer distracting details to inspect later.
Check the original lips, teeth and jaw before uploading. If they already look distorted, replace the source rather than expecting animation to repair it. Keep a copy of the original for comparison with the finished video.
Use your own portrait or ask the subject to approve both the animation and the words they will appear to say. Make the clip’s synthetic nature clear wherever viewers might mistake it for a real recording.
2. Write spoken words, not instructions for the presenter
For a first attempt, aim for roughly 35-45 words. That is a drafting recommendation, not a guaranteed duration: the generated narration determines whether the script fits. Leave room for pauses rather than treating 30 seconds as a target to fill.
Paste only the words the audience should hear. Remove headings, speaker labels and directions such as “smile at the camera”. Do not assume the narration field understands performance instructions or speech markup.
Write numbers according to their meaning. For example, “version three point two” gives the voice a clearer instruction than “v3.2”. If the script runs long, remove a sentence before trying to squeeze it into faster speech.
3. Select Text & voice and listen to your actual script
Open the avatar workspace and sign in. Select Text & voice, enter your script and choose an available voice suitable for its language and delivery. Review any speech-generation charge before confirming the narration.
A voice sample helps you choose a sound; it doesn’t prove your names, figures, or sentences will be spoken correctly. Generate the narration and listen from beginning to end. Check the actual duration, the pronunciation of unfamiliar words, and whether the final phrase ends comfortably.
Fix mistakes now. Use the detailed guide to fix AI voice pronunciation for names, acronyms and awkward numbers. Test a correction in the complete sentence, because a word that works alone may sound different in context.
Once the narration is correct, keep that version unchanged for the avatar render. “Approved” here means you have listened and accepted it, not merely that the tool successfully generated an audio file.
Worked example: a short introduction that discloses the animation
Use one authorised portrait and an available English voice for this 42-word script. It is a reusable example, not a completed test result; check the generated audio’s duration before creating the video.
Welcome to this short demonstration. A talking photo starts with a portrait and a written script. Listen to the voice first, then check the finished video for clear speech, stable facial features and a complete ending. This presenter and voice are synthetic.
Approval check: listen for a natural pause after the opening sentence and a complete delivery of “This presenter and voice are synthetic.” If the read exceeds 30 seconds, shorten the wording and regenerate the narration before proceeding.
Turn your script into an avatar video once the narration is ready.
4. Add the portrait and check the quote for the final inputs
Add your portrait using a file type and size accepted by the upload control. Confirm that the intended face is selected and that the narration attached to the avatar job is the version you approved.
Review the quoted credit cost before creating the clip. Check whether the amount covers speech, animation or both; a narration quote is not automatically the price of a finished video. If you change an input, review the updated quote before submitting.
Submit the avatar job once and follow its progress. If it appears stalled, inspect its status before starting another job. Keep the script, portrait and approved audio together so that a later revision does not accidentally use an older input.
5. Review speech, facial stability and synchronisation separately
When the result is available, use three passes rather than accepting it because the opening looks convincing.
- Listen without watching. Confirm the clip contains the approved words, with no missing opening or cut-off ending.
- Watch with the sound muted. Compare the face with the source portrait. Look for changing teeth, a shifting jaw or distracting head movement.
- Watch normally with sound. Check mouth alignment at the beginning, middle and end, including pauses between sentences.
A recurring issue in talking-head workflows is that head movement and mouth movement need separate judgement. More animation does not necessarily mean better lip-sync. A restrained presenter can be preferable to an expressive one whose movements distract from the message.
Where a download action is available, open the resulting file in the player or editor you intend to use. Confirm its format, dimensions, sound and ending there before planning a larger production around it.
See more of our reporting in Google Top Stories, AI Overviews and AI Mode.
Fix the faulty part before regenerating the video
| What you notice | Check first | What to change |
|---|---|---|
| A name or number sounds wrong | Does the same mistake occur in the narration? | Correct the spoken input, approve new audio, then regenerate the avatar. |
| The narration exceeds 30 seconds | The complete generated audio duration, including pauses. | Shorten the script before creating another video. |
| The audio is correct but the mouth looks distorted | The source portrait and whether the intended narration was used. | Try a clearer portrait while keeping the approved audio unchanged. |
| Synchronisation broke after an audio edit | Whether the replacement audio has different timing. | Use the final narration for a new avatar render rather than laying differently timed speech over the old animation. |
| Captions show phonetic spellings | Whether they were copied from pronunciation-adjusted input. | Restore the correct written spelling in the caption editor; do not change good narration to fix text. |
Keep a clean written script for captions alongside any pronunciation-adjusted speech input. Add or correct captions in a separate editor where necessary, then check them against the final clip.
This walkthrough starts from a still portrait. For changing speech on moving footage, use the AI video lip-sync guide. For a separate audio-led production across named tools, follow the HeyGen and ElevenLabs workflow.
Finish with the version whose words are correct, face remains stable, and the final sentence completes cleanly. Regenerate only after identifying what needs to change.