Text to Speech Avatar: Turn a Script into a Talking Photo

Text to Speech Avatar: Turn a Script into a Talking Photo

A text to speech avatar turns written words into speech and animates a presenter to deliver them. In the DIY AI avatar generator, start with one authorised portrait and a short script. This tutorial covers the text-input workflow, including checks for pronunciation, timing, cost, and the final animation.

The workflow: prepare your portrait and script, select Text & voice, approve the narration, add the portrait, review the avatar quote and create the clip. Keep the narration within the 30-second limit. Finalise what the viewer will hear before paying to animate it.

A talking photo needs a face and a voice, approved separately

The portrait supplies the presenter’s appearance; text-to-speech supplies the narration. The animation then makes the person appear to speak. A photograph alone cannot reproduce its subject’s voice: Microsoft’s text to speech avatar overview describes avatar appearance and voice as separate choices.

A preset voice is enough for this walkthrough. A recognisable personal voice requires a separate, authorised process, such as cloning your own voice. Do not assume that uploading a headshot performs both jobs.



1. Prepare a portrait with an unobstructed mouth

Choose a clear, front-facing photograph of one person, with the whole face visible and some space around the head. Avoid group photographs, extreme angles, heavy shadows and objects covering the mouth. A simple head-and-shoulders composition gives you fewer distracting details to inspect later.

Check the original lips, teeth and jaw before uploading. If they already look distorted, replace the source rather than expecting animation to repair it. Keep a copy of the original for comparison with the finished video.

Use your own portrait or ask the subject to approve both the animation and the words they will appear to say. Make the clip’s synthetic nature clear wherever viewers might mistake it for a real recording.

2. Write spoken words, not instructions for the presenter

For a first attempt, aim for roughly 35-45 words. That is a drafting recommendation, not a guaranteed duration: the generated narration determines whether the script fits. Leave room for pauses rather than treating 30 seconds as a target to fill.

Paste only the words the audience should hear. Remove headings, speaker labels and directions such as “smile at the camera”. Do not assume the narration field understands performance instructions or speech markup.

Write numbers according to their meaning. For example, “version three point two” gives the voice a clearer instruction than “v3.2”. If the script runs long, remove a sentence before trying to squeeze it into faster speech.

3. Select Text & voice and listen to your actual script

Open the avatar workspace and sign in. Select Text & voice, enter your script and choose an available voice suitable for its language and delivery. Review any speech-generation charge before confirming the narration.

A voice sample helps you choose a sound; it doesn’t prove your names, figures, or sentences will be spoken correctly. Generate the narration and listen from beginning to end. Check the actual duration, the pronunciation of unfamiliar words, and whether the final phrase ends comfortably.

Fix mistakes now. Use the detailed guide to fix AI voice pronunciation for names, acronyms and awkward numbers. Test a correction in the complete sentence, because a word that works alone may sound different in context.

Once the narration is correct, keep that version unchanged for the avatar render. “Approved” here means you have listened and accepted it, not merely that the tool successfully generated an audio file.

Worked example: a short introduction that discloses the animation

Use one authorised portrait and an available English voice for this 42-word script. It is a reusable example, not a completed test result; check the generated audio’s duration before creating the video.

Welcome to this short demonstration. A talking photo starts with a portrait and a written script. Listen to the voice first, then check the finished video for clear speech, stable facial features and a complete ending. This presenter and voice are synthetic.

Approval check: listen for a natural pause after the opening sentence and a complete delivery of “This presenter and voice are synthetic.” If the read exceeds 30 seconds, shorten the wording and regenerate the narration before proceeding.

Turn your script into an avatar video once the narration is ready.

4. Add the portrait and check the quote for the final inputs

Add your portrait using a file type and size accepted by the upload control. Confirm that the intended face is selected and that the narration attached to the avatar job is the version you approved.

Review the quoted credit cost before creating the clip. Check whether the amount covers speech, animation or both; a narration quote is not automatically the price of a finished video. If you change an input, review the updated quote before submitting.

Submit the avatar job once and follow its progress. If it appears stalled, inspect its status before starting another job. Keep the script, portrait and approved audio together so that a later revision does not accidentally use an older input.

5. Review speech, facial stability and synchronisation separately

When the result is available, use three passes rather than accepting it because the opening looks convincing.

  1. Listen without watching. Confirm the clip contains the approved words, with no missing opening or cut-off ending.
  2. Watch with the sound muted. Compare the face with the source portrait. Look for changing teeth, a shifting jaw or distracting head movement.
  3. Watch normally with sound. Check mouth alignment at the beginning, middle and end, including pauses between sentences.

A recurring issue in talking-head workflows is that head movement and mouth movement need separate judgement. More animation does not necessarily mean better lip-sync. A restrained presenter can be preferable to an expressive one whose movements distract from the message.

Where a download action is available, open the resulting file in the player or editor you intend to use. Confirm its format, dimensions, sound and ending there before planning a larger production around it.

Your search, your sources
Make DIY AI a preferred source

See more of our reporting in Google Top Stories, AI Overviews and AI Mode.

Add DIY AI on Google

Fix the faulty part before regenerating the video

What you noticeCheck firstWhat to change
A name or number sounds wrongDoes the same mistake occur in the narration?Correct the spoken input, approve new audio, then regenerate the avatar.
The narration exceeds 30 secondsThe complete generated audio duration, including pauses.Shorten the script before creating another video.
The audio is correct but the mouth looks distortedThe source portrait and whether the intended narration was used.Try a clearer portrait while keeping the approved audio unchanged.
Synchronisation broke after an audio editWhether the replacement audio has different timing.Use the final narration for a new avatar render rather than laying differently timed speech over the old animation.
Captions show phonetic spellingsWhether they were copied from pronunciation-adjusted input.Restore the correct written spelling in the caption editor; do not change good narration to fix text.

Keep a clean written script for captions alongside any pronunciation-adjusted speech input. Add or correct captions in a separate editor where necessary, then check them against the final clip.

This walkthrough starts from a still portrait. For changing speech on moving footage, use the AI video lip-sync guide. For a separate audio-led production across named tools, follow the HeyGen and ElevenLabs workflow.

Finish with the version whose words are correct, face remains stable, and the final sentence completes cleanly. Regenerate only after identifying what needs to change.

You Might Also Like:

How to Make an AI Birthday Video with Your Own Photo

AI Birthday Video

By: Steven Jones On:
To make an AI birthday video with your own photo, start with a portrait of you delivering the greeting, write…
12 AI Avatar Prompts for Talking Video Portraits

AI Avatar Prompts

By: Steven Jones On:
These AI avatar prompts are designed to create a clear portrait for a talking video: one person, a readable face…
Steven Jones

Writer: Steven Jones

AI Tools Reviewer and Technical Analyst

Steven Jones is a technology analyst specialising in artificial intelligence, machine learning workflows, and emerging automation tools.

At DIY AI, he focuses on clear, practical guidance for people comparing AI tools in the real world. His work covers text generation, image generation, video tools, data platforms, developer-focused AI products, and the automation workflows that connect them.

Steven's reviews are built around hands-on testing, practical benchmarks, and transparent scoring rather than vendor claims. He looks closely at where each tool performs well, where it falls short, and what those trade-offs mean for creators, teams, and businesses trying to make sensible AI adoption decisions.

He has a particular interest in safety, reliability, output quality, performance metrics, and dataset quality. When he is not reviewing the latest AI model updates, he experiments with prompt engineering techniques and contributes to DIY AI ongoing work on fair, explainable scoring frameworks for AI tools.

Contact

Leave a Comment On: Text To Speech Avatar

Your email address will not be published.