How to Make an AI Spokesperson Video from a Photo
To make an AI spokesperson video from a photo, pair a clear portrait with finished narration, generate the speaking presenter, then check the words and performance together. For a business introduction or customer FAQ, aim for one message and one next step within 20-30 seconds.
You can follow this workflow in the DIY AI Studio avatar generator. The script examples below show what to say; the preparation and review steps help you catch pronunciation, framing and mouth-movement problems before you publish.
- Choose one viewer question and write a short answer.
- Approve the presenter photo and finish the narration.
- Check the credit quote and generate one clip.
- Review the opening, names, mouth movement and call to action.
Give the presenter one job the viewer can act on
A short spokesperson clip can introduce a service, explain an appointment process or answer a common customer question. It is less useful when the viewer needs to inspect a product, follow several screen actions or hear a personal account from a real customer.
Choose the destination before writing. A welcome on your website can point to a nearby button. A video shared elsewhere needs an instruction that still makes sense away from that page. “Click below” becomes ambiguous when someone reposts the clip.
A photo also supplies no recorded performance. The system generates movement around the narration, so don’t expect it to reproduce the subject’s personal gestures. Keep the message factual, and clearly identify an artificial presenter where viewers could mistake it for a real employee or customer.
Use a script pattern that ends with one next step.
Start with roughly 45-60 written words for this brief, then time the actual narration. Pauses, names and spoken initials affect the running time. Cut words before increasing the delivery speed.
| Clip | Script pattern | What to leave out |
|---|---|---|
| Business introduction | Who you help, what they can do, where to start. | The full company history and a list of every service. |
| Customer FAQ | The customer’s question, the answer, any essential conditions, the next action. | Related questions that deserve their own answer. |
| Service announcement | What changes, when it changes, what the customer should do. | Internal reasons that do not affect the customer. |
Business introduction: edit 107 words down to 49
The deliberately overlong draft below gives the viewer several directions:
Welcome to DIY AI. We publish AI tool guides and provide a Studio for creating images, videos and audio. You can explore different creative tasks, compare tools, read tutorials and decide how each option might fit your work. If you want a short presenter video, you can start with a portrait and prepare narration using a recording or a voice you have permission to use. You will also find tools for other types of content, so there are several places to start depending on your project. Browse the guides, explore the Studio tools and open the avatar generator when you are ready to make your first clip.
The edited version keeps the business introduction and one starting point:
DIY AI helps you turn ideas into images, video and audio. Start with one task, choose a Studio tool and check the credit quote before you generate. For a short presenter video, bring a clear portrait and narration you have permission to use. Open the avatar tool to begin.
The edit removes the catalogue of activities and competing requests to browse, compare and explore. It also keeps the credit check before generation. These are written word counts; the exported narration must still be measured against the 20-30 second target.
Worked-video status: The script comparison is complete. The authorised presenter video has not been rendered yet.
Create your spokesperson video using one approved portrait and a finished narration track.
Customer FAQ: answer the question in the first sentence
Can I use my own voice for an avatar video? Yes. Upload a finished recording, or prepare narration with a voice you have permission to use. Listen to the whole track before generating the presenter. Open the avatar tool to prepare your clip.
This answer resolves the uncertainty before explaining the process. Use the same order for questions about booking, delivery, or getting started, and check conditions against your actual service.
Service announcement: make the date and action explicit
From [date], appointments at [business] will start at [time]. Existing bookings remain unchanged. For a new appointment, choose an available time on our booking page. Please check your confirmation email before travelling. Open the booking page to see the updated schedule.
Replace every bracketed field and verify that existing bookings really are unaffected. Use a full date when the clip may remain online beyond the week it is published.
Approve the portrait and narration before paying for animation
Use a portrait of one person with a clearly visible face, even lighting and space around the head and shoulders. Avoid a hand across the mouth, heavy shadows over the lips or a face that occupies only a small part of the image. HeyGen’s photo avatar guidance also recommends clear facial features and a front-facing image with a natural posture.
Have the person approve the message and intended use of their likeness. Check permission for the voice separately. A photograph does not supply the subject’s voice, and selecting a synthetic voice does not make it their real speech.
Listen to the narration without watching an avatar. Check business names, initials, prices and the final instruction. If the voice misreads “DIY AI”, try spelling out the letters in the speech input, then listen again. Keep normal spelling in captions.
Record one speaker without background music. If noise or echo masks the words, re-record or use a cleanup tool, as described in our Adobe Podcast Enhance Speech review. Compare the cleaned version with the original before using it.
Studio currently accepts uploaded narration as MP3 or PCM WAV, lasting 1-30 seconds and no larger than 2 MB. A short WAV can still exceed that size limit. Check both duration and file size after export.
See more of our reporting in Google Top Stories, AI Overviews and AI Mode.
Create one clip and check the price at the duration boundary
- Prepare the audio in Studio. Choose saved narration, upload a recording or create speech from your script. New speech has a separate credit quote. Reusing finished narration avoids another speech-generation charge.
- Add the approved portrait. Upload it or select an available image from Assets.
- Choose the framing. Studio applies padding to fit the selected frame; it does not generate the missing surroundings. Check how the portrait fits before proceeding.
- Review the quote and generate. Submit one clip and follow that job’s progress. A queue delay is not a reason to resubmit the same request.
- Watch and download the result. Completed avatar videos appear in Assets and the selected project. Save the accepted file within the storage period shown in your account.
As of 8 October 2026, Studio’s published avatar-video rate is 4 base credits plus 4 credits per additional 5 seconds:
| Narration duration | Avatar-video credits |
|---|---|
| 20 seconds | 20 |
| More than 20, up to 25 seconds | 24 |
| More than 25, up to 30 seconds | 28 |
A 25.1-second track falls into the final band. Trim unnecessary silence without clipping the last word, then check the current quote. These figures exclude new speech generation and are not charges from a completed example.
Review the opening and the last instruction as carefully as the lips
Creator discussions repeatedly raise a useful problem: plausible mouth movement can still accompany a stiff or mismatched performance. Review the complete delivery at normal speed before inspecting individual frames.
| Check | Accept when | Fix before publishing |
|---|---|---|
| Opening | The first sentence immediately introduces the service or answers the question. | A long greeting, empty pause or unexplained opening claim. |
| Names and facts | Names, dates, numbers and conditions match the approved script. | Mispronunciation or wording that changes the offer. |
| Mouth and expression | The face stays consistent, and speech looks reasonably aligned throughout. | Persistent timing errors, changing teeth or distracting movement during pauses. |
| Call to action | The final instruction is audible and its destination exists. | A cut-off ending or a request with no usable link or button. |
| Captions and framing | Captions match the final audio and remain readable on a phone. | Misspelt names, text covering the mouth or an awkward crop. |
Add and correct captions after accepting the performance; our CapCut review covers an editor for that finishing stage. If narration changes, regenerate the avatar with the corrected track. Replacing the sound alone can leave the mouth following the old words.
Choose a real recording when the person is part of the evidence
Record the actual person for a founder’s personal statement, a customer testimonial or a demonstration where handling the product matters. A generated presenter can explain your message, but cannot establish that someone personally used a service or experienced a result.
For longer lessons, several scenes or recurring presenters, assess the production workflow using our AI talking head generator comparison. For adverts needing product shots, variations and a full edit, use the AI ad video generator guide.
Keep the approved portrait, final script, narration and accepted export together. When the message changes, those files let you revise the clip from a known starting point.