12 AI Avatar Prompts for Talking Video Portraits

12 AI Avatar Prompts for Talking Video Portraits

These AI avatar prompts are designed to create a clear portrait for a talking video: one person, a readable face and enough room around the head. They cover instructors, founder messages, customer support and nine other uses, with a separate narration example and checks for deciding whether a generated image is worth animating.

  1. Create your portrait in the DIY AI image generator using one of the prompts below.
  2. Bring the selected image into the DIY AI avatar generator, where the portrait and narration become a talking video.
  3. Review a short clip before committing to the full script. A convincing still doesn’t prove the mouth will animate naturally.

Talking avatar portrait prompts control appearance; scripts control speech

The image prompt describes the face, clothing, framing, lighting and background. The narration script contains the words the presenter should say. Put each into its corresponding part of the workflow: pasting “front-facing portrait, soft lighting” into a speech field can make the voice read those instructions aloud.

Keep gestures and delivery directions out of the spoken text unless the selected tool explicitly supports them. A portrait instruction such as “calm expression” also cannot guarantee calm delivery. Listen to the chosen voice before rendering.

For products, landscapes and other photographic subjects, use our broader AI photo prompt guide. The templates here deliberately prioritise a face that remains easy to inspect.



A reusable portrait prompt that keeps the face unobstructed

Replace the bracketed fields, then generate one image. Set the actual image dimensions in the generator; writing “4K” in the prompt does not set the file resolution.

Realistic photographic portrait of one fictional adult [face, age range and hairstyle], wearing [plain clothing], against [simple background]. Eye-level camera, face and shoulders directly towards the camera, both eyes looking into the lens. Head-and-shoulders framing, entire head visible with comfortable headroom, neck and shoulders included, hands outside the frame. Relaxed jaw, gently closed lips and a subtle approachable expression. [Soft, even lighting], sharp eyes and mouth, natural skin texture. Hair and accessories leave the eyes, cheeks and lips visible. No other people, reflected faces, text, logos or objects crossing the face.

The closed-mouth expression is a conservative starting point for these templates, not a universal upload requirement. Avoid combining it with “laughing”, “speaking enthusiastically” or other conflicting directions. Keep personality in the clothes and setting while you establish a usable source image.

AI avatar prompt examples: 12 roles with a clear facial view

Each prompt below is a standalone starting point. They are proposed portrait templates, not a set of 12 animation-tested presets. Adapt the person’s appearance to your project; use an authorised reference photograph when the presenter must depict a particular real person.

1. Instructor: leave teaching aids outside the portrait

Realistic portrait of one fictional adult instructor wearing a plain muted blue crew-neck top. Face and shoulders directly towards an eye-level camera, eyes looking into the lens, gently closed lips. Entire head, neck and shoulders visible with generous headroom. Soft, even frontal light, natural skin texture, blurred warm grey teaching-room background. Hands outside frame; no board, writing or other people.

Put diagrams in the lesson edit. A generated board behind the instructor adds details that may need correcting later.

2. Founder message: make the role clear without inventing an identity

Realistic head-and-shoulders portrait of one fictional adult presenter for a founder message, wearing a navy overshirt over an off-white T-shirt. Face and shoulders square to the camera, direct eye contact, relaxed closed mouth. Full head visible with space above the hair. Even diffused light, natural skin texture, softly blurred off-white office wall. No logos, hands or other people.

Introduce a generated person as a virtual presenter. Use the founder’s approved likeness if the video is meant to show the founder themselves.

3. Customer support: keep the microphone away from the lips

Realistic portrait of one fictional adult customer-support presenter in a plain muted teal blouse. Hair neatly clear of the face, both eyes and lips fully visible. Front-facing at eye level, gently closed mouth, relaxed shoulders. Full head and upper chest in frame with comfortable headroom. Soft frontal lighting and a blurred pale grey workspace. No headset, microphone, text or additional people.

The script can establish the support role. A headset is unnecessary and introduces another object around the cheek and mouth.

4. Product onboarding: reserve the interface for a screen recording

Photographic portrait of one fictional adult software onboarding presenter in a plain charcoal crew-neck top. Centred head-and-shoulders view, face directly forward, eye-level camera, relaxed closed lips. Whole head visible with extra space above the hair. Soft balanced lighting, sharp facial features and natural skin texture. Simple pale blue background, hands out of frame, no screens, interface text or other faces.

Pair the finished introduction with a real screen capture so the instructions match the product.

5. Service announcement: choose an expression that fits bad news too

Realistic portrait of one fictional adult service-announcement presenter wearing a plain dark grey shirt. Composed neutral expression, relaxed jaw and closed lips, direct eye contact. Face and shoulders square to the camera, full head and upper chest visible with headroom. Even soft light, natural skin texture, plain warm grey background. No exaggerated smile, hands, logos or other people.

A broad smile may suit a launch but look misplaced beside a delay or outage notice.

6. Webinar host: add event branding after generation

Realistic head-and-shoulders portrait of one fictional adult webinar host in a plain burgundy jacket and neutral top. Face directly forward, eye-level camera, attentive eyes and softly closed lips. Entire head visible with comfortable headroom. Diffused frontal lighting, natural facial texture, softly blurred neutral studio background. Hands outside frame; no stage, audience, event lettering or logos.

Add the event title as editable text in the final video, rather than asking the image model to print it.

7. HR welcome: avoid implying a fictional person is an employee

Photographic portrait of one fictional adult virtual induction presenter wearing a plain sage green knit top. Relaxed, welcoming expression with closed lips, facing the lens directly. Full head, neck and both shoulders visible with spare space above the hair. Soft even lighting, natural skin texture, blurred cream office background. No staff badge, hands, lettering or other people.

Label the presenter appropriately and put the real team’s contact details beside the video.

8. Museum guide: keep exhibits from competing with the face

Realistic portrait of one fictional adult museum audio-guide presenter in a plain brown jacket. Front-facing head-and-shoulders composition, direct eye contact, relaxed closed mouth, whole head visible with headroom. Soft balanced light and natural skin texture. Gently blurred stone-coloured gallery wall with no discernible artwork, statues, visitors or writing. Hands and props outside the frame.

Show the actual exhibit separately, especially when accurate historical details matter.

9. Language tutor: make the lip outline easy to inspect

Realistic photographic portrait of one fictional adult language tutor wearing a plain rust-coloured top. Eye-level camera, face squarely forward, direct gaze, relaxed jaw and gently closed lips. Hair clear of cheeks and mouth. Entire head and shoulders visible with generous headroom. Soft even frontal lighting, sharp lip detail, natural skin texture, plain cream background. No hands, text or other faces.

Review pronunciation and mouth movement separately. Do not use an unchecked avatar as a visual demonstration of speech articulation.

10. Property welcome: separate the presenter from property evidence

Realistic head-and-shoulders portrait of one fictional adult property-welcome presenter wearing a plain navy jacket. Face and shoulders directly towards the camera, eyes on the lens, subtle closed-mouth expression. Full head visible with comfortable headroom. Soft daylight balanced across the face, natural skin texture, neutral light grey background. No property interior, signage, keys, hands or additional people.

Use genuine listing photographs in the edit; a generated room should not become an apparent feature of the property.

11. Podcast introduction: suggest the setting without a foreground mic

Realistic photographic portrait of one fictional adult virtual podcast presenter wearing a plain black crew-neck top. Direct frontal head-and-shoulders view, camera at eye level, attentive eyes and gently closed lips. Whole head visible with headroom. Soft warm light evenly across the face, natural skin texture, blurred acoustic-panel background. No microphone, headphones, hands, lettering or other people.

Keep the source portrait clear. Add programme artwork around the finished clip if needed.

12. Course recap: reuse the original instructor’s face

Using the attached approved instructor portrait as the identity reference, create one realistic head-and-shoulders portrait of the same adult in a plain oatmeal jumper against a light grey background. Preserve facial proportions, age, skin tone and hairstyle. Face directly forward, direct eye contact, relaxed closed lips, full head visible with headroom. Soft even lighting. No hands, text or other people.

This template requires reference-image support. Check for identity changes; if only the lesson changes, reuse the original portrait instead.

Your search, your sources
Make DIY AI a preferred source

See more of our reporting in Google Top Stories, AI Overviews and AI Mode.

Add DIY AI on Google

Worked example: three portrait directions, with animation still to verify

We generated the instructor, founder-message, and support directions together as one composite illustration, using shared instructions for frontal faces, closed lips, and even lighting. This checks the visual direction only. It does not confirm that separate prompt runs will reproduce these faces or that the portraits animate successfully.

Three portrait directions, with animation still to verify
Generated portrait concepts, from left: instructor, founder-message presenter and customer support. Animation validation is pending. Upload one individual portrait to an avatar tool, not this three-person composite.

All three faces have visible eyes and closed lips, with no object covering the mouth. The crop leaves relatively little space above the hair, despite the prompt requesting headroom. For a new source image, ask for a slightly wider crop and inspect the actual result before uploading.

Separate narration script for an instructor test:

Welcome. Before we begin, please open your workbook. We will build one short example together, then pause so you can try it.

This proposed script includes lip-closing sounds in “before”, “please” and “workbook”, giving the reviewer specific moments to inspect. Its rendered duration and lip synchronisation have not been measured.

Animation test pending: no completed talking-video output is available for these portrait concepts. They remain illustrative starting points, not validated source portraits.

Turn your portrait into a talking avatar after you select an individual source image and approve the narration.

Reject portrait problems before paying for animation

Inspect the image at its original size and at the size viewers will see it. HeyGen’s photo avatar guidance recommends clear facial features, visible eyes and lips, and a character large enough in the image to animate. These are useful source-image checks; they do not guarantee results in another engine.

CheckReason to revise the imagePractical fix
One intended faceA mirror, poster or background person adds another face.Simplify the setting and remove reflected or background faces.
Eyes and mouthGlare hides an eye, hair crosses the lips, or the mouth has malformed detail.Correct the portrait before testing speech.
Headroom and shouldersHair touches the top edge, or the intended framing cuts into the chin.Generate a wider source crop. Padding adds canvas space, not missing anatomy.
ExpressionThe source has an exaggerated grin or a visibly strained jaw.Try a relaxed expression while keeping the rest of the prompt stable.
IdentityThe presenter changes between lessons or language versions.Reuse the approved source and voice; compare any edited portrait against the original.

A recurring problem in creator discussions is that small changes to the face, lighting and voice become noticeable across a series. Keep the approved image, exact prompt and chosen voice together. A text description alone is a weak way to preserve the same presenter across repeated generations.

Add speech in Studio and review the first short output

Studio’s avatar workflow accepts saved narration, an uploaded recording or newly generated speech. Use the AI voice generator to prepare narration, or the optional voice-cloning workflow for an authorised voice. Cloning is not required.

  1. Approve the audio first. Listen for names, pronunciation, pauses and a complete ending. Uploaded narration currently needs to be MP3 or PCM WAV, 1-30 seconds long and no larger than 2 MB.
  2. Select the individual portrait. Upload it or choose it from Assets. Check the chosen framing: Studio uses padding rather than expanding the portrait into a wider scene.
  3. Check both costs. Newly generated speech has a separate quote from avatar rendering. Review the current video quote before submitting a short test.
  4. Watch and listen together. Inspect the first spoken word, lip closures, teeth, jaw movement and final pause. Check that the face still resembles the source.

If pronunciation is wrong, revise the narration. If the still already has an obscured mouth, revise the portrait. If both inputs look acceptable but the animation remains unnatural, try another approach rather than repeatedly adding adjectives to the image prompt.

For choosing an engine, see our AI talking-head generator comparison. Keep the source portrait fixed while comparing outputs so you can judge the animation rather than a change of face.

You Might Also Like:

How to Make an AI Spokesperson Video from a Photo

AI Spokesperson Video

By: Steven Jones On:
To make an AI spokesperson video from a photo, pair a clear portrait with finished narration, generate the speaking presenter,…
Steven Jones

Writer: Steven Jones

AI Tools Reviewer and Technical Analyst

Steven Jones is a technology analyst specialising in artificial intelligence, machine learning workflows, and emerging automation tools.

At DIY AI, he focuses on clear, practical guidance for people comparing AI tools in the real world. His work covers text generation, image generation, video tools, data platforms, developer-focused AI products, and the automation workflows that connect them.

Steven's reviews are built around hands-on testing, practical benchmarks, and transparent scoring rather than vendor claims. He looks closely at where each tool performs well, where it falls short, and what those trade-offs mean for creators, teams, and businesses trying to make sensible AI adoption decisions.

He has a particular interest in safety, reliability, output quality, performance metrics, and dataset quality. When he is not reviewing the latest AI model updates, he experiments with prompt engineering techniques and contributes to DIY AI ongoing work on fair, explainable scoring frameworks for AI tools.

Contact

Leave a Comment On: AI Avatar Prompts

Your email address will not be published.