AI Video Prompt Examples: Camera Moves, Actions and Fixes

AI Video Prompt Examples

Start an AI video prompt with one visible action and one clear camera instruction. For text-to-video, describe the scene as well. For image-to-video, let the source image establish the appearance and concentrate on what should move. Add further instructions only when they solve a specific problem.

Try this scene in Studio: copy an example below, choose an enabled video model and check the available mode, settings and credit requirement before generating. Use the image-to-video version only when your selected model offers a source-image input.

Example status: These are original starting prompts and suggested troubleshooting steps, not completed generation tests. No success rates, credit-spend results or model winners are claimed for these examples.

AI video prompts at a glance

Choose the shot you need, then use its full prompt below. The source requirements apply to image-to-video; the text-to-video versions describe those starting scenes instead.

Intended shotSource requirementSubject actionCamera movementFirst adjustment to test
Workshop push-inWide interior with a foreground object and a clear path towards a rear workbenchFurniture stays stillSlow forward dollyShorten the move if room geometry changes
Side-on cycling shotRider and both wheels visible, with space around the bicycleSteady pedallingTrack alongside at matching speedTest a fixed camera separately to isolate rider-motion problems
Overhead hand-and-tile actionOne hand already touching one tile on an uncluttered tableSlide the tile a short distanceLocked overhead shotReduce the slide to a small nudge if contact breaks
Curtain movementCurtain, rail and window frame clearly visibleLower fabric lifts gently and settlesLocked shotRestrict movement to the lower curtain edge


Image-to-video prompts and text-to-video prompts do different jobs

Text-to-video builds the scene: describe the subject, starting position, surroundings, action, camera and lighting. A useful working template is:

[Shot and scene]. [Subject and starting position]. [One main action]. [Camera direction]. [Lighting or visual treatment].

Image-to-video animates a supplied scene: identify what moves, describe its movement, and state the camera behaviour. Avoid rewriting the image as a different setting unless transformation is actually your goal.

[Visible subject] performs [specific movement]. The camera [movement or fixed position]. [Important visible detail] remains unchanged.

Google recommends motion-focused wording for Veo image-to-video. Kling also separates its subject-and-movement approach for image inputs from fuller scene descriptions for text inputs. These are starting structures, not special commands that guarantee compliance.

A reference used to guide a subject is not necessarily the same as an image supplied as the first frame. Check the selected mode before expecting an exact opening composition.

Four copyable AI video prompt examples

Use the version that matches your input mode. Each example includes a source specification, an inspection target and a proposed first revision. None guarantees the first generation will pass.

1. A slow push-in through a workshop

Source image: Use a wide image of an empty printmaking workshop. Include a printing press in the left foreground, a rear workbench and a clear aisle between them. Keep useful straight edges visible so you can judge whether the room changes shape.

Image-to-video prompt:

The camera moves slowly forward towards the rear workbench in one continuous shot. The foreground printing press moves towards the left edge of the frame as the camera passes. The furniture stays still and the lighting remains unchanged.

Text-to-video prompt:

Wide eye-level shot inside an empty printmaking workshop. A printing press occupies the left foreground, with a wooden workbench at the back and a clear central aisle. The camera moves slowly forward along the aisle towards the workbench. Furniture remains still. Soft window light, natural colours, one continuous shot.

What to inspect: Check that the viewpoint advances through the room rather than merely enlarging the whole image. Reject a clip if shelving bends, the workbench changes shape, or the camera cuts to a different view.

First adjustment: Shorten the requested travel first. If the result looks like a zoom, clarify the foreground movement rather than adding more cinematic adjectives. A precisely measured camera path may need controls beyond a text prompt.

2. A side-on tracking shot of a cyclist

Source image: Choose a side view of a cyclist in a clear, traffic-free courtyard. Both wheels, the rider and the ground should be visible. A plain background makes it easier to spot changes in the bicycle or camera direction.

Image-to-video prompt:

The cyclist pedals steadily from left to right. The camera travels alongside at matching speed, keeping the rider and both wheels fully in frame. The brick wall moves past in the background. One continuous side-on shot.

Text-to-video prompt:

Side-on full shot of an adult cyclist riding slowly from left to right through an empty paved courtyard beside a brick wall. The camera travels parallel to the bicycle at matching speed. Keep the rider and both wheels fully visible. Overcast daylight, natural motion, one continuous shot.

What to inspect: Watch the pedals, wheels and ground contact throughout the clip. A rider staying centred is not enough: the background must also behave consistently with the requested tracking movement.

First adjustment: To diagnose a failure, replace the tracking instruction with a fixed-camera version and reduce the riding speed. This is a separate, simpler test, not an accepted tracking shot. Restore camera travel only after deciding whether rider motion itself is usable.

3. An overhead hand moving one tile

Source image: Use an overhead image of one blue wooden tile on a plain tabletop. A fingertip should already touch the tile’s left edge. Avoid extra pieces, written markings or an initial hand position that requires a complicated reach.

Image-to-video prompt:

The finger slides the blue tile a short distance to the right, then stops while maintaining contact with its left edge. The tile stays flat on the table. The overhead camera remains fixed.

Text-to-video prompt:

Direct overhead close-up of a plain tabletop. One blue wooden tile rests near the centre, with a fingertip touching its left edge. The finger slides the tile a short distance to the right and stops, maintaining contact. The tile stays flat. Fixed camera, even daylight, one continuous shot.

What to inspect: Inspect the contact point and the final tile position. Reject a clip if the finger passes through the tile, the tile starts moving before contact, or its shape changes during the slide.

First adjustment: Reduce the movement to a small nudge. If contact still breaks, try a clearer source or a different viewing angle as a separate experiment. Do not assume that adding “perfect hands” will repair the underlying interaction.

4. A locked shot with gentle curtain movement

Source image: Choose an image with the curtain, its attachment point and the window frame visible. This gives you moving fabric to inspect against straight, stationary edges.

Image-to-video prompt:

The lower edge of the curtain lifts gently towards the room, then settles. The curtain remains attached to the rail. The window frame, furniture and camera stay still. Daylight remains steady.

Text-to-video prompt:

Medium-wide view of a quiet room with a lightweight linen curtain beside an open window. A faint breeze lifts the lower curtain edge into the room, then lets it settle. The curtain remains attached to its rail. The camera is fixed, with steady soft daylight and natural fabric texture.

What to inspect: Check the window corners and furniture while the fabric moves. Reject a clip if the whole room appears to breathe, the rail shifts, or the curtain separates from its attachment point.

First adjustment: Ask only for a slight sway at the lower edge. If the background still changes, check for supported masking or regional motion control rather than repeatedly expanding the prompt. This example requests a single shot, not a seamless loop.

Camera movement words: say what should actually change

A pan, a sideways camera move and an orbit describe different shots. Use the movement term plus a plain-language instruction, rather than stacking several terms into the same sentence.

DirectionMeaningPrompt fragment
Locked cameraViewpoint stays stillThe camera remains fixed while the curtain moves.
PanCamera turns left or right from one positionFrom a fixed position, pan slowly right across the workbenches.
TiltCamera turns up or down from one positionTilt upwards from the floor to the workshop ceiling.
Dolly / push-inCamera moves physically closerMove forward along the aisle towards the rear workbench.
Truck / lateral trackingCamera moves sideways; tracking follows the subjectTravel beside the cyclist at matching speed.
Arc / orbitCamera travels around the subjectMove through a shallow arc around the stationary sculpture.
ZoomFraming magnifies without moving the camera positionFrom a fixed position, slowly zoom towards the wall clock.

A rack focus changes what is in focus; it does not require moving the camera. For example: “Focus shifts from the foreground tools to the rear workbench while the camera remains fixed.” Treat this as a request to test, not an available focus-control slider.

Simple versus overloaded prompts: compare the same source

For the overhead tile scene, keep the exact same source image, model and settings. Version A asks for a short slide with a fixed camera. Version B deliberately adds competing demands:

Version B: deliberately overloaded, not recommended as the starting prompt.

The finger slides the blue tile to the right, flips it into the air, catches it and stacks it on another tile while the other fingers tap the table. The camera circles overhead, pushes closer and shifts focus between the hand and tile. Add dramatic moving shadows and a glossy commercial finish. Keep the camera fixed and every detail identical to the source.

The conflict is visible before generating anything: the camera cannot remain fixed while also circling and moving closer. The tile now needs a slide, airborne rotation, catch and stack instead of one contact action. An additional tile must also appear.

This comparison tests instruction workload and consistency, not prompt length alone. A longer prompt can help when extra detail resolves ambiguity. These two versions cannot establish that fewer words always produce better video, and neither has been tested for this article.

A focused revision would keep the original job and reduce its travel:

The fingertip nudges the blue tile slightly to the right while staying in contact. The tile remains flat. The overhead camera stays fixed.

Generate revisions as separately labelled attempts. Do not replace failed originals in the record, and do not count a different, easier shot as proof that the original request succeeded.

Veo prompts: separate visual direction from optional sound

Google’s video generation prompt guide distinguishes subject, action, scene, camera movement and audio direction. Use only the elements relevant to the shot, rather than trying to include every available term.

For a Veo text-to-video attempt, use one of the full scene descriptions above. For image-to-video, start with its motion-only counterpart. To adapt the workshop example for an audio-capable route, add a separate sentence:

Sound: quiet workshop room tone with a faint ventilation hum.

Add that instruction only when the selected model and access route support generated audio. For the initial camera test, keep audio settings fixed across versions. Judge sound separately so a convincing soundtrack does not hide a failed camera move.

Where a Veo interface exposes a dedicated negative-prompt field, Google recommends naming unwanted elements rather than writing commands. For an unwanted-cut problem, a small exclusion such as “cuts, scene changes” belongs in that field, not as an unexplained list pasted into the main prompt. Check that your route actually supports it.

Kling AI prompts: attach each action to a visible subject

Kling’s image-to-video guidance centres on identifying subjects and their movements. For a scene with several objects, make it clear which one acts instead of saying only “move to the right” or “add motion”.

For the tile example, a Kling-oriented starting version is:

The fingertip pushes the blue tile a short distance to the right. The tile slides flat across the tabletop and stops with the fingertip still touching it. The tabletop stays still. Fixed overhead camera.

For text-to-video, add the tabletop, tile, starting hand position and lighting from the full scene description. Do not assume an image-to-video prompt provides enough information when there is no image.

Kling documents camera controls and start/end-frame features separately from prompt wording. Check which features your selected version and interface expose. Their existence in Kling’s own tools does not establish that every Studio integration offers them.

Diagnose the failure before changing the prompt

Use the visible symptom to choose the next experiment. The possibilities below are diagnostic hypotheses, not confirmed causes for a particular generated clip.

Observed problemWhat to investigateNext experiment
The subject barely movesAn unclear action, an ambiguous subject or a source that does not support the requested movementName the moving object and one visible verb; keep the camera fixed
The camera moves the wrong wayConflicting directions or a mismatch between the prompt and a camera presetKeep one camera instruction and remove or align any preset
A hand passes through an objectUnclear starting contact or an interaction the model is not handlingShorten the action; inspect the contact point before adding complexity
Walls, wheels or furniture change shapeThe movement requires too much reconstruction of the sourceReduce travel or angle; compare the full clip with the original image
An unwanted cut appearsSeveral scene descriptions, incompatible framing instructions or model behaviourRequest one continuous shot and remove changes of setting
The clip jumps when repeatedAn end-to-start continuity problem rather than ordinary prompt adherenceTreat loop creation and seam repair as a separate task

If changing wording does not address the same recurring failure, revisit the source, available controls or model. More adjectives are not a substitute for identifying what broke.

Your search, your sources
Make DIY AI a preferred source

See more of our reporting in Google Top Stories, AI Overviews and AI Mode.

Add DIY AI on Google

What a prompt cannot replace?

A request in the text box is not the same as an input or setting exposed by the selected workflow. Check these before spending more credits:

RequirementWhat to use or verify
A specific starting imageA supported image-to-video or first-frame input; text alone does not attach your file
A particular ending compositionAn exposed end-frame or keyframe control; prose can request an ending but cannot supply a missing image input
A prescribed camera path or motion regionA compatible camera, path, mask or regional-motion control where available
A particular duration, resolution or aspect ratioThe actual generation settings; writing “4K” or “vertical” does not change the selected export configuration
Generated speech, sound or lip syncA model and route with the required audio capability; otherwise plan a separate audio workflow

Even a supported control is something to verify in the output, not a reason to skip review. For a single source image, asking to reveal an unseen side also asks the model to invent visual information that the image does not contain.

How to test prompt revisions and record retry costs

Suggested protocol, not completed results: use the four image-to-video scenes above, two fixed prompt versions for each scene and three attempts per version. That is 24 planned attempts per model. Run text-to-video as a separate experiment rather than mixing it into the same-source comparison.

Choose a model that is enabled when you run the test. Keep its version, input mode, source file, duration, resolution, aspect ratio, quality mode and audio settings unchanged. Keep prompt enhancement off or unchanged where possible and record any automatic rewriting you can inspect. If seeds are exposed, use the same three seeds for the paired versions and record them.

RecordDetails to save for every attempt
IdentityShot ID, prompt version, attempt number, date, model/version and access route
Inputs and settingsSource filename, exact prompt, available generation settings, seed and enhancement state; mark hidden fields as not exposed
Submission and outputJob ID, completed/failed status, original video file and any later edit
Actual spendQuoted credits, credits debited and any confirmed refund; do not assume a failed job was free
AcceptancePass/fail against the original shot brief, rejection reason and the timestamp where it appears

Decide the acceptance rules before viewing results. Require the intended action, the requested camera behaviour and stable essential details across the whole clip. A technical failure and a completed but unusable video are different outcomes; retain both in the log. Record additional retries separately rather than silently replacing failures.

Net credits per accepted clip = all credits actually debited for the test, minus confirmed refunds, divided by accepted clips. If nothing passes, report “no accepted outputs” and the spend so far, not a zero cost per usable clip. Keep Studio credits separate from provider credits or money unless you have a verified conversion for that purchase.

Three attempts per version provide a small diagnostic sample, not a reliable model ranking. A JPG contact sheet can document opening, middle and final frames, but only the videos show whether movement and physical contact hold together between those checkpoints.

AI video prompt FAQs

How long should an AI video prompt be?

Use enough detail to define the shot without adding unrelated tasks. There is no universal winning length. Studio’s public guide currently lists a 1,000-character prompt limit, so check the interface limit before pasting a longer example. Prioritise the action and camera direction over decorative adjectives.

Should I use negative prompts?

Only use a dedicated negative-prompt field when your model and interface support it. Otherwise describe the desired result directly: “the camera remains fixed” is clearer than a long list of unwanted camera movements. Neither approach guarantees compliance.

Can the same prompt work in Veo and Kling?

Use the same core shot brief as a starting comparison, but confirm matching input modes and record each model’s settings. These examples have not established equivalent performance. A provider-specific control may matter more than changing a few words.

Can a prompt guarantee an exact logo, face or piece of text?

Do not treat a preservation instruction as a guarantee. Inspect important details across the entire clip. Where exact artwork is essential, plan a workflow that retains or composites the approved asset rather than relying solely on generated reconstruction.

Choose the shot first, then judge the whole clip

Start with one example, define what would make it usable and change one relevant instruction at a time. When a prompt requests a different shot, label it as a new test rather than quietly lowering the acceptance standard.

For tool selection, use our AI image-to-video tools comparison or broader AI video generator comparison. This guide is for directing and diagnosing the shot after that choice.

Try your chosen scene in Studio. Paste the matching prompt, review the settings and credit requirements, and keep the first attempt even if it fails. It gives your next revision something specific to address.

You Might Also Like:

Best AI Video Tools 2026

Best AI Video Generators

By: Steven Jones On:
Updated on: September 30, 2026
Google Flow with Veo 3.1 is the best active AI video generator in the current DIY AI 2026 dataset, scoring…
Best AI Image-to-Video Generators in 2026

Best Image To Video AI

By: Steven Jones On:
Updated on: September 11, 2026
The best AI image-to-video tools turn a still image into believable motion without losing the subject, product, face, or composition…
Steven Jones

Writer: Steven Jones

AI Tools Reviewer and Technical Analyst

Steven Jones is a technology analyst specialising in artificial intelligence, machine learning workflows, and emerging automation tools.

At DIY AI, he focuses on clear, practical guidance for people comparing AI tools in the real world. His work covers text generation, image generation, video tools, data platforms, developer-focused AI products, and the automation workflows that connect them.

Steven's reviews are built around hands-on testing, practical benchmarks, and transparent scoring rather than vendor claims. He looks closely at where each tool performs well, where it falls short, and what those trade-offs mean for creators, teams, and businesses trying to make sensible AI adoption decisions.

He has a particular interest in safety, reliability, output quality, performance metrics, and dataset quality. When he is not reviewing the latest AI model updates, he experiments with prompt engineering techniques and contributes to DIY AI ongoing work on fair, explainable scoring frameworks for AI tools.

Contact

Leave a Comment On: AI Video Prompt Examples

Your email address will not be published.