Skip to main content
Prompts8 min read

AI Video Prompts: 24 Examples and the Formula Behind Them

A prompt library you can copy straight into the app, plus the six ingredients that separate a prompt that works from one that produces a slideshow.

A good AI video prompt names six things: the subject, the action, the setting, the camera, the lighting, and the visual style. Written in that order, in one or two sentences of about 30 to 60 words, it gives the model everything it needs to decide what moves and how the shot is framed. Vague prompts produce still-looking video because the model has no action to animate and no camera instruction to follow. The 24 prompts below are written to that formula and can be pasted directly into VIBE on iOS or Android.

What makes a good AI video prompt

The most common mistake is describing a picture instead of a shot. "A cat on a windowsill" is an image prompt: a video model given that has to invent both the motion and the camera, and it will choose almost none of either. That is why so many first attempts come back looking like a photo with a slight drift. A video prompt needs six ingredients β€” not all six every time, but the more you supply, the less the model has to guess.

  • Subject β€” who or what the shot is about, described concretely. "An elderly fisherman in a yellow oilskin" beats "a man".
  • Action β€” what actually moves. This is the ingredient people skip, and the one that decides whether the clip moves at all.
  • Setting β€” where it happens. Weather, time of day and surface materials all give the model something to render.
  • Camera β€” how the frame behaves: slow dolly push-in, orbit, handheld tracking shot, static locked-off, crane up, low angle. Subject motion and camera motion are separate instructions.
  • Lighting β€” the biggest lever on realism. Golden hour, hard rim light from behind, soft diffused overcast, single tungsten desk lamp, neon ambient glow.
  • Style β€” the treatment that ties it together: cinematic, documentary, 90s cel animation, macro product commercial, claymation, film grain.

The formula

Fill in the blanks and you have a working prompt. Keep it to roughly 30–60 words β€” long enough to be specific, short enough that no single instruction gets buried.

Reusable template

[Subject, described concretely] [doing a specific action] in [setting with weather and time of day], [camera movement and angle], [lighting description], [visual style].

Worked through once: "a barefoot child" + "chasing a paper kite along the shoreline" + "a wide empty beach at low tide, late afternoon" + "camera tracking alongside at knee height" + "warm low sun backlighting the spray" + "16mm film grain, nostalgic summer palette". Joined up, that is one sentence the model can execute without inventing anything important.

Cinematic and realistic prompts

These lean on lighting language and a single unbroken camera move. They work well on Kling 3 Pro, Google Veo 3.1, Sora 2 and Luma Ray 3.2.

Dawn trawler

A lone fisherman in a yellow oilskin hauling wet nets over the rail of a small trawler at dawn, breath visible in the cold air, slow handheld tracking shot following him from behind, flat blue pre-sunrise light with hard specular highlights on the wet rope, muted teal grade, 35mm documentary look.

Rain-soaked square

A woman in a red wool coat crossing an empty cobbled square in heavy rain, umbrella tilted against the wind, slow low-angle tracking shot at knee height, streetlights reflected in the wet stone, deep shadows and hard highlights, neo-noir colour grade with crushed blacks.

Watchmaker

An elderly watchmaker leaning over a cluttered workbench, tweezers lifting a tiny gear into place, dust motes drifting through the beam of a single brass desk lamp, slow push-in ending on a close-up of his hands, warm tungsten key light against a near-black background, shallow depth of field.

Mountain road

A vintage cafΓ©-racer motorcycle leaning hard through a mountain switchback at golden hour, dust kicking off the rear tyre, camera tracking alongside from a parallel road on a long lens, warm backlight flaring across the frame, compressed perspective, cinematic anamorphic look.

Product and ad prompts

Name the surface, the light and the camera move β€” those three decide whether the result looks like a commercial or a phone snap. Google Veo 3.1 with audio switched off is a good default here, because its reference images keep the product recognisable across a set of variations.

Tech reveal

A matte black wireless earbud case opening slowly on a polished concrete surface, lid rotating to reveal the earbuds inside, camera orbiting ninety degrees around it at desk height, soft overhead key light with one hard rim light from the right, shallow depth of field, premium tech commercial aesthetic.

Food macro

A bar of dark chocolate snapping in half in extreme slow motion, cocoa dust scattering into a hard shaft of side light, deep black background, macro lens locked off, high-contrast studio product lighting, luxury confectionery commercial.

Skincare on stone

A frosted glass serum bottle standing on wet river stones, a single drop falling from the pipette and rippling the shallow water beneath, slow push-in from a low angle, cool overcast morning daylight, fine mist drifting in the background, clean natural beauty-brand aesthetic.

Sports energy

A pair of running shoes on a rain-slicked athletics track, camera rising from ground level to reveal the finish line beyond, water spray lifting off the surface, cold blue stadium floodlights with hard shadows, high-contrast grade, energetic sportswear commercial.

Nature and animal prompts

Wildlife prompts benefit from a stated lens and a documentary reference, because those carry a whole set of framing conventions with them. Weather is the cheapest way to add motion to a landscape β€” wind, rain and drifting cloud give the model something to animate when the subject itself is still.

Snow leopard

A snow leopard picking its way along a rocky ridge in falling snow, tail counterbalancing each step, breath steaming in the cold, slow telephoto tracking shot from across the valley, flat overcast grey light, natural colour, wildlife documentary style.

Hummingbird macro

A hummingbird hovering at a red trumpet flower in extreme slow motion, wings blurred into a disc, tongue extending into the bloom, macro lens locked off, warm morning sun backlighting the translucent petals, heavily blurred green garden bokeh behind.

Storm over wheat

Dark storm clouds building over a wide wheat field, wind rolling visible waves through the crop, static wide shot at eye level, late afternoon sunlight breaking through a gap in the cloud base and sweeping across the field, natural colour, patient landscape cinematography.

Dolphins at the bow

A pod of dolphins riding the bow wave of a boat, filmed from just above the waterline, bodies breaking the surface and diving again, spray catching the light, bright midday sun refracting through the blue water, handheld feel, natural documentary grade.

Anime and animation prompts

Name the medium explicitly, and be specific: "90s cel animation with visible line work" gives the model far more to hold onto than "anime style". Models vary a lot here, so this is the category where running the same prompt on two of them pays off most.

Summer hillside

An anime schoolgirl running down a narrow hillside path at sunset, uniform ribbon and hair streaming behind her, tall grass brushing past the frame, camera tracking alongside her, warm orange sky with heavy cloud detail, 90s cel animation look with visible line work and hand-painted backgrounds.

Felt stop-motion

A stop-motion felt fox waking up and stretching in a hollow log, moss and tiny mushrooms around the entrance, handmade fabric texture visible in every surface, slow dolly in from outside the log, soft diffused daylight, whimsical children’s picture-book palette.

Neon courier

A cyberpunk courier on a hoverbike weaving between neon towers in the rain, jacket snapping in the airflow, camera chasing directly behind at speed with heavy motion lines, magenta and cyan signage lighting reflecting off wet surfaces, modern anime style, high contrast.

Claymation kitchen

A clay-animated chef flipping a pancake in a tiny cluttered kitchen, thumbprints and tool marks visible in the clay, pans and jars crowding the shelves behind, locked-off camera at counter height, warm practical light from a small window, playful claymation style with a slightly jerky frame rate.

Social and viral prompts

Short-form has its own grammar: vertical frame, the first second doing all the work, and a camera that is either locked off overhead or right in the action. Set 9:16 in the aspect-ratio picker rather than only asking for it in the prompt β€” the picker is what controls the output frame.

Overhead unboxing

A pair of hands unboxing a glossy black mystery package on a bright white desk, tape peeling back and the lid lifting away, camera locked off directly overhead, crisp even shadowless lighting, fast confident movements, satisfying tactile ASMR aesthetic, vertical framing.

Night street food

A street-food vendor tossing noodles in a wok over an open flame at a night market, fire flaring up into the frame with each toss, camera close and handheld just above the pan, warm orange firelight against a dark busy street, vertical framing, high-energy food content.

Outfit transition

A person standing in a plain grey tracksuit in front of a concrete wall, a fast whip-pan blurs the frame sideways, and they emerge in a full black evening outfit in the same position, identical background and lighting throughout, single-take transition, vertical framing.

Miniature cooking

Tiny hands assembling a miniature cheeseburger on a doll-sized grill, patty sizzling, cheese slowly melting over the edge, extreme macro lens locked off at grill height, warm kitchen light from the left, shallow depth of field, vertical framing, oddly satisfying miniature-food style.

Image-to-video prompts describe motion, not the scene

This is where most people go wrong. When you upload a photo, the model already knows what the scene looks like β€” composition, subject, lighting and colour are all fixed by the image, and describing them again wastes the prompt. What the model does not know is what should move. So an image-to-video prompt is a set of motion instructions: what the subject does, what the environment does, and what the camera does. "Animate this image" produces almost nothing, because it contains none of the three.

Portrait β€” subtle life

Slow dolly push-in toward the subject, her hair drifting in a light breeze, one slow blink and a small smile forming, shoulders rising with a breath, background gently falling further out of focus as the camera moves.

Product β€” controlled orbit

The camera orbits the product 180 degrees clockwise at a fixed height while the product itself stays perfectly still, the highlight travelling slowly across its surface and the reflection shifting as the camera moves.

Still life β€” ambient motion

Steam begins rising from the cup and curling toward the top of the frame, the surface of the liquid ripples once, sunlight shifts slowly across the table from left to right, camera holds completely static throughout.

Landscape β€” reveal

The camera cranes up and pulls back to reveal the valley beyond the subject, mist drifting through the low ground, cloud shadows sweeping across the hillside, the subject staying in the lower third of the frame throughout.

Two instructions, not one

Subject motion and camera motion are independent, and models treat them as separate problems. "She turns her head toward the camera" and "slow dolly push-in" produce very different results from either one alone. Combining them deliberately is what makes a clip look directed rather than generated β€” and if a clip feels static, one of the two is usually missing.

What to leave out of a prompt

Cutting these out is usually a bigger improvement than adding anything.

  • Negatives. Writing "no cars, no people" often puts cars and people in the shot, because the model reads the nouns and not the "no". PixVerse 6 and Google Veo 3 Fast have a dedicated negative-prompt field; everywhere else, just describe what you do want.
  • Duration and aspect ratio as words. "10 second video, 9:16" in the prompt does nothing. Both are pickers in the app, and the picker is what the model receives.
  • Several unrelated actions. Clips are short. One action completes convincingly; three compete for the same few seconds and none of them finish.
  • On-screen text and logos. AI video models render legible text unreliably. Add titles, captions and branding in an editor afterwards, where they will be sharp and spelled correctly.
  • Real people by name, and anything that violates the content policy β€” prompts naming public figures, or asking for sexual or graphic content, are blocked before generation.
  • Contradictory camera directions. "Static locked-off shot with a sweeping crane move" leaves the model to pick one, and you will not know which until it comes back.
  • Padding adjectives. "Beautiful, amazing, stunning, high quality, 8k, masterpiece" adds nothing a video model can act on. Specific nouns and verbs do the work superlatives only pretend to.

Do different models need different prompts?

Yes, in three specific ways, and knowing them saves a lot of wasted generations.

Audio is the first. Some models always generate a soundtrack β€” Sora 2, Seedance 2.5, Flux 3 and LTX 2.5 Fast among them β€” so direct the sound rather than let the model choose it: write the line of dialogue, name the ambience, describe the effect. Google Veo 3.1 and Kling 3 Pro make audio optional. Others produce none at all, including Luma Ray 3.2 and Kling v2.5 Turbo Pro, and on those, sound words in the prompt do nothing.

Length is the second. Google Veo 3.1 generates 4, 6 or 8 second clips, Sora 2 generates 4, 8 or 12, Kling 3 Pro runs any whole second from 3 to 15, and Seedance 2.5 reaches 30. A prompt with a beginning, a middle and an end will not fit into 4 seconds, so match the ambition of the prompt to the duration you picked.

Structure is the third. Kling 3 Pro accepts a story of up to six shots, each with its own prompt, so there a shot list is a legitimate prompt format; on every other model it gets compressed into one confused shot. A few models change the rules entirely β€” Grok Imagine 1.5 is image-to-video only and treats the prompt as optional motion guidance, and the Talking Avatar model has no prompt field at all.

A checklist before you hit generate

  1. 1Is there a specific action, expressed as a verb? This is the most common cause of a still-looking clip.
  2. 2Is there a camera instruction? Add one even if it is "static locked-off shot", so the model is not guessing.
  3. 3Is the lighting described? Time of day, direction and quality of light do more for realism than any other clause.
  4. 4Is the style named at the end? One clear reference beats three competing ones.
  5. 5Did you set the aspect ratio and duration in the pickers rather than in the text?
  6. 6Draft on a free model first β€” Seedance Pro Fast, LTX 2 Distilled, PRUNA V or LTX 2.5 Fast β€” fix the prompt, then finish on a flagship.

AI video prompts β€” frequently asked questions

How do I write a good AI video prompt?

Name six things in one or two sentences: the subject, the action it performs, the setting, the camera movement, the lighting, and the visual style. Roughly 30 to 60 words is the useful range. The action and the camera movement matter most β€” a prompt without either gives the model nothing to animate, which is why it comes back looking like a still photo.

Why is my AI video not moving?

Almost always because the prompt describes a scene rather than an event. Add an explicit action verb ("steam rises", "she turns toward the camera", "the leaves shift in the wind") and a separate camera instruction ("slow dolly push-in", "orbit around the subject"). This matters most in image-to-video, where the photo has already supplied the scene and the prompt exists only to describe motion.

How long should an AI video prompt be?

About 30 to 60 words for most shots. Shorter than that and you are leaving decisions to the model; much longer and individual instructions start getting diluted, especially in a clip of only a few seconds. The exception is Kling 3 Pro, which accepts a structured story of up to six shots with a separate prompt for each.

Do different AI models need different prompts?

Yes, in three ways. Models that always generate audio, such as Sora 2 and Seedance 2.5, should have the dialogue and ambience written into the prompt; models with no audio ignore sound words entirely. Duration limits differ β€” 8 seconds on Google Veo 3.1, 12 on Sora 2, 15 on Kling 3 Pro, 30 on Seedance 2.5 β€” so the prompt has to fit the clip. And Kling 3 Pro is the only one that accepts a multi-shot shot list as a prompt format.

Should I use negative prompts?

Only where the model has a dedicated field for them. In VIBE, PixVerse 6 and Google Veo 3 Fast accept a negative prompt. On every other model, writing "no cars" in the main prompt tends to add cars, because the model reads the noun and not the negation. Describe what you want in frame instead.

Can I use the same prompt on more than one model?

Yes, and it is the fastest way to learn which model suits your subject. Because VIBE includes 35+ AI video models in one app on a shared token balance, you can run one prompt through two or three models and compare the output directly. Expect to adjust the audio instructions and the duration between models, but the core description usually transfers.

What is the best AI video prompt for TikTok or Reels?

A single clear action, a close or overhead camera, and vertical framing set in the aspect-ratio picker. Short-form clips have about one second to establish what is happening, so front-load the visual hook and leave the setup out. The social prompts in this article are written to that pattern.

Do I need a prompt for every model?

No. The Talking Avatar model has no prompt field at all β€” you supply a photo and a voice recording and the audio drives the whole video. Grok Imagine 1.5 is image-to-video only and treats the prompt as optional guidance for the motion. Every other video model in the app takes a text prompt, and several also accept an image alongside it.

Copy a prompt, generate it in 30 seconds

Download VIBE free on iOS and Android. Paste any prompt from this page, pick a model, and generate β€” with free models included and 35+ AI video models in one app.

Download VIBE AI Video Generator on the App StoreGet VIBE AI Video Generator on Google Play