Why does the same prompt look great on one AI video model and terrible on another?
Because AI video models are not interchangeable. Each one was trained and tuned for a different job: some hold a camera move steady, some nail lip-sync, some are simply cheap enough to iterate on. Sending every shot to your favourite model is the single biggest reason AI video output feels inconsistent.
I ran the same three briefs (a product orbit, a talking-head line, and a moody city b-roll) through five models over the last fortnight. The quality gap between them was smaller than the routing gap. Picking the right model per shot beat picking the "best" model overall, every time.
Model routing is the practice of assigning each shot in a video to the model best suited to that shot type, instead of generating everything on one platform. It is a workflow decision, not a subscription decision, and it typically cuts cost and reshoot count at the same time.
Which AI video model is best for each type of shot?
Match the shot to the model's strength. Use a lip-sync-native model for dialogue, a high-control model for camera moves, a cheap fast model for background b-roll, and a leaderboard-topping model when a single hero shot has to carry the whole piece. No model in July 2026 wins all four categories.
Here is the routing table I now work from:
--- Dialogue and talking-head: a model with native lip-sync. Seedance 2.5 is the current pick because the audio is generated with the mouth movement rather than dubbed over it. Dubbed lip-sync reads as fake within two seconds on a phone screen.
--- Camera moves and directed shots: Runway Gen-4.5. It responds to structured prompting (dolly, orbit, rack focus) more predictably than the leaderboard leaders, and it slots into an editing workflow rather than a chat window.
--- Background b-roll and filler: the cheapest tier you can tolerate. Veo 3.1 Lite and Runway Gen-4 Turbo both sit around US$0.05 per second of output. Nobody pauses on your third cutaway.
--- Hero shot or thumbnail moment: Gemini Omni Flash. As of late July 2026 it leads the LMArena text-to-video leaderboard at 1527 Elo, roughly 45 points clear of second place, and it supports conversational editing so you can refine rather than reroll.
--- Long or 4K sequences: Seedance 2.5 again, which added native 4K and 30-second outputs plus a 3D pre-visualisation step for planning camera movement before the final render.
The models change every few weeks. The routing categories do not. Build the habit around the categories.
How much does AI video actually cost per finished minute?
Far more than the per-second sticker price, because you never use the first generation. At a realistic 8:1 generate-to-keep ratio, one finished minute means roughly 480 seconds of output. At US$0.10 per second that is about US$48, or roughly HK$375. At US$0.05 per second it drops to about HK$187.
Published list prices as of late July 2026, for output seconds:
--- Veo 3.1 Lite: about US$0.05 per second. Runway Gen-4 Turbo: about US$0.05 per second (5 credits per second).
--- Kling: about US$0.07 per second, with Kling 3.0 nearer US$0.10.
--- Sora 2 base: about US$0.10 per second at 720p. Sora 2 Pro: about US$0.30 to US$0.50, reaching US$0.70 at 1080p Pro.
--- Gemini Omni Flash: about US$0.10 per second (US$17.50 per 1M video output tokens), halving to roughly US$0.05 through the Batch API.
--- Runway Gen-4.5: 12 credits per second, about US$0.12. Veo 3.1 Standard: about US$0.40 per second at 1080p.
Run the arithmetic before you pick a subscription. A 60-second product video shot entirely on Veo 3.1 Standard costs roughly HK$1,498 in generation alone at that same 8:1 ratio. Routed properly, with one hero shot and the rest on a cheap tier, the same minute lands nearer HK$250.
Treat every figure here as a list price on a fast-moving market. Check the vendor's own pricing page on the day you commit budget.
What does a good AI video prompt look like?
A good prompt names five things in a fixed order: subject, action, camera, lighting, and format. Vague prompts produce vague motion. The models are not guessing your intent from vibes; they are matching your words to training captions, and caption language is technical.
Here is the template I use for every shot. Copy it as-is and fill the brackets.
Try this prompt:
SUBJECT: [one concrete subject, with material and colour, e.g. "a matte black ceramic coffee cup with steam"]
ACTION: [one physical action only, e.g. "steam rising slowly, liquid still"]
CAMERA: [one move only, e.g. "slow 90-degree orbit left to right, eye level, 50mm lens"]
LIGHTING: [source, direction, temperature, e.g. "single soft key light from camera left, warm 3200K, deep falloff to black"]
FORMAT: [duration, ratio, resolution, e.g. "6 seconds, 16:9, 1080p, no text overlay, no logo"]
NEGATIVE: [what must not appear, e.g. "no hands, no reflections of a studio, no camera shake"]
Two rules make this template work. One action, one camera move per clip. Asking for a pan and a zoom and a subject that turns around produces mush; generate two clips and cut between them. Never let lighting be implicit. Lighting is the variable that decides whether five separate clips look like one shoot.
If you already write structured prompts for text models, this is the same discipline applied to a different output. Our piece on context engineering covers why explicit structure outperforms clever phrasing across every model type.
How do you keep multiple AI clips looking like one video?
Lock three variables across every generation: lighting description, lens and height, and colour palette. Copy those three lines verbatim between prompts and change only the subject and action. Most "AI video looks disjointed" complaints trace back to rewriting the lighting line each time.
Beyond the prompt, use whatever the platform calls its consistency feature. Style reference, character lock, and seed reuse all do the same structural job: they constrain the model to a previous output instead of starting fresh. Runway and Kling both expose this; the chat-based tools often bury it.
Then finish in an editor, not in the generator. Grade all clips to a single LUT, add one grain layer across the whole timeline, and match audio. Fifteen minutes of grading hides more model inconsistency than fifty extra generations will.
The still-image side moved too, which matters for storyboards. Meta launched Muse Image on 7 July 2026, currently around third on text-to-image and second on image editing, and Google shipped Nano Banana 2 Lite. Generating your key frames as stills first, then animating the approved frames, is dramatically cheaper than iterating in video.
Where does AI video still fall apart in 2026?
Four places, reliably. Named products, text on screen, hands doing precise work, and continuity across more than about ten seconds. These are not prompt problems you can solve with better phrasing; they are current model limits, and pretending otherwise wastes budget.
--- Your actual product. No model has seen your SKU. A generated approximation of a real product you sell is worse than a phone-shot clip of the real thing. Film the product; generate the environment.
--- On-screen text. Legible text is still unreliable at small sizes. Add titles in your editor, not in the prompt.
--- Precise hand work. Fingers manipulating small objects remain the most common failure. Cut away or crop out.
--- Length. Most models still perform best on short clips. Even where 30-second outputs exist, build a sequence from multiple shots rather than betting on one long perfect generation.
Honest framing beats hype here. AI video is now genuinely good at atmosphere, motion and environment, and still weak at specificity. Plan the edit around that split and the output stops feeling like a gamble.
Try this now: the 20-minute routing test
Take one shot you actually need this week. Generate it three times using the template above: once on a cheap tier, once on Runway Gen-4.5, once on the current leaderboard leader. Same five lines, same lighting text, only the platform changes.
Then look at cost per usable clip, not cost per second. In my runs the cheap tier won on b-roll by a wide margin, and lost badly on anything with a directed camera move. That single test replaces months of subscription guesswork.
Log the result in a two-column note: shot type on the left, winning model on the right. That note is your routing table, and it is worth more than any leaderboard, because it reflects the shots you actually shoot. If you want a broader sense of how differently models behave under load, our guide to AI effort settings covers the same principle on the text side.
The takeaway
Stop shopping for the best AI video model. Start routing shots to the model that handles them. The routing table costs you one afternoon to build, survives every model release, and cuts a typical finished minute from four figures in Hong Kong dollars down to low hundreds.
Technology should reduce the number of decisions you make under pressure, not add to them. We understand AI. We understand you better. With UD by your side, AI doesn't feel cold.
Reviewed by the UD AI team.
🎬 Turn the Technique Into a Workflow
You have the routing table. The next step is building it into a production process that runs every week without you rebuilding it from scratch. We'll walk you through every step, from tool selection and prompt libraries to review gates and delivery.