vezo0 cr
← ALL ENGINES

// ENGINE 05 — KUAISHOU

Kling 3

Fast, cinematic video with audio & lip-sync. Talking characters — lip-sync, expressive faces, and action that lands, up to 15 seconds.

CREDITS
60 cr
QUALITY
great
SPEED
fast
DURATIONS
5 / 10 / 15s
ASPECTS
9:16 · 16:9 · 1:1
IMAGE INPUT
first frame
01

GENERATE

kling-3 — render session
|
Aspect
Duration
02

PROMPT SAMPLES

PROMPT_01

A street vendor grins at the camera and says "best noodles in the city — come find out," steam rising from the wok behind him, night market bustle.

PROMPT_02

A young woman in a denim jacket talks directly to camera in her kitchen, mid-laugh, morning light through the window — casual vlog energy.

PROMPT_03

A breakdancer hits a freeze in the middle of a subway platform crowd, camera whip-pans to catch the landing, onlookers cheer.

PROMPT_04

Close-up of a barista sliding a latte across the counter, winking: "on the house." Café chatter and clinking cups behind.

PROMPT_05

A gym coach points at the lens between sets and says "no excuses today," then racks the bar — plates clanging, room tone alive.

PROMPT_06

Two friends split-screen a phone; one gasps "you have to see this," the other leans in — vertical 9:16, reaction-cam framing.

03

ABOUT

Kling 3 is Kuaishou's video engine, built by a company that serves one of the largest short-video audiences on earth — and its strengths map directly onto that world: characters who speak with convincing lip-sync, expressive faces, and energetic motion that stays sharp instead of smearing.

The lip-sync is the headline. Put a spoken line in quotes in your prompt and the character delivers it — mouth shapes, timing, and the small facial movements around speech that make a delivery read as real. No other engine in the studio attempts dialogue this directly, which makes Kling the default for spokesperson-style creative.

With audio generated in-shot and durations up to 15 seconds, it carries a complete spoken beat — a testimonial, a punchline, a hype moment — in one render. Text-to-video frames in 9:16, 16:9, or 1:1, and its fast renders keep the iteration loop short when you're workshopping a line read. For realism without dialogue there is Veo 3.1; for silent emotional close-ups, Hailuo 2.3.

04

STRENGTHS

  • Convincing lip-sync: quoted dialogue in the prompt is actually delivered on camera
  • Expressive faces that track the emotional read of the scene
  • Energetic motion — dance, sport, action — that stays crisp instead of smearing
  • Up to 15 seconds: a full spoken beat or story slot in one render
  • Vertical-first 9:16 framing built for short-form feeds
05

PROMPTING TIPS

  1. TIP_01Put dialogue in quotation marks — "best noodles in the city" — and Kling will lip-sync the line rather than paraphrase it.
  2. TIP_02Describe the delivery, not just the words: grinning, deadpan, mid-laugh. The face follows the adverb.
  3. TIP_03Keep spoken lines under ~15 words for a 10-second shot so the delivery doesn't rush.
  4. TIP_04For UGC-style ads, name the setting and camera style — "talking to camera in her kitchen, casual vlog energy" — the vernacular matters.
06

USE CASES

01

UGC-style ads

A character looks into the lens and says the line. Kling's lip-sync makes spokesperson-style creative possible without a shoot.

02

Character dialogue

Two-shots and reaction beats where faces do the storytelling — expressions track the emotional read of the prompt.

03

Action and sports

Fast, dynamic movement — parkour, dance, a skate trick — rendered with the crispness short-form feeds reward.

04

Vertical-first stories

Built for 9:16: fifteen seconds is a full story slot, not a clip of one.