Seed Audio 1.0
Sign up
Back to blog

Kling 4.0: Thirty-Second Clips, Stereo Sound, and a Flash Preview

What Kling announced on September 28, 2026: a 30-second native clip, stereo audio, and which controls are still marked coming soon.

Sep 29, 2026SeedAudio Editorial Team
Kling 4.0: Thirty-Second Clips, Stereo Sound, and a Flash Preview

Kling 4.0 is a video model that keeps the soundtrack inside the shot. On September 28, 2026, @Kling_ai posted on X that the full model is coming this October, and that Kling 4.0 Flash is already open to Ultra Yearly subscribers. That post names a month and an access group. It does not describe what the model can do.

The capability list below is the longer official introduction that IT Home reprinted on September 29. Phoenix and National Business Daily carried the same brief that morning. A control counts as available only when the studio you are using actually exposes it. Two items in that introduction are explicitly marked as not live yet.

This note was checked on September 29, 2026.

What you can open today

Two names are in the announcement, and they are not the same release.

  • Kling 4.0 is scheduled for October 2026. No day is named. Chinese business coverage describes the current stage as an internal test.
  • Kling 4.0 Flash opened on September 28 to a small group. English write-ups of the X post call that group Ultra Yearly subscribers. The Chinese introduction calls it 黑金年卡, the black-gold yearly plan. The official line for Flash is speed and cost on high-frequency jobs. It does not say Flash includes the 30-second length, the 15 references, or 10-bit HDR.

Two features in the same introduction are labeled coming soon even for the full model: 4K and 1080p 10-bit HDR output, and repeat extensions out to two minutes.

What the official brief lists

ControlWhat Kling’s introduction saysStatus in that text
Length of one passUp to 30 seconds, with a long take and a continuous sceneListed for Kling 4.0
ExtensionRepeat extensions, up to 2 minutesComing soon
References

Up to 15 multimodal items in one job: as many as 10 images, 5 video clips, and 7 subjects

Listed for Kling 4.0
KeyframesUp to 10, for character state, scene changes, and story beatsListed for Kling 4.0
Picture4K and 1080p 10-bit HDR, plus a 21:9 frameHDR marked coming soon; 21:9 is listed with it
SoundHigh-quality audio, dual-channel stereo, tighter lip syncListed for Kling 4.0
PromptUp to 8,000 tokens, with the model allowed to expand the briefListed for Kling 4.0
FlashFaster, cheaper passes for high-frequency workOpen September 28 to yearly black-gold members

The reference math is easy to misread. Fifteen is the total for one job. Ten images, five clips, and seven subjects are ceilings inside that total. They do not add up to a 22-item upload.

Motion is described as an upgrade to fast movement, continuous action, follow shots, and orbits. Editing can add, change, or remove a subject or the background, and can change style, weather, color, material, and shot size. A separate “creative replicate” pass is supposed to read a reference film’s camera language and story structure, then make a new commercial clip for social, product, fashion, or electronics ads. Named looks are cinematic, animation, fantasy, pixel, and felt.

The same note says the creation page now takes image, video, and audio references in one box, switches between a grid and a list, keeps preview, edit, and new material on one timeline, and adds a canvas where an agent sits beside the nodes.

Sound stays inside the picture

The audio claims are about the clip, not a session.

Kling says the model supports high-quality audio and dual-channel stereo, so the voice has more depth and sits closer to production sound. Lip sync between voice, dialogue, and performance is supposed to be tighter. Languages named in that section are Chinese, English, Japanese, Korean, Spanish, Portuguese, German, French, and Hindi. Accents and dialects named beside them are Beijing, Taiwan, Northeast China, Sichuan, Cantonese, American English, British English, and Indian English. The same paragraph says multilingual text, emoji, and logos should stay readable while the shot moves.

That is a picture with a voice in it. Dialogue, room, and music arrive as one bounce. You cannot drop the rain under a whisper, or replace one market’s line, from that file. When the job is the soundtrack, write the scene in the Seed Audio generator and leave the picture model for the picture.

What one early test reported

APPSO published a hands-on on September 29. Treat it as a review, not as a second spec.

They describe stable follow, pan, and high-angle moves in a large action scene; a bamboo duel and a rooftop fight that kept the same faces through fast cuts; a 30-second dramatic scene; and a Cantonese family dinner where a father’s line matched the mouth and the stereo split the room from the street. They also placed 10 keyframes on a road-movie shot and used a 21:9 frame.

Two of their claims sit ahead of the reprinted introduction. They treat 10-bit HDR as footage you can already grade in DaVinci Resolve or Premiere. IT Home’s reprint still marks 4K and 1080p 10-bit HDR as coming soon. They also mention extra reference slots and extra looks, including 3D cyber, vlog, and documentary, that the reprint does not list. When the review and the reprint disagree, use the reprint.

Their note on references is still practical. Say which clip owns camera and performance, and which still owns the face, the clothes, and the room. Otherwise the model imports props you did not ask for.

A brief you can write before the public launch

Length: 30 seconds
Aspect: 21:9
Language: Cantonese
Audio: stereo; room on the left, street on the right; lip sync on
Keyframes:
- 00:00 wide, cold window light, dinner table
- 00:08 father looks down and starts the line
- 00:18 mother stops her chopsticks
References:
- still A: father's face and clothes
- still B: mother's face and clothes
- still C: the room
Do not borrow faces or furniture from any reference video.

If a Flash preview comes back, listen for three failures: a voice that misses the mouth, stereo that is just a doubled mono track, and a face that changes between keyframes. Those are the claims the October model is supposed to hold. A vague “cinematic dinner” will not tell you whether it did.

What is still unpublished

  • No calendar day inside October.
  • No price, and no statement that credits per clip stay the same.
  • No public API model id for 4.0 or 4.0 Flash in the announcement. A September 28 read of Kling’s API price table, reported by CellCog, still listed 3.0-era models only. Confirm that on Kling’s own pricing page before you wire a job to a model name.
  • No published matrix of what Flash can do versus the October model.

Sources