Wan 3.0 is live in Actoria: sound, long takes, references
We added Wan 3.0 to the Actoria model selector. It brings three things our other models don't combine in one place: native audio on every clip, single takes up to 30 seconds, and up to 10 reference images steering a scene. The film above was generated entirely on Wan 3.0 with our synthetic actress Mia — sound included, no separate foley pass. Here's what the model changes in practice and when to pick it.
What Wan 3.0 adds
Native sound. Every clip returns with its own audio track — rain, room tone, street noise, the small effects that sell a shot. In the film above, the rain hiss in the alley and the tram bell on the rooftop came straight out of the model; we didn't add a single sound in post.
Long takes. A single generated shot can run 5, 10, 20 or 30 seconds. That's the difference between a slideshow of moments and an actual camera move: a tracking shot that settles, an orbit that completes, a push-in that lands exactly when the action does.
References. Up to 10 images can steer one scene. Actoria uses them the way a set photographer would: your actors' portraits ride along as identity anchors, and extra slots can pin a location, a costume or a prop. This is what keeps a character recognizable across every shot — the same idea behind why AI characters drift, now with more anchors per scene.
Where it sits in the lineup
| Model | 720p, with sound | Max take | References per scene | Real-person actors |
|---|---|---|---|---|
| Wan 3.0 | 12 tokens/s | 30 s | up to 10 | yes, with consent |
| Seedance 2.5 | 16 tokens/s | 30 s | up to 4 | yes, with consent |
| Seedance Lite | 8 tokens/s | 30 s | — | yes, with consent |
At 720p Wan 3.0 is now the cheapest sound-capable model in the selector after Lite — and the cheapest with multi-reference control, full stop. A 5-second 720p scene is 60 tokens; the whole three-scene film in this post cost 198 tokens including storyboard frames and assembly. 1080p is available at 24 tokens per second when a shot deserves it.
How the launch film was made
The pipeline is the same one we use for every film on this blog — pick a model, approve frames, generate, assemble. The only change was the selector. That's the point of a multi-model studio: when a new model lands, your actors, scripts and workflow come along unchanged.
When to pick Wan 3.0
- Atmosphere-heavy scenes where ambient sound does half the work — weather, streets, interiors.
- Slow, deliberate camera moves that need 10–30 seconds to complete.
- Multi-reference scenes: one character plus a fixed location or prop that must stay exact.
- Budget drafts with sound — at 12 tokens/s it undercuts Seedance 2.5 by a quarter.
New model, same cast: Mia walked out of a Seedance film and into a Wan 3.0 one without losing her face.
Frequently asked questions
Is Wan 3.0 available to everyone?
Yes — it's in the model selector for every account, per scene, per project and in templates. Pricing is the standard token grid: 12 tokens per second at 720p, 24 at 1080p.
Can I use my verified face on it?
Yes. Real-person actors work on Wan 3.0 under the same rule as everywhere in Actoria: the person verifies their own face first, and the consent they give explicitly covers Wan's provider. Verified faces ride into scenes as identity references — same mechanics as synthetic actors.
Does the sound replace Actoria's sound layer?
No — it complements it. Wan 3.0 brings ambience and effects natively; the separate sound layer is still there when you want a designed soundscape on any model's footage.