Wan 3.0 online — sound, long takes, your cast

Wan 3.0 is the newest generation of the Wan video family — and you don't need a GPU or a ComfyUI graph to run it. In Actoria it works right in the browser: text-to-video and image-to-video with native audio, single takes up to 30 seconds, and up to 10 reference images per scene. Plug it into the same studio you already know: consistent actors, storyboard frames, multi-scene films.

Everything on this page runs on Actoria.ai — the AI video generator where your cast never breaks character.

What's new in Wan 3.0?

Three things stand out in practice. Native sound: scenes come back with ambience and effects already in the track — rain on an awning, a tram bell, room tone — no separate foley pass. Long takes: a single generated shot can run 5, 10, 20 or 30 seconds, so slow camera moves finally have room to breathe. And references: up to 10 images can steer one scene, which is how faces, wardrobe and props stay recognizable from shot to shot.

In Actoria, Wan 3.0 sits in the model selector next to the Seedance family: pick it per scene, per project or per template. At 720p it is also the cheapest sound-capable model in the lineup — 12 tokens per second versus 16 for Seedance 2.5.

Wan 3.0 without the setup

The Wan family made its name as the open model you could run yourself — if you had the VRAM, the patience and a weekend. Wan 3.0 in Actoria is the other path: open the site, describe a scene, get a clip with sound in minutes. No checkpoints to download, no nodes to wire, no queue on your own GPU, and it works the same on a laptop and a phone.

References are the superpower

Up to 10 reference images per scene change how you direct. Your actors' portraits ride along as identity anchors, so the same synthetic character keeps their face across every Wan 3.0 shot; extra references can pin down a location, a costume or a key prop. Combined with Actoria's storyboard-first flow — approve the frame, then animate it — you get control that a bare prompt never gives.

Real people are welcome too — under the same rule as everywhere in Actoria: a real face appears only after its owner verifies it with a live camera check, and the consent covers Wan's provider explicitly. Your verified self, invited friends, synthetic actors and constructor characters all share the same scenes.

How it works

01
Pick Wan 3.0
In the scene's model selector choose “Wan 3.0 · with audio”.
02
Set the frame
Compose the first frame with your actors, or go straight from text.
03
Generate with sound
Pick 5–30 seconds and render; the clip returns with its own audio track.

Why creators pick Actoria

Native audio
Ambience and effects arrive inside the clip — no separate sound pass.
Takes up to 30 s
5, 10, 20 or 30 seconds per shot — room for real camera moves.
10 references per scene
Faces, wardrobe and props stay consistent shot after shot.
720p and 1080p
12 tokens/s at 720p — the cheapest sound-capable model here; 1080p when it matters.

Frequently asked questions

What is Wan 3.0?

The next generation of the Wan video model family, following the widely used Wan 2.2: text-to-video and image-to-video with native audio, longer takes and multi-image reference control. In Actoria it runs in the cloud — no local install.

How much does Wan 3.0 cost in Actoria?

12 tokens per second at 720p and 24 at 1080p — a 5-second 720p scene with sound is 60 tokens. On current plans that is roughly $0.86–1.35 per 5-second clip, cheaper per second than Seedance 2.5.

Can I use my verified face on Wan 3.0?

Yes. Real-person actors work on Wan 3.0 under Actoria's standard rule: the person verifies their own face with a live camera check first, and the consent explicitly covers Wan's provider. After that your face — or an invited friend's — rides into scenes as an identity reference, same as a synthetic actor's portrait.

Wan 3.0 or Seedance 2.5 — which one?

Try both — the selector is per scene. Wan 3.0 wins on price per second, take length and multi-reference control; Seedance 2.5 is our default for exact on-screen text and real-person actors. Many films mix them scene by scene.

Does Wan 3.0 here differ from self-hosted Wan?

You get the model without the ops: no VRAM requirements, no graphs, no updates to chase. Around it you get what a bare model doesn't have — persistent actors, storyboard approval before you pay for motion, scene-by-scene projects and a final assembled film.

Keep exploring

Actoria.ai — create actors once, keep the same faces in every scene.