Wan 3.0 in motion: a rainy alley, a street stall and a rooftop dawn — with native sound
·15 s·Wan 3.0·Cast: Mia
The newest Wan generation is now a model option in Actoria: text-to-video and image-to-video with native audio, single takes up to 30 seconds, up to 10 reference images per scene, 720p and 1080p — and it's the cheapest sound-capable model in the selector.
How this video was made
Generated in Actoria from a three-paragraph script: each paragraph became a five-second scene, the synthetic actors were created once and kept their faces in every shot, and the scenes were stitched into one film.
Scenes
- 0s A rain-lit neon alley at night: one long tracking shot beside Mia, rain hiss and distant thunder in the native audio track.
- 5s A street food stall under a striped awning: slow push-in past paper lanterns as Mia warms her hands on a paper cup, rain drumming above.
- 10s Rooftop at dawn: the camera orbits Mia from silhouette into morning light while a tram bell rings somewhere below.
Try Wan 3.0 on your next scene
Open the model selector, pick “Wan 3.0 · with audio” and give a long take room to breathe.


