Video → prompt: any clip becomes a shot list

·5 min read·Guides·by Actoria

Every AI filmmaker has a folder of clips they wish they could prompt like. Actoria's **From video** feature closes that gap: upload a clip, and it comes back as a timestamped shot list — who does what at which second, where the camera goes, when the line lands. It costs 5 tokens and takes about a minute. To test it honestly, we fed it a 45-second scene everyone knows from a classic action film — and then ran the extracted prompt, completely raw, on three different models. The Wan 3.0 take is above; the other two, the full prompt and the exact bill are below.

Generated in Actoria · Wan 3.0 · 25 sOpen the video page

How From video works

1
Upload a clip
Any video up to 60 seconds / 30 MB — a movie scene, an ad you admire, your own draft footage.
2
Get the shot list
The feature watches it and writes a private template: timestamped beats, numbered characters (Hero 1, Hero 2…), camera moves as text. 5 tokens, refunded if parsing fails.
3
Run it — raw or edited
The template opens ready to generate: cast your own actors into the Hero slots, tweak beats, or launch it exactly as extracted.

The important design choice: it extracts structure, not pixels. No faces are cloned, no footage is reused — you get the direction of the scene (its rhythm, blocking and camera grammar) as editable text, and your own cast performs it.

The prompt it extracted

This is the untouched output for our 45-second test clip — nine beats, two numbered characters, camera moves included. We didn't rewrite a word before generating:

TimeExtracted direction
0–2Hero 1 speaks to Hero 2, asking about Sarah Connor. Hero 2 is writing on a form. A man in dark sunglasses and a black leather jacket stands at a counter, looking at a police officer. The police officer is seated behind the counter, writing on a form.
2–6Hero 2 looks up from his writing and speaks to Hero 1, refusing to let him see Sarah Connor. Hero 2 then resumes writing.
6–10Hero 1 asks Hero 2 where Sarah Connor is. Hero 2 looks up and speaks to Hero 1.
10–15Hero 2 tells Hero 1 that it may take a while and he can wait on a bench. Hero 1 turns his head slightly to the right.
15–20Hero 1 looks up, his face sweaty. He then looks down and to the right.
20–25Hero 1 looks directly at the camera and says, “I'll be back.”
25–32Hero 1 walks away from the counter, his back to the camera. He walks out of frame. The camera pans to the right, showing the outside of the building at night.
32–37Hero 2 is writing on a form with a pencil. The camera zooms in on his hand as he writes.
37–45Hero 2 looks at the camera. Suddenly, Hero 1 smashes through the counter, destroying it. Hero 2 is thrown back. Papers fly everywhere.
Yes — this is the same shot list behind our I'll Be Back experiment. That article's prompt started life right here: extracted by From video, then hand-trimmed to 30 seconds. Today we're running the raw 45-second original instead.

One prompt, three models

Same text, three engines, zero editing. Watch how each one interprets identical direction:

Wan 3.0 — one continuous 25-second take with native sound. The pause before the line and the pan to the street survive intact.
Seedance 2.5 — one 25-second take, the spoken line lands word for word; the warmest faces of the three.
MiniMax H3 — the budget render on our own GPU: a single 12-second take covering the core beats (its single-take ceiling is 15 seconds).

What it all cost

StepTakeRate (720p, sound)Tokens≈ USD*
From video → prompt45-s source clipflat5$0.07
Wan 3.0one 25-s take12 tok/s300$4.30
Seedance 2.5one 25-s take16 tok/s400$5.80
MiniMax H3one 12-s take8 tok/s96$1.40

*At the annual Basic rate. The token numbers are real charges from these exact generations, not list-price math. The whole experiment — prompt plus three films — came to 801 tokens, about $11.50, cheaper than a cinema ticket for a scene that once needed a police-station set and a stunt car.

What we learned running it raw

On ethics: From video reads a scene the way a film student does — structure, rhythm, camera. It doesn't copy footage, likenesses or dialogue tracks, and the source clip never leaves your private library. Study the classics; cast your own heroes.
Great scenes were always free to learn from. Now they're one upload away from being your storyboard.

Frequently asked questions

What clips can I upload?

Anything up to 60 seconds and 30 MB: film scenes, ads, TikToks, your own phone footage. The result is a private template in your library — visible only to you.

Does it clone the people in the video?

No. Characters come back as numbered role slots (Hero 1, Hero 2) with text descriptions. You cast them from your own actor library — synthetic characters, or verified real people with consent.

How much does the extraction cost?

5 tokens per clip, refunded automatically if the parse fails. Generation is then priced by whichever model you pick — the table above shows real numbers for a 25-second scene.

Can I edit the prompt before generating?

Yes — the template opens in the editor first: change beats, merge scenes, swap the setting, keep only the camera moves. Raw runs are for science; edited runs are for shipping.

Turn your favorite scene into a shot list
Upload a clip, get the prompt for 5 tokens, and re-shoot it with your own cast tonight.
Create your first video

More from the blog

All articles