Video → prompt: any clip becomes a shot list
Every AI filmmaker has a folder of clips they wish they could prompt like. Actoria's **From video** feature closes that gap: upload a clip, and it comes back as a timestamped shot list — who does what at which second, where the camera goes, when the line lands. It costs 5 tokens and takes about a minute. To test it honestly, we fed it a 45-second scene everyone knows from a classic action film — and then ran the extracted prompt, completely raw, on three different models. The Wan 3.0 take is above; the other two, the full prompt and the exact bill are below.
How From video works
The important design choice: it extracts structure, not pixels. No faces are cloned, no footage is reused — you get the direction of the scene (its rhythm, blocking and camera grammar) as editable text, and your own cast performs it.
The prompt it extracted
This is the untouched output for our 45-second test clip — nine beats, two numbered characters, camera moves included. We didn't rewrite a word before generating:
| Time | Extracted direction |
|---|---|
| 0–2 | Hero 1 speaks to Hero 2, asking about Sarah Connor. Hero 2 is writing on a form. A man in dark sunglasses and a black leather jacket stands at a counter, looking at a police officer. The police officer is seated behind the counter, writing on a form. |
| 2–6 | Hero 2 looks up from his writing and speaks to Hero 1, refusing to let him see Sarah Connor. Hero 2 then resumes writing. |
| 6–10 | Hero 1 asks Hero 2 where Sarah Connor is. Hero 2 looks up and speaks to Hero 1. |
| 10–15 | Hero 2 tells Hero 1 that it may take a while and he can wait on a bench. Hero 1 turns his head slightly to the right. |
| 15–20 | Hero 1 looks up, his face sweaty. He then looks down and to the right. |
| 20–25 | Hero 1 looks directly at the camera and says, “I'll be back.” |
| 25–32 | Hero 1 walks away from the counter, his back to the camera. He walks out of frame. The camera pans to the right, showing the outside of the building at night. |
| 32–37 | Hero 2 is writing on a form with a pencil. The camera zooms in on his hand as he writes. |
| 37–45 | Hero 2 looks at the camera. Suddenly, Hero 1 smashes through the counter, destroying it. Hero 2 is thrown back. Papers fly everywhere. |
One prompt, three models
Same text, three engines, zero editing. Watch how each one interprets identical direction:
What it all cost
| Step | Take | Rate (720p, sound) | Tokens | ≈ USD* |
|---|---|---|---|---|
| From video → prompt | 45-s source clip | flat | 5 | $0.07 |
| Wan 3.0 | one 25-s take | 12 tok/s | 300 | $4.30 |
| Seedance 2.5 | one 25-s take | 16 tok/s | 400 | $5.80 |
| MiniMax H3 | one 12-s take | 8 tok/s | 96 | $1.40 |
*At the annual Basic rate. The token numbers are real charges from these exact generations, not list-price math. The whole experiment — prompt plus three films — came to 801 tokens, about $11.50, cheaper than a cinema ticket for a scene that once needed a police-station set and a stunt car.
What we learned running it raw
- Timestamps transfer. All three models kept the beat order, and the long models (Wan 3.0, Seedance 2.5) held the 4-second bureaucratic pause that makes the famous line work.
- Numbered heroes prevent identity swaps — nine beats of “Hero 1 / Hero 2” and nobody traded jackets mid-scene.
- Know your model's take ceiling. Wan 3.0 and Seedance 2.5 run up to 30 s in one shot; H3 tops out at 15 s, so a long shot list means either trimming to the core beats (what we did) or splitting into takes on a quiet beat.
- Raw is a starting point. The extraction is faithful, sometimes too faithful — trimming beats 15–20 (the sweaty pause) made the 30-second version in our previous article punchier.
Great scenes were always free to learn from. Now they're one upload away from being your storyboard.
Frequently asked questions
What clips can I upload?
Anything up to 60 seconds and 30 MB: film scenes, ads, TikToks, your own phone footage. The result is a private template in your library — visible only to you.
Does it clone the people in the video?
No. Characters come back as numbered role slots (Hero 1, Hero 2) with text descriptions. You cast them from your own actor library — synthetic characters, or verified real people with consent.
How much does the extraction cost?
5 tokens per clip, refunded automatically if the parse fails. Generation is then priced by whichever model you pick — the table above shows real numbers for a 25-second scene.
Can I edit the prompt before generating?
Yes — the template opens in the editor first: change beats, merge scenes, swap the setting, keep only the camera moves. Raw runs are for science; edited runs are for shipping.