One line in, a script out: what the idea-to-scenes generator does

Typing one line and getting a shot list back is genuinely useful. It is also the step where people accept the machine's first draft, which is the mistake.
"Two avocados argue about which one has more healthy fats." That is a video idea, and it is about as much as anyone wants to type. What it is not is a script: no shots, no order, no idea who says what, no sense of how long any of it runs.
The gap between those two things is mechanical enough to automate, and that is what the idea-to-scenes step does. You give it the line, it comes back with scenes.
What it actually produces
Not prose. A structured set of scenes, each one a shot with its own description, its own dialogue tagged by speaker, and a length that follows the dialogue rather than a guess. Which means what comes back is immediately editable in the same place you would have written it by hand - not a wall of text you have to cut up first.

It also assigns speakers to the characters you cast. That is why casting comes first: given a cast, the generator writes for those characters by name, and the names are what bind lines to reference images at render time. Given no cast, it writes for nobody in particular and you have to do that work afterwards.
Three things to fix every time
The first draft is a draft. In practice the same three things are worth a pass:
- Trim the lines. Generated dialogue runs long - it explains where a real joke lands. Cut it to the shortest version that still parses.
- Check the beat order. A generated script often puts the best line in the middle. Reorder scenes so it lands last.
- Delete a scene. Almost every three-scene draft is a two-scene video with an establishing shot nobody needs.
None of that is a criticism of the generator; it is what a first draft is. The value is that you are editing instead of staring at an empty box.
Where a one-line idea is not enough
The quality of what comes back tracks the specificity of what goes in, and one word is not specific. "Funny video about cats" produces exactly the video that prompt deserves. The line that works is a situation with a tension in it: two parties, one disagreement, one place.
- Weak: "office humour".
- Better: "a capybara CEO announces quarterly results to a room of very quiet interns".
- Weak: "cooking video".
- Better: "a strawberry dad discovers his baby is a banana, at breakfast".
The pattern in the good ones: a who, a where, and something that has just gone wrong. Everything else - the shots, the reactions, the ending - is what the generator is for.
Then write the still, not the clip
Once the script reads well, generate a still preview for the scenes you are least sure about. It is the cheapest way to find out whether the shot description means to the model what it means to you, and it doubles as the reference the video model animates. Approve a frame, then pay for motion.