Structure a video prompt around shot, action, camera movement, setting and duration. Free to use in your browser; no sign-up.
A video prompt is a shot description. Unlike a still image prompt, it has to specify movement: what the subject does, what the camera does, and over what duration. Prompts that describe only a scene tend to return either a static shot or unpredictable motion, because nothing in the request said what should change.
Without a key this tool composes a draft from proven formats, filled with what you typed. That runs entirely in your browser and costs nothing. Add a key and the same inputs go to Claude instead, which writes rather than assembles.
Your key is stored only in this browser and is never sent to us, so we cannot see it. Requests go straight from your browser to Anthropic. It is masked here and obfuscated in storage, though anyone with developer tools on this machine can still read their own key back.
Keys come from console.anthropic.com. Usage on your own key is billed to you by Anthropic, not by us.
Ferndale Cold Brew is a small roastery that wants a six second opening shot for a product video: a bottle on a counter, morning light, something that feels handmade rather than corporate. The temptation is to describe the mood. The prompt instead describes a shot, which is what the model can actually render.
The prompt reads: 'Medium close-up, 50mm, waist height. A dark glass bottle of cold brew coffee stands on a scratched oak counter in a small cafe. Condensation beads on the glass and one drop runs down the side. Behind it, out of focus, a person in an apron moves left to right across frame. Camera: slow push in, roughly ten centimetres over the shot, tripod smooth, no handheld shake. Light: hard morning sun from camera left through a window, long shadow falling to the right, warm white balance. Six seconds. Shallow depth of field, filmic, slight grain. No text, no logo.'
Every clause in that prompt is doing a job. The lens and the height fix the framing so the bottle sits where it was intended rather than centred by default. Subject motion and camera motion are stated separately, which stops the two collapsing into an unrequested drift: the condensation and the person in the background move, the camera does one slow push and nothing else. The light is given a direction and not just a time of day, because direction is what makes the next shot cut together with this one. The duration is short enough to stay coherent. And the two exclusions at the end deal with the failure everybody meets, which is invented lettering appearing on the label.
The negative case is worth naming precisely. A weaker prompt would be: 'A beautiful cinematic shot of our cold brew coffee in a cosy cafe, morning vibes, showing the craft that goes into every bottle, ending with the logo.' It has no frame, so the model chooses one. It has no camera instruction, so the shot may drift, orbit or sit still, and it will differ on every generation. 'Morning vibes' gives no light direction, which means the second shot will not match this one. 'The craft that goes into every bottle' asks for a narrative that cannot be rendered in one shot, so the output averages several ideas into a scene that looks like coffee advertising and shows nothing specific. And requesting a logo asks the model to draw text, which is the one thing it is least reliable at.
The working version is longer and duller to read, which is the correct trade. A shot prompt is a set of instructions for a camera operator who cannot ask you questions, and everything you leave out gets decided by something other than you.
Do not generate video where accuracy about your own business matters. Your premises, your staff, your actual product on its actual shelf: a generated approximation of those is a small lie that customers notice, particularly local ones who have been inside the building. A phone on a tripod for ten minutes produces a truer shot and usually a faster one. Generated footage earns its place for the abstract, the illustrative and the impossible to film, not for the things sitting in front of you.
Skip it when text has to be legible. Prices, dates, phone numbers, package copy, anything a viewer will read and act on: generated frames render lettering unreliably, and the label on a bottle will come back as convincing nonsense. Shoot the plate, or generate the background and add the text in an editor afterwards where you control it. The same caution applies to hands, logos, brand colours and any recognisable third-party product.
Hold off entirely if there is no shot list. Prompting is a per-shot activity, and without a list you are generating fragments and hoping they assemble. Write the sequence first, even roughly: what the viewer sees, in what order, for how long, and what the cut between each pair of shots is doing. That list will tell you which shots you can film, which you can licence, and which are genuinely worth generating, which is usually fewer than expected.
Then there are the situations where the constraint is not craft. Regulated claims, before and after imagery in health or beauty, anything showing an outcome a customer might expect: generated footage is a poor place to be imprecise, and some platforms require disclosure of synthetic media. Check the rules for where the video will run before you spend the credits rather than after. If a shot would need a model release when filmed, generating a person to avoid that conversation does not remove the underlying question of what the footage implies.
A reliable prompt reads in roughly the order a camera department would work: framing, subject, action, setting, camera, light, duration, style. Framing means shot size and lens, such as wide, medium or close-up, plus a focal length, since a 24mm and an 85mm produce entirely different relationships between subject and background. Subject and action are the what and the movement, kept to one continuous thing. Setting places it and fixes the era and texture of the room.
Camera is the field most often missed, and it needs its own sentence: static tripod, slow push in, handheld follow, tilt up, orbit left. Light wants a direction and a quality, hard or soft, from where, at what colour temperature, because that is what carries across shots. State the duration even where the tool has its own setting, and put style terms last and keep them short. Long style stacks pull the frame away from everything specified before them. Finally, exclusions: a brief list of what should not appear, most often text, logos and extra people.
Generated clips hold together for a short window and then start to degrade: hands change, backgrounds reorganise, a subject's clothing shifts. The practical convention is to plan in shots of a few seconds and to cut, which is also how the sequence would be shot with a real camera. A finished thirty second piece is comfortably six to ten shots, and building it that way gives you a rhythm rather than one long take that slowly falls apart.
Write the cut into the plan rather than discovering it in the edit. A push in cutting to a static shot reads as arrival. Two moving shots cut together fight each other unless the movement continues in the same direction. Because the model has no memory between generations, continuity is your job: the same light direction, the same lens language, the same time of day stated identically in each prompt. Generate several takes of each shot, since the variance between generations of one prompt is high, and choose in the edit. Budget for that in time and in credits, because the first version of any shot is rarely the one you use.
An image prompt describes a state. A video prompt describes a change, and everything difficult follows from that. In a still, the composition is the whole result, so weight goes into subject, framing, light and style. In a moving shot, two independent things are changing at once, what happens in front of the lens and what the lens itself is doing, and a prompt that does not separate them leaves the model to decide. That is the single most common cause of shots that look almost right and cannot be cut with anything else.
Time introduces a second difference: pacing. A still has no wrong duration, while a shot has a length that either fits the edit or does not. And a third: consistency between generations. One image can stand alone, but a video is a sequence, so the prompt has to carry repeated elements that hold the shots together. In practice this means keeping a small block of fixed text, covering the lighting, the lens, the grade and the era, that you paste into every prompt in the set, and varying only the subject, action and camera lines beneath it.
Treat them as rushes rather than as finished output. Bring everything into an editor, watch each generation at full size before deciding, and look specifically at the places these models fail: hands, faces at the edge of frame, any straight line in the background, and the last second of the clip. Trimming the final half second often rescues a shot that otherwise breaks down at the end.
Then unify what the generations did differently. A light grade across the whole sequence will pull mismatched clips closer together than any amount of reprompting, and a consistent grain or a subtle stabilisation pass does similar work. Sound matters more than most people expect, because generated footage arrives silent and silence is what makes it feel synthetic; room tone and a couple of specific effects do a great deal. Keep the prompts alongside the clips in a plain file so a shot can be regenerated later with the same settings, and check the disclosure requirements of the platform you are publishing to before it goes out.
Common questions about video prompt generator output, answered without the sales pitch.
Describe one shot: subject, the action, the setting, camera movement, framing, lighting and duration. Keep subject motion and camera motion as separate statements.
Build a structured prompt with context, constraints, audience and output format. Free to use in your browser; no sign-up.
Open tool →Prompts & creativeBuild a visual prompt from subject, composition, lighting, style and exclusions. Free to use in your browser; no sign-up.
Open tool →No pitch deck, no retainer talk on the first call. Bring the problem: we'll tell you honestly whether we're the right people for it.
We’ll review it, then arrange a founder call if it is a fit. Or call +91 95182 76146: Monday to Saturday, 10:00 - 18:00 IST (UTC+5:30)