Build a visual prompt from subject, composition, lighting, style and exclusions. Free to use in your browser; no sign-up.
An image prompt is a description of a photograph or illustration that does not exist yet, written for a model that has no idea what you pictured. The reliable structure is subject, then action, then setting, then framing and lens, then light, then style. Vague prompts do not produce surprising images; they produce average ones.
Without a key this tool composes a draft from proven formats, filled with what you typed. That runs entirely in your browser and costs nothing. Add a key and the same inputs go to Claude instead, which writes rather than assembles.
Your key is stored only in this browser and is never sent to us, so we cannot see it. Requests go straight from your browser to Anthropic. It is masked here and obfuscated in storage, though anyone with developer tools on this machine can still read their own key back.
Keys come from console.anthropic.com. Usage on your own key is billed to you by Anthropic, not by us.
Take Ashwood Ceramics, a two person pottery studio selling mugs and serving bowls at markets and through its own site, needing a wide hero image for a page about a new stoneware range. There is no budget for a photographer this month and the pots are not glazed yet, so the image has to be generated, and it has to look like the studio rather than like a catalogue.
The prompt: "A pair of hands lifting a freshly thrown stoneware bowl from a potter's wheel, wet clay still visible on the fingers, in a small working studio with wooden shelves of unglazed pots behind. Three quarter view from slightly above, close shot, 50mm lens, shallow depth of field with the shelves softly out of focus. A single large window to the left giving soft overcast daylight, gentle shadows, no artificial light. Muted earth tones, warm greys and oatmeal. Documentary photography, natural and unstyled, fine grain. No text, no logos, no faces, no bright colours, no glossy studio lighting." Aspect ratio set to 16:9 at generation time rather than cropped afterwards.
It works because almost every clause changes the picture. Hands and a bowl on a wheel is a subject doing something, which gives the model a scene instead of a category. The framing block, three quarter view, slightly above, close, 50mm, shallow depth of field, decides the composition before any adjective gets a say, and it is the part most people leave out entirely. The lighting clause names a single source with a direction, which is why the result has consistent shadows rather than the ambient flatness a prompt without lighting produces. The palette is given as three specific colour names rather than as a mood. Faces are excluded deliberately, since hands avoid the uncanny results faces often produce and sidestep any question about depicting a person who does not exist. And the exclusions clear out what models add by default: invented text, invented logos, and the glossy commercial lighting that would make a working studio look like a showroom.
The prompt the studio wrote first: "beautiful professional photo of a modern ceramics studio, artisan handmade pottery, cozy aesthetic, cinematic, highly detailed, 8k, trending, award winning, masterpiece"
Nothing in that describes a shot. Professional, cinematic and award winning are quality adjectives, and quality adjectives push a model towards the most generic version of an appealing image rather than towards a particular one. There is no camera position, so the model picks a wide establishing view every time. There is no light source, so the scene comes back evenly lit and lifeless. Cozy and modern pull in opposite directions and cinematic pulls against both. There are no exclusions, so the shelves fill with invented signage and the wheel acquires a decorative plant. The result is competent, anonymous and immediately recognisable to any reader as a generated image, which is the one thing the hero of a handmade goods page cannot afford to be.
If the image needs to show your actual product, premises, team or work, generate nothing. A stoneware range that exists has to be photographed, because a customer comparing the hero image with the mug that arrives will notice, and a generated approximation of a real product edges into misrepresentation. The same applies to premises, to before and after work, and to anything a buyer would treat as evidence. A phone camera by a window will serve you better than the best prompt you could write.
If legible text has to appear in the image, plan on adding it afterwards. Text rendering in generated images remains unreliable, particularly for anything longer than a couple of words, and a nearly correct logo or a misspelled shop sign is more damaging than no text at all. Generate the background as a plain composition with deliberate empty space, then set the type in a design tool where you control the font, the spelling and the contrast. That also makes the asset reusable when the wording changes.
If people are recognisable, or the style is, stop and check. Prompting for a real person's likeness, or for the manner of a living named artist, raises questions about rights and permissions that vary by platform and jurisdiction and are not settled by the fact that the tool complied. Client work often carries its own constraints as well: some brands have photography guidelines that generated imagery will not satisfy, and some sectors restrict imagery outright, particularly where a picture could imply a result. Ask before generating rather than after.
And when the shot is genuinely ordinary, licensed stock is often the faster answer. A generic desk, a road, a plate of food or a city skyline exists in thousands of properly shot versions, and finding one takes less time than iterating a prompt towards something merely acceptable. Generation earns its place when the specific scene you want does not exist in stock, or when you need a series of images that share a consistent look, which is the case a stock library handles worst.
Most models weight earlier terms more heavily than later ones, so the sequence of a prompt is a statement of priorities rather than a list. The order that behaves predictably is subject, then what the subject is doing, then the setting, then framing and lens, then lighting, then style, then exclusions. Putting a style reference first tends to produce an image that is mostly style with a vague subject somewhere inside it, which is a common reason a prompt comes back atmospheric and useless.
Length is a diminishing return rather than a virtue. Adding clauses helps while each new clause changes something visible, and starts to hurt once you are adding words that only sound impressive, because attention spreads across everything you wrote and the terms that matter get diluted. A practical test before adding a phrase: name what would look different if you removed it. If you cannot, cut it. Quality adjectives such as beautiful, professional, stunning and masterpiece almost never survive that test, since they describe how you want to feel about the image rather than anything the image contains.
Photographic vocabulary is the most reliable set of controls available, because it maps onto real distinctions the model has seen labelled thousands of times. Shot distance comes first: extreme close up, close up, medium shot, wide, establishing. Then angle: eye level, low angle looking up, high angle looking down, overhead flat lay. Then focal length, which changes the geometry rather than just the zoom, with wide focal lengths exaggerating depth and long ones compressing it and separating the subject from its background. Then depth of field, stated as shallow or deep rather than as an aperture number you do not need.
Lighting is the second lever and the one most often missing entirely. Name a source, a direction and a quality: soft overcast daylight, hard midday sun from behind, a single window to the left, warm evening backlight, a practical lamp visible in frame. That one clause is usually the difference between a flat image and one with shape. Colour behaves the same way. Three named colours will steer a palette; a mood word such as vibrant or moody will not, because it has meant something different in every image ever described that way.
These three describe the same picture to three different readers, and confusing them wastes effort. A creative brief is written for a person and carries the things a model cannot use: why the image exists, where it will sit, what the brand does not do, who signs it off, and the deadline. A prompt is written for a model and should contain only what is visible in the frame, which is why briefs pasted into a prompt box produce poor results: purpose, audience and rationale are not visual facts, and the model tries to render them anyway.
Alt text runs the other way. It is written for a screen reader user and for anyone whose image failed to load, so it describes the meaning of the image in one plain sentence rather than cataloguing its composition. Reusing a prompt as alt text produces something absurdly long and full of lens specifications no reader needs. Write the alt text separately, from what the image finally shows, keep it to a sentence, and describe the point rather than the technique. If the image is purely decorative, an empty alt attribute is the correct choice rather than a description nobody wanted.
Change one thing at a time. The instinct after a disappointing result is to rewrite the whole prompt, which discards whatever was working and makes the next result equally hard to explain. Adjust the framing clause alone, generate, then adjust the lighting alone, and keep the versions you did not choose, because a prompt you liked three iterations ago is often the one you want back. Where the tool exposes a seed, holding it fixed while you change a single clause is the closest thing to a controlled comparison available, and varying the seed with the prompt fixed shows how much of the result was chance.
Set a limit in advance, something like ten attempts or twenty minutes, and treat reaching it as information rather than failure. Some images do not come out of a prompt because the scene is too specific, involves accurate text, needs a real object, or depends on a relationship between several elements the model keeps rearranging. At that point the answer is a photograph, a stock licence, an illustrator, or a simpler composition. Finish by checking the output rather than the prompt: count fingers and limbs, read any incidental text, look at reflections and at hands holding objects, and view the image at the size it will really be seen.
Common questions about image prompt generator output, answered without the sales pitch.
Work in order: subject, action, setting, framing and lens, lighting, style, and exclusions. Concrete visual nouns do more than adjectives, and framing plus lighting are the two elements most often missing.
Build a structured prompt with context, constraints, audience and output format. Free to use in your browser; no sign-up.
Open tool →Prompts & creativeStructure a video prompt around shot, action, camera movement, setting and duration. Free to use in your browser; no sign-up.
Open tool →No pitch deck, no retainer talk on the first call. Bring the problem: we'll tell you honestly whether we're the right people for it.
We’ll review it, then arrange a founder call if it is a fit. Or call +91 95182 76146: Monday to Saturday, 10:00 - 18:00 IST (UTC+5:30)