Diagnosis
| Problem | Evidence in your prompt | Confidence |
|---|---|---|
| Three mixed styles | "cyberpunk watercolor oil painting" names three media. The model blends them. | Likely |
| Never night and day at once | "at night" and "bright sunny day" contradict each other. The model picks one or averages them. | Likely |
| Messy and cluttered | "minimalist but very intricate" contradicts itself. Three held objects plus a sign plus a forest is a lot for one image. | Likely |
| Too many fingers | The girl holds a sword, a bouquet and a lantern, which needs several complex hand grips. Many generators handle hands badly. | Likely |
| Gibberish sign text | Many generators render long text unreliably, and a 5-word all-caps phrase is a hard case. | Likely |
| "no bad hands" does nothing | Many tools read "hands" as something to draw and ignore the "no". It can even draw more attention to hands. | Possible |
| Filler words add nothing | "beautiful stunning masterpiece 8k ultra detailed" is abstract and gives the model nothing to draw. | Possible |
Minimal fix
Changes are in bold.
beautiful stunning masterpiece 8k ultra detailedgirl in a red dress standing in a forest at night,bright sunny day,holding a lantern only,a sword and a bouquet of flowers andwatercolor painting style,cyberpunk watercolor oil painting style, minimalist but very intricate,a wooden sign reading "WELCOME"that says 'WELCOME TO THE ENCHANTED WOODS', no bad hands
Clean rewrite
A girl in a red dress standing on a forest path at night, holding a glowing lantern in one hand, the other arm relaxed at her side. Full-body shot, centered, with tall dark trees behind her and a small wooden sign reading "WELCOME" to her left. Warm lantern light on her face and dress, deep blue moonlit shadows. Watercolor painting, soft washes, simple background.
If your tool has a negative prompt field, put this there: extra fingers, deformed hands, blurry text, mixed styles.
Variants to test
- Night and day together. Keep the rest, but change the light to "dusk, last orange sunset glow on the horizon, deep blue sky above, lantern lit". Dusk is the closest a single image gets to night and day at once. Look for whether the light now looks coherent.
- Keep the sword and flowers. Use "holding a lantern in her left hand, a sword sheathed at her belt", and drop the bouquet. Look for whether the hands stay correct when one hand is holding the lantern and the other is empty.
For a cyberpunk look, make that a separate test. Use "oil painting, neon pink and teal rim light" in place of watercolor. Don't combine it with the other styles.
Test plan
- Change one thing at a time and generate 3–4 images per version with the same settings.
- If hands are still wrong, crop or hide them: "hands behind her back", or a waist-up shot with the lantern at the edge of the frame.
- For the sign, either shorten it to one word (as above) or leave it blank ("a blank wooden sign") and add the text in an image editor afterward. This is the most reliable method.
What I couldn't tell
- Which tool you're using. This matters for several things:
- Midjourney and Stable Diffusion use negative prompts differently (
--noand a negative field). Tools like DALL·E or ChatGPT image generation usually have no negative field, so describe what you want instead. - Newer models handle text and hands much better than older ones, so your results may improve just by switching.
- Midjourney and Stable Diffusion use negative prompts differently (
- What your images look like. If you share one, I can tell you more precisely what went wrong.
- Check your tool's docs for how it handles negative prompts, text rendering and aspect ratio. A tall ratio like 2:3 suits a full-body standing figure.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. No image was generated: the model only wrote text prompts, so how well they work depends on the image tool you use. The model could not see any generated image, so its diagnosis is a set of hypotheses to test.
What it does
Works like a debugger: collect the exact prompt, the tool and what went wrong, then read the prompt against a table of common problems (ignored parts, bland results, clutter, mixed styles, contradictions, wrong framing, flat light, color leaking, distorted hands, garbled text, unwanted elements, run-to-run differences). For each problem it names the evidence in the prompt and how sure it is, then shows a minimal fix with the changes marked, a clean rewrite and two variants that each test one thing.
Also covers
A test plan (change one variable, generate several images per version, compare under the same settings), advice on when the tool or model is the limit, and what it could not tell because it cannot see the image or the tool's behavior.
Good for
A prompt that "does not work", pictures that keep missing a detail, and learning which words matter.
Low risk: pure instructions with no scripts, no network access and no file writes. It works from the prompt and your description of the result; unless you share the image it cannot see it, and it cannot know how your tool weighs words, so every diagnosis is a hypothesis to test. It will not help bypass a tool's safety filters or terms, imitate a real person to deceive, or produce sexual content involving real people or minors. Check the tool's documentation for its negative-prompt and text-rendering behavior. Original skill by AIBars (MIT). Tried once on an invented messy prompt; no image was generated.