Diagnosis
| Problem | Evidence in your prompt | Confidence |
|---|---|---|
| Three mixed styles | "cyberpunk watercolor oil painting" names three media. The model blends them. | Likely |
| Never night and day at once | "at night" and "bright sunny day" contradict each other. The model picks one or averages them. | Likely |
| Messy and cluttered | "minimalist but very intricate" contradicts itself. Three held objects plus a sign plus a forest is a lot for one image. | Likely |
| Too many fingers | The girl holds a sword, a bouquet and a lantern, which needs several complex hand grips. Many generators handle hands badly. | Likely |
| Gibberish sign text | Many generators render long text unreliably, and a 5-word all-caps phrase is a hard case. | Likely |
| "no bad hands" does nothing | Many tools read "hands" as something to draw and ignore the "no". It can even draw more attention to hands. | Possible |
| Filler words add nothing | "beautiful stunning masterpiece 8k ultra detailed" is abstract and gives the model nothing to draw. | Possible |
Minimal fix
Changes are in bold.
beautiful stunning masterpiece 8k ultra detailedgirl in a red dress standing in a forest at night,bright sunny day,holding a lantern only,a sword and a bouquet of flowers andwatercolor painting style,cyberpunk watercolor oil painting style, minimalist but very intricate,a wooden sign reading "WELCOME"that says 'WELCOME TO THE ENCHANTED WOODS', no bad hands
Clean rewrite
A girl in a red dress standing on a forest path at night, holding a glowing lantern in one hand, the other arm relaxed at her side. Full-body shot, centered, with tall dark trees behind her and a small wooden sign reading "WELCOME" to her left. Warm lantern light on her face and dress, deep blue moonlit shadows. Watercolor painting, soft washes, simple background.
If your tool has a negative prompt field, put this there: extra fingers, deformed hands, blurry text, mixed styles.
Variants to test
- Night and day together. Keep the rest, but change the light to "dusk, last orange sunset glow on the horizon, deep blue sky above, lantern lit". Dusk is the closest a single image gets to night and day at once. Look for whether the light now looks coherent.
- Keep the sword and flowers. Use "holding a lantern in her left hand, a sword sheathed at her belt", and drop the bouquet. Look for whether the hands stay correct when one hand is holding the lantern and the other is empty.
For a cyberpunk look, make that a separate test. Use "oil painting, neon pink and teal rim light" in place of watercolor. Don't combine it with the other styles.
Test plan
- Change one thing at a time and generate 3–4 images per version with the same settings.
- If hands are still wrong, crop or hide them: "hands behind her back", or a waist-up shot with the lantern at the edge of the frame.
- For the sign, either shorten it to one word (as above) or leave it blank ("a blank wooden sign") and add the text in an image editor afterward. This is the most reliable method.
What I couldn't tell
- Which tool you're using. This matters for several things:
- Midjourney and Stable Diffusion use negative prompts differently (
--noand a negative field). Tools like DALL·E or ChatGPT image generation usually have no negative field, so describe what you want instead. - Newer models handle text and hands much better than older ones, so your results may improve just by switching.
- Midjourney and Stable Diffusion use negative prompts differently (
- What your images look like. If you share one, I can tell you more precisely what went wrong.
- Check your tool's docs for how it handles negative prompts, text rendering and aspect ratio. A tall ratio like 2:3 suits a full-body standing figure.
Real run in an isolated folder with only this skill installed. Only the Skill and Read tools were enabled, so no command was run and no file was written. The scenario is fictional. No image was generated: the model only wrote text prompts, so how well they work depends on the image tool you use. The model could not see any generated image, so its diagnosis is a set of hypotheses to test.
它做什么
像调试程序一样工作:先收集完整的提示词、所用工具和哪里不对,再对照一张常见问题表来读提示词(部分内容被忽略、结果平淡、画面杂乱、风格混搭、自相矛盾、构图不对、光线平、颜色串色、手部变形、文字乱码、多出不想要的元素、每次结果不一样)。对每个问题,它会指出提示词里的证据和把握程度,然后给出标出改动的最小修改、一版干净的重写,以及两个各只测试一件事的变体。
还涵盖
测试计划(一次只改一个变量、每个版本生成多张、在相同设置下对比)、什么时候是工具或模型本身的限制,以及因为看不到图片或工具的行为而无法判断的部分。
适合什么场景
“不好使”的提示词、总缺某个细节的图,以及想弄清哪些词真正起作用。
低风险:纯指令文件,没有脚本,不联网、不写文件。它依据提示词和你对结果的描述来判断;你不分享图片的话它看不到图,也不知道你的工具如何给词加权,所以每个诊断都是需要测试的假设。它不会帮人绕过工具的安全过滤或使用条款、冒用真实人物来欺骗,或制作涉及真人或未成年人的色情内容。关于反向提示词和文字渲染的行为,请查看所用工具的文档。AIBars 原创(MIT)。已用一条虚构的混乱提示词试用过一次;没有生成图片。