When a batch fails, the fastest emotional reaction is to blame the model. “It doesn’t understand materials.” “It can’t hold composition.” “It ignores negative space.” Sometimes those statements are true. More often the prompt is carrying a soft instruction, a competing signal, or a missing constraint that the model is simply obeying. I used to spend hours changing models and parameters before I admitted the language itself was the problem.
I now run a short debugging checklist before I switch tools, rewrite the entire prompt, or decide the model is incapable. The checklist is deliberately mechanical. It forces me to test the structure before I abandon it. This is the exact sequence I use on almost every commercial job when the output is not matching the brief.
Why Blaming the Model Is Usually Premature
Models are inconsistent, but they are rarely random. When the same failure repeats across a batch, there is almost always a corresponding weakness in the prompt: a relative term instead of a rule, a style layer competing with a structural requirement, an exclusion that arrives too late, or a product description that leaves room for interpretation. Changing models without fixing those weaknesses simply moves the failure to a new interface.
The checklist exists to catch those weaknesses in order of likelihood and cost. It is faster to tighten one block than to restart the entire job in a different tool.

The Checklist
I run these questions in order. I do not skip ahead. I do not generate a large new batch until the relevant step is resolved.
1. Is the product description specific enough to prevent invention?
I read the product block out loud. Does it name the exact form, material, finish, and critical proportions? Or does it leave gaps the model has to fill?
Vague: “premium ceramic bottle”
Specific: “matte ceramic bottle with subtle texture, precise cylindrical form, no gloss, no plastic sheen”
If the product is drifting, I tighten this block first and run a four-image test before touching anything else. Most geometry and material failures start here.
2. Are the camera and framing instructions written as rules or as suggestions?
I look for relative language: “slightly above,” “three-quarter view,” “off-center,” “clean composition.”
I replace each one with a concrete constraint: height landmark on the product, directional angle, explicit placement zone.
Then I generate a small grid and check whether the product sits at a consistent height and position.
If the camera is still drifting, the language is still soft. I do not proceed until the grid holds.
3. Is the negative-space (or any hard spatial) requirement stated as a non-negotiable constraint?
“Room for text” and “open composition” fail regularly.
“Clear empty space occupying the right third of the frame, no objects, no shadows, no texture” fails far less often.
I confirm the spatial rule sits in the composition block, not buried after style language.
I test it under the actual crops the client will use. If the space collapses under crop, the constraint is still too weak or the product placement is too aggressive.
4. Is any style or mood language competing with the structural blocks?
I count the atmospheric adjectives and aesthetic references. If there are more than two or three in the first working prompt, I strip them out and re-test the foundation alone.
Style language is added only after the product, camera, and space are stable.
When a previously working structure suddenly fails after I added “cinematic soft light” or “luxury minimalism,” the new language is the prime suspect. I remove it and verify.
5. Are the exclusions short, early enough, and limited to high-priority failures?
Long negative lists can fight the positive instructions. I keep exclusions to the failures I have already seen on this type of job and place them after the composition block.
If a recurring unwanted element is still appearing, I move its exclusion earlier or make it more specific rather than adding five new negatives.
6. Am I changing more than one variable at a time?
When a batch is close but not right, the urge is to adjust material, light, camera, and style in a single new prompt. That makes diagnosis impossible.
I change only the block that corresponds to the current failure, generate a small diagnostic batch, and score only that failure.
If the failure is reduced, I lock the change. If it is not, I tighten further or re-examine the translation of the brief. Multi-variable changes are postponed until the single-variable tests are clean.
7. Have I actually scored the batch against the brief, or only against my taste?
I re-read the original commercial requirements: product fidelity, consistency across the set, usable space for the planned placements, material response that will not trigger client notes.
I look at the rejected frames only through those criteria. Attractive frames that fail the brief are still failures.
If the batch is failing a criterion I never stated clearly in the prompt, the prompt is incomplete. I add the missing constraint and re-test.
How the Checklist Changes the Workflow
Running the seven questions usually takes ten to twenty minutes. It almost always surfaces a soft instruction or a competing signal that I can fix without abandoning the model or the overall structure. The most common fixes are:
Tightening the product material line with a clear exclusion
Replacing relative camera language with a height landmark and placement zone
Restating negative space as a hard fraction plus “no objects”
Removing early style pressure that was overriding structure
Only after the checklist is clear do I consider changing models or major parameters. In the majority of commercial cases the model was capable; the prompt was not yet precise enough.
A Real Debugging Sequence
A five-image set of a brushed metal bottle kept failing client review for “inconsistent finish and unreliable text space.” My first reaction was to blame the model’s material handling. Instead I ran the checklist.
Product description was long and contained competing terms. I shortened it to “brushed stainless steel, soft directional grain, no mirror finish.”
Camera language was “slight three-quarter view, product off-center.” I replaced it with a height landmark and “product in the left two-thirds.”
Negative space was “room for headline.” I restated it as a hard right-third constraint.
Style language included “premium product photography.” I removed it for the diagnostic pass.
A six-image test batch cleared both the finish and the space issues. I restored a minimal mood note afterward. The client approved the revised set. The model had not changed. The prompt had.
What I No Longer Do When a Batch Fails
I no longer switch models as the first move.
I no longer add more descriptive language on top of an already long prompt.
I no longer generate large new batches while the failure mode is still undiagnosed.
I no longer treat “the model ignored me” as the default explanation.
All of those responses cost time and rarely fixed the underlying softness in the instructions.
The Practical Outcome

The checklist has not made every prompt succeed on the first try. It has made the failures shorter, more diagnostic, and far less expensive. Most commercial problems that look like model limitations turn out to be prompt structure problems once the language is forced into concrete rules and tested one block at a time.
Before I blame the model, I run the seven questions in order. I tighten the specific block that controls the failure. I confirm the fix with a small diagnostic batch. Only then do I decide whether the tool itself is the limiting factor.
Test the prompt before you distrust the model. The structure is usually the faster variable to change. The demo can look like a model failure. The checklist usually proves it was still a language problem.
Brands and cases are illustrative. For AI tool features, rules and availability, refer to their official sites.

Leave your thoughts here, too.