I still have the folder. Thirty-seven generations, three different models, and one very patient client who eventually asked if we should just reshoot the product photo the old way. The image that started it looked almost right. That was the trap.
It was early 2023. A small DTC brand selling matte-black water bottles needed a single hero lifestyle shot for their main product page and a matching set of three supporting frames for paid social. The brief was simple on paper: bottle on a clean concrete ledge, soft overcast daylight, slight three-quarter view, enough empty space on the left for a short headline, no hands, no logos other than the subtle embossed mark already on the bottle. The existing product photo was clean but flat. They wanted something that felt lived-in without looking staged.
I opened Midjourney, wrote what I thought was a careful prompt, and generated the first batch. One frame came back that made me sit up. The light was soft, the concrete had the right texture, the bottle sat at a pleasing angle. I sent it to the client with a note that said we were close. They replied twenty minutes later: “Love the mood. The bottle proportions feel off and the embossed logo is distorted. Can you fix it?”
I did what almost everyone does the first time this happens. I added more words to the prompt.

The Spiral Begins
My second prompt was longer. I restated the bottle shape, added “accurate proportions,” “correct embossed logo,” and a few extra lighting notes. The new batch looked sharper in the grid. Two frames seemed better. I sent them. The client came back with the same notes plus a new one: the concrete ledge now felt too warm and the bottle finish had shifted from matte to slightly satin.
I added more language. “True matte black finish, neutral cool daylight, precise cylindrical proportions, undistorted embossed mark.” I raised the stylize value, then lowered it. I tried a different aspect ratio. I switched to a new seed range. I brought in a reference image and weighted it. Each round produced images that fixed one problem and introduced two others. The logo would lock in one generation and melt in the next. The proportions would stabilize and the lighting would drift. Negative space would appear and then get filled with soft shadows or a conveniently placed plant.
By generation twenty I was no longer solving the brief. I was chasing the last good frame and trying to reproduce it through sheer volume. That is the moment most prompting sessions turn expensive.
What I Was Actually Doing Wrong
Looking back, the mistakes were obvious and completely predictable.
I treated the almost-good image as a starting point instead of evidence that the foundation was incomplete. The first frame had lucked into attractive light and surface texture. It had not locked the product geometry or the logo treatment. Every new prompt I wrote assumed the model understood what “fix the proportions” meant in relation to the previous output. Models do not work that way. Each generation is largely independent unless you deliberately constrain it.
I kept adding corrective language on top of an already overloaded prompt. The model began to average competing instructions. “Accurate proportions” fought with the residual style pressure from earlier aesthetic words. “Undistorted logo” competed with the lighting and material descriptors. The result was the classic failure mode: images that looked more detailed and less usable.
I also never stopped to re-evaluate the brief itself. The client’s original product photo showed a specific bottle height-to-width ratio and a very subtle emboss. My early prompts described the bottle in general terms and then tried to correct the output after the fact. That sequence almost always loses.
The Point Where I Stopped Digging
Around generation thirty I did something I should have done at generation five. I closed the chat, opened a blank prompt, and wrote only the four-part foundation I now use on every job: product and material first, scene and lighting second, composition and hard constraints third, short exclusions fourth. No style language. No reference to the previous almost-good frames. No “fix the logo” instructions.
The new batch was less immediately pretty. Several frames were flat. Two of them, however, held the bottle proportions and the embossed mark well enough to survive a tight crop and a headline overlay. I sent those two with a short note explaining that we were starting from structure instead of chasing the earlier mood. The client approved one with a minor request for slightly cooler light. I adjusted the lighting block only, generated a final small set, and delivered.
Total time from the decision to restart: about forty minutes. Total time spent in the earlier spiral: closer to three hours.
The Lesson That Stuck
More prompting is not the same as better prompting. When an image is almost right but fails a hard commercial requirement—product geometry, logo integrity, usable negative space—the fastest path is usually to strip the prompt back to the minimum structure that enforces those requirements, not to keep stacking corrective adjectives.
I now treat any “almost” image as a diagnostic signal rather than a draft to be iterated. If the product proportions are wrong, I do not ask the model to fix them on top of the existing prompt. I rewrite the product block until the geometry holds, then rebuild the rest. If the logo distorts, I describe the mark more precisely or remove the demand for visible branding until the bottle itself is stable. Chasing the last good frame almost always costs more than restarting from a clean foundation.
This is also why I keep the Failed Generations folder. The thirty-seven frames from that water-bottle job are more instructive than the final approved image. They show the exact sequence of over-correction that most of us fall into the first time a client says “close, but not quite.”
Practical Rules I Use Now
When a generation is close but fails a commercial constraint, I force myself through a short checklist before I touch the prompt again:
Is the product description specific enough that the model has no room to invent proportions?
Have I stated the hard spatial requirements (negative space, camera height, framing) as constraints rather than preferences?
Am I trying to correct the output with new language instead of rebuilding the foundation?
Have I added so many style or reference terms that they are competing with the structural instructions?
If the answer to the last two questions is yes, I start a new prompt. It feels slower in the moment. It is almost always faster by the end of the day.
I also set a hard limit. If I have generated more than twelve images on the same structural prompt without a clear improvement in the failure mode I care about, I stop and rewrite. Volume is not a strategy when the foundation is wrong.
What the Client Actually Needed
The final approved image was not the most beautiful one I generated that day. It was the one that kept the bottle honest, left the left third open for their headline, and did not require the art director to explain why the logo looked melted. That is the only standard that mattered.
The earlier spiral produced several frames that would have performed well in a Discord showcase or a “look what AI can do” post. None of them would have survived the product page. The difference between those two outcomes is the entire reason this site exists.

The Takeaway I Still Repeat
When the first good-looking frame appears, do not celebrate. Check the hard requirements first. If any of them fail, do not add more words to the same prompt. Rebuild the foundation until those requirements hold, then layer mood and style on top of a structure that already works.
I burned an afternoon learning that lesson on a single water bottle. I keep the failed generations so I do not have to learn it again the expensive way.
The model was not the only thing I got wrong that day. The prompting strategy was. Once I stopped trying to fix the image with more language and started enforcing the brief with less, the work moved.
Test the structure before you trust the next generation. That is the only approach that has consistently survived the brief for me since.
Brands and cases are illustrative. For AI tool features, rules and availability, refer to their official sites.

Leave your thoughts here, too.