I needed the same product on five different backgrounds without the bottle changing shape, finish, or label placement. The brief was straightforward and therefore revealing: a frosted glass serum bottle, locked product description, five clean background territories, and a hard requirement that the product remain consistent enough for a carousel. No major lighting overhauls. No new props. Just reliable background swaps that did not force a full re-generation of the hero object.
I ran the same locked product through Midjourney, Flux, and Stable Diffusion under controlled conditions. Forty generations per tool, identical product block, matched variation structure, same scoring criteria. The question was simple: which tool kept the product most stable while actually changing the background.
Here’s what the test showed.
The Brief and the Scoring Rules

The product was a frosted glass dropper bottle with a minimal label and matte silver cap. The five backgrounds were:
Clean light stone surface
Soft neutral linen
Pale wood grain
Cool gray architectural surface
Simple soft-shadow infinite backdrop
Each background needed soft, even light that did not fight the frosted glass. The product had to stay recognizable as the same object across the set. Label text and dropper proportions were not allowed to drift.
I scored every generation on four binary checks before looking at overall attractiveness:
Product geometry holds against the locked reference
Material response (frosted glass + matte cap) remains consistent
Background actually changes in a usable way
Frame is clean enough for social or site use without heavy repair
Only frames that passed the first two checks counted as usable for the set.
Test Conditions
All three tools received the identical product block at the start of every prompt. Background and surface language changed; product language did not. I used the same four-part foundation structure: product and material first, scene and lighting second, composition and constraints third, short exclusions fourth. Style language stayed minimal.
I generated in small batches, reviewed against the scoring rules, adjusted one variable at a time, and continued until I reached forty attempts per tool. Wall-clock time included prompting, generation, and scoring. No external background-removal or compositing tools were allowed. The test measured what the models could deliver directly.
Midjourney Results
Midjourney produced the most immediately polished frames and the weakest product stability under background changes.
Early generations looked attractive. Light and surface texture responded quickly. When the background shifted, however, the bottle often shifted with it. Frosted glass picked up new specular behavior. Label placement drifted. Cap proportions softened or sharpened depending on the surface. I could get strong individual images; building a five-background set that still looked like the same bottle required heavy cherry-picking and still left visible inconsistency.
Out of forty generations I marked eleven as product-usable across different backgrounds. From those eleven I could assemble a set of five only by accepting small but noticeable variations in glass response and label clarity. Time to a workable shortlist: approximately 100 minutes.
Midjourney’s strength was aesthetic coherence per frame. Its cost was product drift whenever the background changed.
Flux Results
Flux delivered the highest number of product-stable frames across background changes.
Once the product block was locked, Flux kept geometry and material response more consistent when I swapped surfaces. The frosted glass stayed diffuse. The cap finish held. Label placement remained reliable enough for side-by-side comparison. Backgrounds changed cleanly without pulling the product into new lighting behavior as often as Midjourney did. The frames were sometimes slightly less “finished” in atmosphere, but the product itself was more trustworthy.
Out of forty generations I marked nineteen as product-usable. Building the five-background set was straightforward. Time to shortlist: approximately 85 minutes.
Flux’s strength in this test was product lock under variation. Its weaker point was that some backgrounds required an extra prompting pass to reach the same tonal polish Midjourney produced more easily.
Stable Diffusion Results
Stable Diffusion sat between the two on pure product stability and required the most workflow discipline.
With a tight character-style lock on the product and careful seed management, Stable Diffusion could hold geometry and material well. Without that discipline the drift increased. Background changes were possible, but the model was more sensitive to how the surface language was written. Some backgrounds introduced subtle proportion shifts or glass response changes that only became obvious in the side-by-side grid.
Out of forty generations I marked fourteen as product-usable. The usable rate improved in the later batches once I tightened the workflow. Time to shortlist: approximately 110 minutes, including the extra locking passes.
Stable Diffusion’s strength was controllable consistency when the workflow was tight. Its cost was higher setup effort and more sensitivity to prompt and seed hygiene.
Direct Comparison
On the metric that mattered most for this job—product stability across background changes—Flux produced the highest yield of usable frames. Midjourney produced the most attractive individual frames and the most product drift. Stable Diffusion produced solid results when the process was carefully controlled and required the most overhead to reach that control.
For a five-background set that had to read as the same bottle in a carousel, Flux got me to a clean deliverable fastest and with the fewest compromises on product fidelity. If the job had been a single hero image with a dramatic background, Midjourney’s aesthetic strength would have been more relevant. If I already had a heavily locked local workflow and needed maximum control, Stable Diffusion remained viable.
What Actually Drove the Differences
The largest practical difference was how each model treated the locked product block when new surface and light language was introduced. Flux appeared to protect the early product tokens more reliably under moderate scene changes. Midjourney more readily allowed the new scene language to influence material response and subtle proportions. Stable Diffusion could match or exceed Flux’s stability, but only after additional workflow constraints that the other two did not require in this test.
None of the tools delivered perfect invariance. All three still needed a human side-by-side check. The difference was how many frames survived that check.
How I Choose Now for Background Work
When the brief requires the same product on multiple backgrounds and consistency is the primary commercial risk, I start with Flux. The higher usable rate under background variation saves more time than the slightly higher polish of the alternatives.
When the background itself is the hero and product consistency is secondary, I start with Midjourney and accept the extra selection work.
When I need maximum local control and already have a locked pipeline, Stable Diffusion remains in the mix.
I keep the product block identical across any comparison. The moment the product language starts changing between tools, the test stops measuring background handling and starts measuring prompt noise.

Limitations
This was one product type (frosted glass bottle), five relatively simple backgrounds, and no major lighting redesigns. More complex scenes, reflective metals, or extreme style shifts would surface different failure modes. External compositing or inpainting workflows would change the numbers for all three tools. I tested direct generation under the time pressure and process I actually use for client sets.
The conclusion is therefore specific. For reliable product background changes where the bottle must remain the same object, Flux currently gives me the highest percentage of usable frames per generation batch in the workflow I trust.
The Takeaway
Background changes are easy until the product has to stay honest. The tool that protects the product while still accepting new surfaces is the one that survives this particular brief. Attractive drift is still drift. The set has to read as the same product or the carousel concept weakens.
I ran the forty-generation tests on all three tools so the next time a brief asks for the same bottle on multiple backgrounds I already know which path produces more keepers. Test the product lock under real background variation before you trust the single beautiful frame. The demo can look consistent. The set has to prove it.
Brands and cases are illustrative. For AI tool features, rules and availability, refer to their official sites.

Leave your thoughts here, too.