← All posts
·5 min read·Oranokai

The failure modes that don't show up in demo videos — held-product label deformation, one-prompt-for-every-category, and generic filler in slots that should carry real product data. What causes them and how to actually fix them.

Most AI product photography content stops at "upload a photo, get a studio shot back." That's true for the easy 80% of cases. It's the other 20% — the renders that look almost right, that pass a quick glance and fail on close inspection — that actually cost sellers money, because they ship.

These three mistakes are the ones we see most often, and they're not prompt-wording problems. They're structural: the same generation approach applied to every product, regardless of what the product actually is.

Mistake 1: One interaction pose for every category

The standard "lifestyle in use" prompt for AI product photography looks something like this:

A person actively using [product] in a natural real-world setting, clearly interacting with the product.

That phrasing works well for a huge range of products — skincare, small electronics, kitchen gadgets, anything genuinely held in a hand. It fails, specifically and predictably, for anything that isn't held: footwear, clothing, bags. Feed a running shoe through a "held and interacted with" prompt and you get exactly what you asked for — a shoe held up in someone's hand and inspected like a piece of jewelry. Compositionally, it reads as wrong the instant a shopper sees it, even if they couldn't articulate why.

The fix isn't a better prompt. It's a different prompt per interaction type. Products fall into a small number of real interaction categories:

  • Held — skincare, cosmetics, small electronics, tools you'd pick up and examine
  • Worn — footwear, clothing, bags, jewelry actually worn on the body
  • Used at a surface — kitchen tools, desk items, anything used in place rather than carried

A shoe needs "worn naturally, in motion or standing" language, not "held and inspected." Get this one variable right and the rest of the lifestyle shot — lighting, framing, product visibility — can stay exactly the same template it was before.

Mistake 2: Generic filler in slots that should carry real product data

Amazon's dimension/size-spec slot exists to answer one question: does this actually fit? For a water bottle or a backpack, "product size and form" is a reasonable photographic brief. For a skincare jar, it produces something like "Herbal Extract Enriched" or "Airtight Screw Cap" in that slot — true statements about the product, completely wrong content for a slot whose entire job is communicating scale and fit.

The underlying cause is almost always the same: nothing in the generation pipeline told the copy step that this slot means something different for this category. Left to its own defaults, an AI copywriter fills every slot with whatever generic benefit claims are left over, in the order they occur to it — not by what that specific slot is supposed to prove.

The fix is category-aware slot content, not category-aware prompts alone:

CategoryWhat the "spec" slot should actually show
Skincare, cosmeticsKey active ingredients — ingredient name + what it does
FootwearReal sizing/fit facts — size range, width, true-to-size vs. runs small
ElectronicsWhat's in the box, or a comparison grid against the model it replaces
Everything elseActual dimensions, actual scale

If your pipeline can't tell a skincare product from a shoe when it decides what goes in that slot, it's going to keep producing plausible-sounding, useless content there — and a plausible-sounding wrong answer is harder to catch in QA than an obviously broken one.

Mistake 3: Label deformation on held-product shots

This is the hardest of the three, because it's a real limitation of the underlying models, not a prompting gap. When a generative model redraws a product being held in someone's hand, it's redrawing the entire scene, including the product's own printed label — and printed text is exactly the kind of fine detail these models are weakest at reproducing faithfully. The failure looks like a jar label reading "SKIN BRIGHTENING PACKAGING" instead of "SKIN BRIGHTENING CREAM," or fine print that's legible-shaped but not actually legible.

Three things measurably help, in order of effort:

  1. Use a product-lock model, not a general scene-generation model, for anything with a product in-frame. Models explicitly built to preserve a reference image's exact details (rather than reinterpret it into a new scene) garble far less — the tradeoff is usually cost, since fidelity-focused models tend to run more expensive than free-form generation.
  2. Test whether a different model class helps for this specific failure. Not every model handles held-product label fidelity the same way, and it's worth an isolated, single-image test on the exact failure case before assuming it's unsolvable — some newer models handle in-hand product text meaningfully better than the defaults most pipelines reach for first.
  3. Never rely on a held-product shot to carry your fine print. Give your close-up/detail slot that job explicitly, with a macro shot where label text is the actual subject, not incidental detail in a wider scene. A held-product lifestyle shot's job is context and desirability, not label legibility — asking one image to do both is asking for the failure mode above.

The pattern underneath all three

None of these are "the AI is bad" problems. They're all the same structural gap: a pipeline built around one default case (a genuinely handheld product, sold in one standardized category, with no fine print that matters), applied uniformly to every product that comes through it. The fix in every case is the same shape — stop treating "product photography" as one problem, and start treating it as a small number of distinct problems that happen to share a rendering pipeline.

If you're evaluating an AI product photography tool, ask it directly: does the interaction pose change for worn items? Does the spec slot know the difference between a skincare jar and a water bottle? What happens specifically when a product is held in a hand — is that tested, or assumed to work? The answers tell you more about whether a tool will hold up at catalog scale than any single sample render will.

Try Oranokai — the free tier gives you 20 credits, enough to run these exact tests against your own product before committing to a full catalog.

Want to try Oranokai? Free tier gives you 20 credits — enough to test on your own product.