# The Reproduction Problem

# The Reproduction Problem

Feb 26, 2026
# The Reproduction Problem
00:00
19:21
There is a particular frustration that anyone who has worked seriously with AI image generation eventually encounters. It arrives not at the beginning, when everything feels possible, but later — after the initial wonder has settled and the real work begins. You have an image. A specific image. And you need to reproduce it. Not interpret it. Not reimagine it in a complementary style. Reproduce it, with the kind of fidelity that makes a client nod instead of wince. This is when you discover that the tools built for creation are quietly, fundamentally ill-suited for reproduction. And understanding why requires making a distinction that sounds philosophical until you realize it's entirely practical: the difference between judgment and discernment. --- Judgment, as most AI systems practice it, is evaluative. It asks whether something is permitted, appropriate, on-brand, safe. It operates like an editor trained primarily on what to cut. This is not a flaw — moderation standards exist for good reasons, and evaluative filtering does necessary work. But judgment is not the same as structural analysis, and when a system defaults to one while attempting the other, something quietly breaks. Discernment asks a different set of questions entirely. Not *is this acceptable* but *what is this, precisely*. Not *what does this image look like* but *what makes this image structurally itself*. It is analytical where judgment is evaluative, taxonomic where judgment is moral. And it is, almost entirely, what current image-to-prompt systems lack. The gap becomes visible the moment you ask a system to reproduce something specific. --- Take a marketing composite that appears, at first glance, almost laughably simple: a 3×3 grid of nine different people, each holding the same book, each photographed in a different environment. The kind of social-proof image you've seen a thousand times. Scroll past it and you register diversity, warmth, accessibility. Which is precisely what it's designed to communicate, and precisely what makes it treacherous to reproduce. A system operating in aesthetic mode reads this image the way a casual viewer does. It notes the expressions, the varied demographics, the approachable body language, the warm lighting. It writes a prompt that will generate something that *feels* like the original — diverse, warm, accessible. And then it produces nine tiles in which the book cover drifts across every single one. Typography shifts. The spine changes weight. The panel layout, which in the original was deliberately asymmetric — blank red on the left, full cover on the right — gets quietly normalized into something more balanced, more orderly, more wrong. The system didn't fail aesthetically. It failed structurally. It looked at the image and saw what the image was *saying* rather than what the image *was*. What the image actually was, beneath its social-proof surface, was a precision engineering problem. The true invariants had nothing to do with warmth or diversity. They were: the exact book cover, reproduced identically across all nine tiles. The asymmetric panel layout, maintained without exception. The grid geometry, consistent to the pixel. The object perspective, held within tolerance. The people were interchangeable. The backgrounds were interchangeable. The lighting was negotiable. The book was not. It was the anchor around which everything else had to organize itself — and a system that couldn't identify that hierarchy had no hope of reproducing the image, regardless of how beautifully it described what it saw. --- This hierarchy isn't unique to book grids. Anyone who has spent time in commercial product photography understands it viscerally, even if they've never named it. Consider what happens when a spirits brand commissions a campaign series — twelve images, one for each expression in a whisky range, each bottle photographed in a different environment. Misty highland moorland for the peated single malt. Warm library light for the aged sherry cask. Cool slate and rain for the coastal release. The environments are the variables, chosen to evoke the character of each whisky. The bottle is not a variable. The label must be identical across every shot — same typeface weight, same foil registration, same capsule color rendering under wildly different lighting conditions. The photographer doesn't achieve this by describing the bottle more accurately to an assistant. They achieve it by treating the bottle as infrastructure. It ships to every location. It sits on the same custom-built mount. It is lit with its own dedicated source, independent of the ambient environment built around it. The world changes. The bottle doesn't. When brands attempt to replicate this kind of series using AI generation — either to extend the campaign or to prototype it before commissioning — they encounter the same failure the book grid exposed. The system reads each prompt and re-synthesizes the bottle from scratch, because that's what generative systems do. The label drifts. The foil shifts from gold to brass to something the brand manager can only describe as "wrong." The capsule changes diameter in ways that are imperceptible in isolation and mortifying in a series. No individual image is obviously broken. The set is unusable. The fix, again, is object dominance. The bottle is declared the anchor. Every environmental element — the light, the surface, the atmospheric haze — is specified in relation to it rather than independently. The prompt stops describing a scene and starts describing a scene *organized around a fixed object*. It is a small grammatical shift with significant structural consequences. The system stops treating the label as something it needs to imagine and starts treating it as something it needs to preserve. That reordering is the difference between a campaign and a collection of near-misses. What the packaging photographer knows, and what the AI system must be taught, is that some objects in an image are not subjects. They are facts. And facts don't get reinterpreted — they get protected. --- At the top of any reproduction hierarchy sit the structural invariants: product packaging, typography, asymmetric layout rules, grid geometry, object orientation. These are locked. Below them are contextual variables — people, environments, lighting conditions — which can shift within reason. At the bottom is decorative noise: the minor background details and micro-atmospheric differences that neither define the image nor threaten its integrity. A system that treats all three tiers as equally weighted will drift. The drift is imperceptible at first and catastrophic at scale. The correction, when it comes, feels almost insultingly simple: establish object dominance. Declare the anchor. Build the prompt not around what the image depicts but around what the image cannot afford to lose. Once a system understands that the book is the constraint and everything else adapts to it, the output stabilizes in a way that no amount of descriptive precision could achieve. You haven't made the system more creative. You've made it more disciplined. The distinction matters. --- There is a version of this problem that is purely technical, and smarter engineers than I am are working on it — building systems that extract structural invariants automatically, that detect meaningful asymmetry rather than correcting for it, that enforce negative constraints as rigorously as positive ones. That work is real and necessary. But there is another version of the problem that is cognitive, and it applies to anyone using these tools right now, today, with the systems as they currently exist. It is the habit of describing images the way we experience them — emotionally, atmospherically, from the outside — rather than the way they are built. We reach for mood when we should reach for structure. We describe what an image *communicates* when we should be mapping what it *requires*. The shift is not natural. We are not trained to see this way. Looking at a warm social-proof grid and thinking first about panel asymmetry and object-locking tolerance runs against the grain of how images are designed to be received. Which is precisely why it's the skill worth developing. --- Image generation is, by now, almost easy. The models are capable, the interfaces are accessible, the outputs are frequently astonishing. But reproduction — faithful, tolerance-level, structurally precise reproduction — remains genuinely hard, and it will stay hard until the tools are built around a different question than the one most of them currently ask. The question most systems ask is: *what should this image look like?* The question that reproduction demands is: *what must this image not lose?* Answering the second question requires discernment — not the evaluative, compliance-oriented filtering that AI systems have been carefully trained toward, but the cold, taxonomic, structurally rigorous kind that separates invariants from variables and protects the former without apology. It is less glamorous than creativity. It is less legible than judgment. It produces no interesting edge cases and offers no opportunity for the system to demonstrate its aesthetic sensibility. It just works. And in a field that has spent considerable energy celebrating what these tools can imagine, there is something quietly radical about insisting on what they must remember.
¿Te gusta esta publicación?

Comprar Clarence Coggins un Book

Más de Clarence Coggins