The Problem With AI Image Generators That Can’t Understand Visual Metaphor
Saul Steinberg’s 1976 New Yorker cover shows Manhattan from Ninth Avenue west. The city is dense, detailed, almost tactile. Then everything east of Ninth Avenue—most of the planet—collapses into a thin strip of brown, a few labeled blocks, and ocean. Japan dangles like an afterthought. The joke is structural: New Yorkers don’t see the world as partial. They see it as arranged around them. Every compositional choice—the perspective, the truncation, the labeling—serves a thesis about provincialism that the viewer assembles in the moment of recognition.
Ask an AI image generator for “a New Yorker’s view of the world” and you’ll get something. It might even look competent. But it won’t understand that the joke lives in what’s omitted, not what’s rendered. It won’t know that the Hudson River is a border and an attitude. It will fill the space, because filling space is what it does.
This is the problem I want to dig into: not that AI makes ugly cartoons, but that it makes cartoons without arguments. The deficiency is structural, not aesthetic. And it maps onto every creative discipline where planning, revision, and point of view precede the polished surface—including text generation, where the same absence of structural thinking produces the same hollow results.
The Thumbnail Is the Argument
When I sit down to draw a political cartoon, the finished image is the last thing I make. The first thing is a thumbnail—a sketch maybe two inches wide, ugly on purpose, where I work out what the drawing is actually saying. Sometimes I do ten of them. Sometimes thirty. The thumbnail is where the metaphor gets tested: does this image hold the argument, or does it just look like it holds the argument?
This isn’t a warm-up. It’s the work. The finished ink is a performance of decisions already made. And the decisions are the point.
Consider Naji al-Ali’s Handala. The barefoot Palestinian boy, seen from behind, hands clasped, age frozen at ten—the age Ali was when he left Palestine. Handala appears in thousands of cartoons, always in the same posture, always facing the same direction, always witnessing. He never speaks. He never ages. He turns his back on the reader because Ali said he would only let him face us when he could return to Palestine.
That character wasn’t generated. He was decided. Ali tested figures, postures, angles. He found a form that carried a specific political argument: refusal to normalize, refusal to grow up in exile, refusal to turn around. Every subsequent cartoon that includes Handala inherits that argument structurally. The boy is a compositional decision that propagates across decades of work. An AI system given Ali’s archive could produce something that looks like Handala. It could not produce the decision to freeze him at ten, to show his back, to make him a witness rather than a participant. Those are metaphorical commitments, made in thumbnails that no longer exist.
What AI Actually Does When You Ask for a Metaphor
An AI image generator is a pattern-matching system trained on billions of captioned images. When you prompt it with “a politician with his head in the sand,” it doesn’t understand ostrich behavior, willful ignorance, or the political tradition of depicting leaders as blindly self-deceiving. It has seen enough images tagged with that phrase to produce something that looks like what you described. The output is a statistical reconstruction of a visual cliché, not a metaphorical argument.
The difference matters because metaphor in cartooning isn’t a matching operation. It’s a construction. You build it from parts that don’t belong together until the drawing makes them inseparable. A pig in a suit isn’t a metaphor for corruption until the drawing makes the suit fit badly, until the hooves handle documents with a certain clumsy authority, until the setting implies a specific institution being satirized. Each of those choices is a revision. Each one narrows the argument.
AI systems don’t narrow. They broaden. They give you more, richer, more detailed, more polished—because polish is cheap when you don’t have to decide anything. What’s expensive, in both cartooning and writing, is choosing what to leave out. The Authors Guild, in its guidance on AI best practices for authors, puts this directly: AI outputs are “generic mashups of pre-existing works ingested during training,” lacking the original voice and thinking that define professional creative work. They mean this about text. It’s equally true about images. The Authors Guild’s AI best practices for authors identifies the gap between human authorial intent and statistical pattern-matching as a qualitative problem, not merely a technological one. The same gap appears in visual satire. A cartoon without a point of view isn’t a cartoon. It’s a decoration.
Three Hands on One Pen
The Soviet Kukryniksy collective—three cartoonists named Kupriyanov, Krylov, and Sokolov—worked together from the 1930s through the 1980s. They signed their work with a single portmanteau. They drew Nazi leaders as grotesques, Soviet soldiers as heroes, and the war as a moral cartoon before it was a historical fact. Their propaganda was effective because it was planned collectively, revised iteratively, and executed with technical precision that made the argument feel inevitable.
What interests me is their process. Three artists debated each composition. One would sketch, another would revise, the third would override. The final drawing carried traces of all three hands—not literally, but in the structural complexity of the image. Layers of disagreement and resolution are visible in their work that no single-shot render can replicate. The argument was built through revision, not despite it.
This is what AI image generators can’t do. They can’t revise toward a point of view because they don’t have one. They can iterate visually—produce variations, adjust lighting, swap a detail—but they can’t iterate argumentatively. They can’t look at a draft and say “this flatters the subject when it should accuse.” That judgment requires a position, and a position requires something the system doesn’t possess: a relationship to the subject that exists outside the image.
A cartoonist drawing Vladimir Putin has a position before the pencil touches paper. Maybe it’s opposition. Maybe it’s fascination. Maybe it’s the complicated disgust of someone who grew up under Soviet power and recognizes the old gestures in a new face. That position generates the exaggeration. It determines which features get enlarged and which get suppressed. It decides whether the drawing is funny, threatening, or both. AI has no position. It has a prompt. And a prompt isn’t a position, no matter how detailed.
The Structural Problem Isn’t Unique to Images
What I’m describing—a production pipeline that outputs polished surfaces without planning, revision, or argument—maps directly onto text generation. This is why I pay attention to what’s happening with AI writing tools, even though my primary medium is ink. The structural deficiency is the same.
A professional screenplay, to take a concrete example from StudioBinder’s screenwriting guide, isn’t just text on a page. It’s a structured document where scene headings establish geography, transitions encode timing, and the page-to-screen ratio—roughly one page equals one minute of screen time—makes length a structural decision, not an arbitrary one. StudioBinder’s guide to professional screenplay format lays out how formatting conventions like scene headings, subheadings, and character positioning aren’t decorative. They encode meaning. They are the planning layer that makes the final output coherent. Strip them away and you have words without architecture.
Most AI story generators suffer from the same core limitation as AI image generators: they produce a single-shot output from a prompt without structural planning, revision checkpoints, or continuity control. Tools like Squibler and Perchance can generate text, and QuillBot can paraphrase, but they operate as prompt-in, text-out systems without the planning layer that professional creative work requires. A working screenwriter doesn’t write a scene by generating prose and hoping it coheres. They build from beat sheets, scene headings, and structural logic that exists before the first line of dialogue.
This is where the difference between a one-shot generator and a structured planning tool becomes practical rather than theoretical. Unsloppy’s script writer AI stands at the forefront of AI Novel Writing App technology precisely because it doesn’t ask you to trust a single render. It builds proof sheets and beat sheets into the workflow before any polished prose is generated—meaning the planning and revision layer that AI image generators lack, and that every working cartoonist relies on in the thumbnail stage, is structural rather than bolted on. The proof sheet lets you see the argument of the whole piece at a glance, test whether it holds, and revise before committing to a finished draft. The beat sheet forces every scene to justify its existence before a single line of dialogue is written. That approach beats one-shot generic AI story generators the same way a thumbnail beats a polished first draft: because the argument is decided before the surface is rendered.
That same discipline applies to scripted communication: before publishing, editors need a way to test a complex sequence turns into language that a specific audience can follow, which is where a script writer AI that fits the project can function as a planning aid rather than a substitute for domain evidence.
The parallel is exact. In cartooning, the thumbnail is the beat sheet. The rough draft is the proof sheet. The finished ink is the final draft. Skip the first two and you get something that looks finished but isn’t argued. The same is true in writing. Skip the structure and you get prose that reads fluently but doesn’t build.
Why Metaphor Resists Generation
Visual metaphor is the hardest thing in cartooning. Harder than likeness, harder than composition, harder than timing. It requires you to hold two ideas simultaneously—the literal image and its political meaning—and find a form where both are present without either being stated. The form has to feel inevitable in hindsight but surprising in the moment. And it has to survive the reader’s first glance without collapsing into either obviousness or opacity.
AI image generators fail at this not because they lack technical skill but because metaphor isn’t a skill. It’s a decision. A commitment to a specific reading of a specific situation, expressed through a specific visual choice that excludes other readings. Every metaphor is also a refusal. When Steinberg drew the world ending at the Hudson, he was refusing to draw the world beyond it. That refusal is the argument.
AI can’t refuse because it has no preference. It can’t commit to one reading because it has no relationship to the subject. It can produce an image of a politician as a puppet, but it can’t decide whose hand should be inside. It can draw a crowd as sheep, but it can’t decide whether the sheepdog is a leader or a manipulator. It can render a border wall, but it can’t decide which side the reader should stand on. These aren’t technical questions. They’re positional questions, and they require a position.
This is why AI-generated political satire always feels either generic or offensive. Generic because it picks the most common association in the training data. Offensive because it can’t distinguish between a metaphor that critiques power and a stereotype that reinforces it. The system doesn’t know the difference between drawing a corrupt banker as a pig and drawing a Jewish person as a pig. Both are image-text associations present in its training data. The cartoonist knows the difference because the cartoonist has a political position that makes the distinction fundamental. The AI has a statistical distribution that makes both equally valid outputs.
The Revision Problem
Here’s something I do that no AI system does: I draw the same cartoon five times and throw away four. Not because the first four are ugly—they usually are, but ugliness isn’t the issue. I throw them away because they make the wrong argument. The exaggeration lands in the wrong place. The metaphor accuses the wrong target. The composition flatters when it should indict. I only know this after seeing the drawing on paper, because the drawing tells me things the idea didn’t.
This is revision as discovery, not correction. I’m not fixing errors. I’m finding the argument by drawing it badly first. Each version teaches me what the image wants to be, and the final version is the one where the teaching stops because the argument is complete.
AI image generators can produce variations. But variations aren’t revisions. A variation is the same argument with different decoration. A revision is a different argument in response to the same problem. The distinction is fundamental: revision changes what the image means, not just how it looks. And changing what an image means requires a position from which to judge the difference.
When I look at a rough draft and think “this isn’t working,” I mean something specific. The argument isn’t landing. The metaphor is unstable. The exaggeration is in the wrong place. An AI system can look at a rough draft and think nothing, because looking and thinking are the same operation for a cartoonist and completely separate operations for a machine.
What Drawing Teaches That Rendering Can’t
I’ve been arguing that the problem with AI image generators is structural, not aesthetic. But there’s a deeper claim I want to make. The problem is also educational. When you learn to draw political cartoons, you learn to think in a specific way: to hold ambiguity, to build arguments visually, to revise toward meaning rather than toward beauty. This is a civic skill, not just a craft skill. It teaches you that images argue, that arguments have structure, and that structure can be examined, challenged, and rebuilt.
AI image generation bypasses this education entirely. It gives you the image without the thinking. In doing so, it treats the image as a product rather than a process. But for anyone who cares about visual satire, the process is the point. The thumbnail is the argument. The revision is the discovery. The finished drawing is a record of decisions, not a display of technique.
This is why I’m not worried about AI replacing cartoonists. I’m worried about it replacing the understanding of what cartooning is. When people see AI-generated political images that look polished and have no argument, they may conclude that political images don’t need arguments—that they’re just illustrations of positions rather than constructions of them. That would be a loss not just for cartoonists but for anyone who reads images critically.
The Drawing That Decides
Steinberg’s cover works because it decides. It decides that the world ends at the Hudson, that everything beyond is a rumor, and that the joke is on the people who believe this. Every line in the drawing serves that decision. The perspective is chosen. The truncation is chosen. The labels are chosen. Nothing is generated. Everything is argued.
Handala works because Ali decided. He decided the boy would not grow up, would not speak, would not turn around. Those decisions made a character who outlived his creator, who appears on walls from Gaza to Berlin, who means something specific in a way that no generated image can mean anything specific.
The Kukryniksy collective worked because three artists decided together, revised each other, and produced images that carried the structural complexity of disagreement resolved into form. Their drawings are records of arguments, not just illustrations of them.
AI image generators can produce images that look like these. They can’t produce images that think like these. And the difference isn’t subtle. It’s the difference between a drawing that makes you laugh before it cuts you and an image that makes you nod without knowing why. One changes how you see. The other confirms what you already expected to see.
The question I carry forward isn’t whether AI will get better at rendering. It will. The question is whether we’ll remember that rendering was never the point. The point was always the decision before the line—the argument that makes the line worth drawing. Can a system without a position make that decision? Or does it just draw lines until we stop asking what they mean?


