Blog

The Caption Is the Last Thing You Draw: Why AI Can Write Your Pitch but Not Your Punchline

James Thurber’s 1939 cartoon “If Grant Had Been Drinking at Appomattox” shows a disheveled Union general slouched in a chair, rambling incoherently at a bewildered Robert E. Lee. The caption reads: “And so, General Lee, if you will just—uh—just sort of let me sort of sneak back into—uh—sort of sort of…” That trailing, stammering, perfectly flat sentence is the joke. The drawing is funny. The caption is devastating. Remove the caption and you have a man who looks unwell talking to a man who looks uncomfortable. Add the caption and you have the Civil War renegotiated by a drunk. The line does not explain the image. It completes it. And the person who wrote that line was the same person who drew that chair, that slouch, that confused Confederate posture. One hand, one brain, one timing.

The caption is the last thing you draw. I mean that literally. In my own process—and in the process of every working cartoonist I have ever watched—the caption arrives after the image is nearly finished. Sometimes after it is completely finished. Sometimes after it has gone to the editor and the editor has said “is there a caption?” and I have stared at the panel for forty minutes and written six versions and crossed out five. The caption is not a label. It is not a summary. It is the place where the visual argument gets its final twist—or its deliberate, devastating flatness. It is the moment where the cartoonist stops being an illustrator and becomes a writer.

Which is why the current enthusiasm for letting AI text models generate cartoon captions strikes me as a category error with real consequences. Not apocalyptic ones. Not “the death of cartooning” ones. Something more specific and more depressing: the slow erosion of a craft tradition that runs from Thurber through Charles Addams through Gary Larson and into the best webcomics working today—a tradition in which the relationship between drawn image and written language is not a division of labor but a single act of thinking.

What the Caption Actually Does

Consider Charles Addams’ captionless panels for The New Yorker, published throughout the 1950s and 1960s. A woman in a fur coat pours tea for a plant. A family on skis glides past a snowman they have clearly just murdered. A man in a suit reads a newspaper in a room where a small octopus sits in an armchair doing the same. No captions. No titles. The drawings are complete arguments. Addams’ decision to withhold the caption is itself a caption—an act of trust in the reader and a refusal to supply the verbal twist that would make the image digestible. The silence is the joke’s timing.

Now consider what happens if you ask an AI text model to generate a caption for an Addams panel. It will produce something competent and cheerful. “Just another quiet afternoon at home.” “Even the houseplant enjoys a good cup of tea.” “The whole family gathers around the morning paper.” These are not wrong. They are worse than wrong. They are adequate. They flatten the strangeness into a gag that explains itself, which is the one thing a great Addams panel never does. The model does not know what the drawing is not saying, and the not-saying is where the art lives.

Gary Larson understood this from the opposite direction. His Far Side panels almost always carried captions, and the captions were load-bearing structural elements. “Midvale School for the Gifted” shows a child pushing on a door marked PULL. The caption names the institution. The drawing shows the action. The joke exists in the gap between the name’s promise and the action’s failure. Remove the caption and you have a child struggling with a door. Remove the drawing and you have a school name. Larson’s 1986 Far Side Farm Calendar pushed this further—monthly panels with captions that recontextualized agricultural scenes into something surreal and specific. The caption “Cows listening to a recording of a bulldozer” does not describe the image. It reframes it. The drawing shows cows standing in a field. The caption tells you what they are hearing. The joke lives in the listener’s knowledge, not the cow’s expression.

Larson wrote those captions himself. He has spoken about the agony of caption revision—the process of writing twenty versions of a sentence and choosing the one that sounds most like a man muttering to himself. That muttering quality, that sense of a specific human voice with a specific rhythm, is what an AI text model structurally cannot produce. It does not mutter. It generates. It optimizes for plausibility. Plausibility is the enemy of the great caption, which is almost always slightly implausible in its phrasing, slightly off in its rhythm, slightly more or less than what the situation seems to call for.

The Negotiation Between Two Modes

The relationship between drawn image and written label has always been a negotiation between two modes of thinking. The visual mode is spatial, simultaneous, gestural. The verbal mode is temporal, sequential, syntactic. When a cartoonist writes a caption, they are not translating the image into words. They are placing a verbal object next to a visual object and letting the friction between them produce meaning. The caption can contradict the image, as in Thurber’s drinking Grant. It can understate the image, as in Addams’ silent panels. It can recontextualize the image, as in Larson’s listening cows. It can flatly describe the image in a way that makes the description itself the joke, as in many of Roz Chast’s New Yorker panels, where the caption’s clinical tone is the comedic instrument.

This negotiation is a creative-writing act. It is governed by voice, timing, tonal precision, and the specific relationship between form and expression that creative writing as a discipline has always studied. Purdue University’s creative writing framework treats this kind of form-and-expression relationship as core to the discipline, not peripheral to it—distinct from professional and technical writing tasks, with its own foundational skills and genre conventions. The cartoon caption belongs to the creative side of that boundary. It is closer to poetry than to copywriting, closer to a punchline than to a headline, and it requires the same quality of attention a poet gives to a line break: where does this stop, and what does the stopping do?

When a cartoonist delegates the caption to an AI text model, they are not merely outsourcing a sentence. They are surrendering the moment where the visual argument gets its final twist. They are giving up the line break. And they are doing so in a context where the alternative—writing the caption yourself—is not the hard part of the job but the part that makes the job worth doing.

What AI Text Tools Actually Do Well

None of this means AI text tools have no place in a cartoonist’s workflow. They do. The place is just not inside the cartoon. It is around it.

A working cartoonist spends a surprising amount of time writing prose that is not captions: grant proposals, artist statements, pitch letters to editors, editorial correspondence, exhibition text, newsletter copy, commission briefs, and long-form visual essays that accompany serialized cartoon projects. This writing is necessary, sometimes voluminous, and almost always done under time pressure. It is also, crucially, writing where the cartoonist’s personal voice matters but where the structure and organization of the prose matter more than any single sentence’s comedic timing. A grant proposal does not need a punchline. It needs a clear statement of project scope, a budget narrative, a work sample list, and a biographical paragraph. A pitch letter to an opinion editor needs a concise summary of the cartoon’s angle, a visual description, and a publication timeline.

This is the kind of writing where a tool designed for structured drafting can genuinely help. If you are organizing a long-form visual essay—a serialized webcomic with accompanying prose commentary, say, or a graphic journalism project that needs a written introduction for each installment—you need something that helps you manage structure, keep track of section flow, and draft the connective tissue between visual segments without asking the tool to invent your argument or your voice. When the scattered notes, thumbnail sketches, and section outlines for a serialized cartoon project need to become a coherent editorial structure, long-form drafting tools like Unsloppy can serve that organizing function: they hold the scaffolding while you write the sentences that matter, and they do not pretend to be funny. The line between using a tool for structural drafting of the writing-around-the-cartoon and using it for the writing-inside-the-cartoon is the line that matters. One is a tool. The other is an abdication.

The distinction matters ethically, not just aesthetically. The Authors Guild has published guidance for writers navigating AI tools, and their position is clear on the core tension: a creator’s original voice, thinking, and creativity are what make their work theirs, and commercially available large language models were trained on pirated, unlicensed creative work without compensating authors. The Authors Guild’s AI best practices establish that AI should assist with peripheral professional writing tasks—emails, administrative documents, research organization—rather than core creative work. Their framework maps directly onto the cartoonist’s situation: the caption is core creative work. The grant proposal is peripheral professional writing. The ethical and practical line falls in the same place.

The A/B Testing Problem

There is a newer phenomenon that complicates this argument, and I want to address it directly. Some contemporary webcomic artists now A/B test captions on social media: they post a panel with one caption, note the engagement, then repost with an alternate caption and compare. The practice treats the caption as a variable to be optimized rather than a decision to be made. The logic is platform logic: engagement metrics reward whichever phrasing produces the most immediate reaction, which is not the same as the phrasing that produces the best joke.

I am not arguing that A/B testing captions is the same as generating them with AI. It is not. The cartoonist is still writing both versions. But the two practices share a structural assumption: that the caption is a separable, optimizable component of the cartoon rather than an integral element of the visual argument. This assumption is wrong, and it produces a specific kind of degradation—captions calibrated for reaction rather than for rightness. A caption that gets more likes is not necessarily a better caption. It is a caption selected for the platform’s reward structure, which favors clarity, relatability, and immediate comprehension. All qualities that great captions tend to undermine.

Thurber’s stammering Grant caption would not have performed well in an A/B test. It is too long, too awkward, too committed to a specific rhythmic failure that mirrors the character’s drunkenness. An alternate caption like “Grant struggles to negotiate surrender” would be clearer, more relatable, more immediately comprehensible, and completely without comic value. The platform would prefer it. Thurber would not have written it. The fact that Thurber’s caption has been reprinted for eighty-five years and the alternate would have been forgotten in eighty-five minutes is not a counterargument. It is the argument.

The Caption Is Where You Are Most Yourself

I want to be specific about what is lost when a cartoonist stops writing their own captions, because the loss is not obvious to anyone who has not spent time staring at a finished drawing and trying to find the sentence that completes it.

What is lost is the cartoonist’s voice. Not their drawing style, which is visible and identifiable and relatively easy to preserve, but their verbal personality—the thing that makes a Thurber cartoon a Thurber cartoon rather than a competently drawn New Yorker gag. Thurber’s drawings are technically loose. His captions are technically perfect. The perfection is not in the grammar. It is in the timing. The sentence “And so, General Lee, if you will just—uh—just sort of let me sort of sneak back into—uh—sort of sort of” is timed like a comedian’s aside. It builds, it stumbles, it trails off. The drawing cannot do that. The caption does it for the drawing. And the person who timed that stumble was the same person who drew the slouch, because the stumble and the slouch are the same joke expressed in two media.

This is why the caption is the last thing you draw. Not because it comes last in the production sequence, though it often does. Because the caption is where the cartoonist’s whole sensibility—visual and verbal, spatial and temporal, gestural and syntactic—converges into a single act of expression. You can outsource your bookkeeping. You can outsource your newsletter. You can outsource your grant proposal and your pitch letter and your exhibition statement. You cannot outsource the moment where your drawing turns to the reader and says the thing that only you would say, in the way that only you would say it, at the exact moment only you would choose to say it.

What Remains Yours

The practical question for a working cartoonist in 2026 is not whether to use AI text tools. That decision has already been made by the economics of freelance creative labor, which require you to produce more administrative writing than any human can sustain without help. The practical question is where the line is. My answer, after twenty years of drawing and captioning, is simple: the line is at the panel’s edge. Everything outside the panel—grants, pitches, correspondence, project proposals, the structural scaffolding of a long-form visual essay—can be drafted with AI assistance, reviewed by you, and sent into the world with your name on it because the thinking is yours even if the typing had help. Everything inside the panel—the caption, the title, the speech balloon text, the label on the sign, the word on the protest placard—must be written by the same hand that drew the line it sits next to. Not because AI cannot produce a plausible sentence. Because the plausible sentence is not the one you want. You want the implausible one. The one that sounds like you muttering to yourself at midnight, staring at a drawing of a drunk general, trying to find the exact number of “sort of”s that makes the Civil War funny.

Thurber found it. It took him three.

So here is the question I keep coming back to: when you look at a cartoon you drew and read the caption beneath it, do you want to recognize your own voice—or someone else’s best guess at what your voice might sound like? The caption is the one place in the whole production chain where no one and nothing can substitute for the specific, irreducible, slightly strange way you talk to yourself while looking at your own drawing. Guard it. The rest is negotiable. That is not.

Comments Off on The Caption Is the Last Thing You Draw: Why AI Can Write Your Pitch but Not Your Punchline