Imagination, Aligned

What AI can and can’t do with a story — and what is actually under threatWhat AI Can and Can’t Do With a Story (Image via MJ)Something shifted in the literary world this year, and you can date it by the temperature of the arguments. In May, Olga Tokarczuk, who won the Nobel Prize in 2018, told an…

What AI can and can’t do with a story — and what is actually under threatWhat AI Can and Can’t Do With a Story (Image via MJ)Something shifted in the literary world this year, and you can date it by the temperature of the arguments. In May, Olga Tokarczuk, who won the Nobel Prize in 2018, told an audience in Poznań that while working on her forthcoming novel she had asked an AI model what songs her characters might have danced to a few decades ago, and that the model gave her a few titles.In the same breath she said this book would likely be her last long novel, because the years such a work demands no longer make sense “from a purely economic perspective.” The world, she added, “with its destructive momentum, no longer deserves long, demanding novels.”The reaction was immediate and disproportionate. Headlines announced that a Nobel laureate had “used AI to write her latest novel.” Tokarczuk had to issue a clarification: she had written the book herself and had turned to the model only to look things up faster, the way most people now do. What complicated the defense was a second admission — that she had taken to addressing the model affectionately, calling it kochana, “my dear” — the feminine form, the way one might speak to a female collaborator (“kochana, how might we develop this beautifully?”) — and it was this hint of something warmer than a query box, as much as the bare fact of contact, that drew the heat.What is striking is not what she did, which amounts to a literate person using a search engine that talks back, but the fury that met the admission. The intensity was the real news. It told you that a line had been crossed in the collective imagination, even if no line had actually been crossed on the page.We have seen this film before. It ran in the visual arts in 2022 and 2023, when image generators arrived and the same emotional sequence played out: disbelief, then outrage, then accusation, then a slow, grudging normalization. The literary version is simply running a few years late. But there is one difference worth holding onto, because the rest of this essay turns on it.With images, the tell was on the surface — the six-fingered hand, the waxy skin, the melted background — and the models largely overcame it. There was a deeper tell too, harder to name and still with us: a certain average prettiness, a default “AI look” that pulled every image toward the same idea of the beautiful. That sameness is the visual analogue of what we will meet later in language as mode collapse, and it was reinforced from below, by the mass of users who reached for it precisely because it matched what they already took beauty to be.With prose, the surface tell is migrating in the opposite direction. It is sinking from the surface into the structure. And there, in the depths, it is far harder to see — and harder still to be rid of.The waveConsider the two scandals that bracketed Tokarczuk’s. In March, around ten thousand authors put their names to a book called Don’t Steal This Book, unveiled at the London Book Fair. Eighty-eight pages of names — Kazuo Ishiguro, Richard Osman, Mick Herron, Alan Moore, Marian Keyes, Philippa Gregory — followed by blank pages. The emptiness was the argument.Organized by Ed Newton-Rex, himself a refugee from the AI music industry, the protest was aimed at a proposed British exception that would let AI companies train on copyrighted books without permission or payment. The grievance here is about livelihood and consent, not about quality. Nobody in that book was worried that a model would out-write Ishiguro. They were worried it would be built out of him and then compete with him.In June came the subtler and more revealing case. Granta announced it would stop publishing the winners of the Commonwealth Short Story Prize after readers raised suspicions about “The Serpent in the Grove” by Jamir Nazir, the Caribbean regional winner. The supposed evidence of machine authorship was almost entirely surface: items grouped in threes, the now-notorious “not X, but Y” cadence.Nazir denied it, attributing his style to a speech-to-text method he uses because of a health condition. Nobody could prove anything. Granta’s publisher, Sigrid Rausing, conceded that the judges might have rewarded “an instance of AI plagiarism,” but admitted, “we don’t yet know, and perhaps we never will know.”That sentence is the hinge of the whole debate. Surface detection has collapsed. The rhetorical tics everyone has learned to recognize — the tricolon, the antithesis, the em-dash, the words “delve” and “tapestry” — are exactly the things that a newer model, or a careful human editor, can strip out in an afternoon.This is not speculation. Fine-tuning an AI on human-authored books has been shown to drop its detection rate on creative writing from 97% to 3% — while making its output the kind that readers actually prefer. If our only instruments are stylistic, we are reduced to Rausing’s shrug. We may never know.There are quieter signs that AI is already on the shelves: novels published with stray AI prompts accidentally left in the text, the steady reports of self-published genre titles flagged as largely machine-made. The anxiety, loudest among people who call themselves professional writers, is not irrational. But it is worth noticing that the word “professional” names a livelihood, not a level of literary quality. Keep that distinction in your pocket. We will need it at the end.The shrugRunning underneath the outrage is a calmer current. The largest survey of the writing profession to date now puts the share of writers using AI tools in some part of their process at a clear majority — 61% — and while fiction authors are the most wary, even among them, at 42%, the practice is no longer exotic.The data here is thinner than one would like, so a caveat is in order: a contemporaneous Authors Guild poll, sampling more traditionally literary members, found only 13% using it — a reminder that “writer” covers very different economies. There is a respectable case — the science-fiction novelist Debbie Urbanski makes it in Literary Hub — that novelists should treat these models as collaborators kept on a short leash, each assigned one narrow task, with the human firmly holding the wheel; others make the same argument from the writer’s desk. Tokarczuk’s own defense lives here too: a tool for faster research, nothing more.Acceptance and outrage are not actually in contradiction. They are responses to two different uses, and the line between them is precisely where the research draws it: the difference between using a model to find something — a fact, a title, a word, a direction — and using it to author something. To see why that line is real and not merely a matter of etiquette, we have to leave the anecdotes behind and look at what these systems actually do when you ask them to make a story.What the machine can and cannot doThe most interesting recent work stops asking whether AI prose sounds artificial and asks instead whether it is built differently. A team led by Jenna Russell at the University of Maryland did exactly this in a study called StoryScope.They generated more than sixty thousand stories from five leading models alongside authentic human-written storieы, then scored each one along hundreds of structural features — not diction or rhythm, but architecture: who drives the plot, how the timeline is arranged, whether the theme is stated or implied.Stripped of all stylistic information, these structural features alone could separate human from machine with 93% accuracy. More tellingly, the signal survived a deliberate stylistic rewrite almost intact. You can sand off the surface; the skeleton remains.And the skeleton has a consistent shape. AI stories explain themselves. Their narrators state the moral roughly 77% of the time, against 52% for humans. They prefer tidy single-track plots with the protagonist firmly in charge of the resolution (69% versus 46%) and far fewer loose subplots. They render emotion through the body — the tightening chest, the cold sweat — about 81% of the time, where humans are far likelier to simply name a feeling.They are also strikingly unspecific: where a human writer will name a real book, author, place, or brand, the machine reaches for the vague allusion, so humans point to particular named works and authors at nearly twice the rate (47% to 24%). A different habit, not the same one: humans also step out from behind the story to address the reader directly about four times as often (28% to 7%). And the people inside the stories differ too — the protagonists in the human stories are morally ambivalent more often than not (59% to 38%), while the ones the machine invents rarely are.Plot all of this as geometry — each story a point in the space of these measured features — and the five AI models cluster together in a tight, shared region, while the human stories scatter widely around them. By the study’s own measure of rarity, the human story is the most unusual of six versions of the same prompt 58% of the time.Why is this? The most plausible answer, and the one I find hardest to argue away, is that the flatness is not a sign of immature technology but the literary cost of alignment. A second study this spring, Narrative Flattening, traced the effect to its source. It took a single model at four successive checkpoints — the raw base model (OLMo-32B), then each stage of the post-training that turns a text predictor into a helpful, well-behaved assistant (supervised fine-tuning, then preference optimization, then reinforcement learning) — and had each version continue the openings of real stories drawn from three different pools: professionally edited literary fiction, stories posted on public platforms, and fiction written to order from explicit prompts.Measured against those human baselines, the flattening worsened at every stage of training. The thematic range narrowed, the emotional intensity drained, the stories grew more alike. And it was most noticeable for the professionally edited literary fiction — the corpus whose human distribution sits farthest from the model’s default, and therefore has the most distance to lose.Read that against what we know about how these assistants are made. The same process that teaches a model to be reassuring, to resolve your problem, to make its reasoning explicit, to avoid anything that might unsettle you, is the process that makes it a timid storyteller. Because fiction depends on the opposite virtues. A good story withholds, it refuses to state its meaning, it leaves you uneasy and does not apologize. An aligned model has been trained, at considerable expense, never to do any of those things. A good assistant makes a bad novelist and storyteller, alas.There is a neat confirmation of this in the research on Verbalized Sampling. If you stop asking a model for the answer and instead ask it to lay out a spread of possible answers, each tagged with a probability, and then take one from the unlikely tail, you can recover a large part of the lost diversity — in creative writing, between 1.6 and 2.1 times more — with no measured loss of quality. One caution is worth stating, because the technique is easy to mystify: the probabilities the model writes down are not a direct readout of its internal likelihoods but its own generated approximation of them — useful in a practical sense rather than bearing any relation to reality.As for why the trick works, the authors are less coy than one might expect. They trace the flatness to a “typicality bias” in the human preference data the models are tuned on — annotators systematically reward the more familiar text — and argue, with a formal model behind them, that asking for a distribution rather than a single answer lets the response approximate the broader spread the model learned in pretraining, before alignment narrowed it: different prompts collapse to different modes, and the distributional prompt collapses to a wider one.What the result shows either way is that the range was never absent, only suppressed. The diversity is still in there, underneath the training that taught the model to default to the expected, the typical, the thing an average reader would nod at. Which is to say: we flattened these models in our own image of what a safe, agreeable answer looks like.Time Is a Flat CircleAlignment explains the flatness in general. But at least one aspect of it is rooted deeper still, and deserves attention of its own, because it concerns the architectural foundations of large language models. The machine cannot bend time. And it cannot bend time because it does not live in it.Look again at the StoryScope features that most reliably mark a human author. They are overwhelmingly temporal. Humans jump across time. They use flashback and flash-forward, withhold a revelation and then detonate it so that you have to reread everything that came before in a new light. The study even has a name for that last move — the depth of recontextualization a surprise forces on the earlier scenes — and humans score far higher on it.AI, by contrast, tells the story straight: first clue to final reveal, cause neatly preceding effect. A human mystery might open at the funeral and spiral backward through decades. The machine opens with the first clue.Widen the lens from the single story to the whole writing life and the same gap appears. A separate study this year, Temporal Flattening, tracked 412 real authors across more than a decade of their own work (2012–2024) and found that human writers measurably drift: their concerns, their emotional register, the texture of their thought all shift over time.The models do not, even when you feed them their own past writing as context. One objection has to be headed off here, because it is the obvious one: this is not the claim that a model never changes. GPT-3.5 and its successors plainly differ. But that change is imposed from the outside, by a lab retraining the system between releases — it is not the model living through time and being altered by what it has written.Within any single fixed version, the model is stateless, each request begins from nowhere. The researchers found this temporal signature so distinctive that it alone could separate human from machine with 94% accuracy.Put the two scales together and you arrive at something close to a definition. A flashback is memory, foreshadowing is dread, the slow rewriting of the past in light of the present is what it feels like to have lived through something and understood it too late. These are not narrative techniques a model has failed to master. They are operations of a mind that persists through time — and whatever a model is, it does not persist that way at all. It has no past to remember and no future to fear, so it has no reason to fold the timeline.Its flatness, in the end, is timelessness. This is also why, as yet another recent study found, you can ask the same model the same prompt enough times and watch it drift toward a single average story. With no memory, there is nothing to individuate one telling from the next.What is automated, and what is threatenedLet’s now return to the outrage with a cooler head. The fear is that AI will write our literature. The evidence points to something more local, less alarming, but also, in a way, less flattering to us. AI does not automate writing. It automates the formula — and in doing so it reveals how much of the market already is formula.The clues are all there in the same studies. The flattening, remember, was worst for literary fiction: the further human writing sits from the formula, the further the machine falls short of it. Which means the human–machine gap is smallest where the writing is already most conventional.StoryScope’s detector worked least well on one genre above all: mystery and detective fiction, the most machine-like category it measured.Out in the actual marketplace, the AI flood is not lapping at the literary novel. It is rising fastest in self-published genre fiction, in the commodity end of romance and thriller and procedural, the easily formalized forms where a satisfying, well-made, unsurprising story is exactly what the contract with the reader calls for.It would be a mistake — itself a kind of flattening — to conclude that “AI takes genre fiction.” The finest genre writing is the rule-breaking kind. Ursula Le Guin, John le Carré, Patricia Highsmith are as safe from imitation as any writer on the Booker Prize shortlist, because their books defy the reader’s expectations. What AI can actually do is reproduce the median of every genre. And the median, commercially, is concentrated in mass-market category publishing, which is precisely where most working writers earn most of their living.The machine threatens the lower floors of the professional writing economy: the midlist, the work-for-hire — the steady output that never aspired to exceptionality but did let many of its authors pay the rent. The authors whose names open Don’t Steal This Book are right that the danger exists, but, I think, wrong about its nature.The machine cannot automate the creation of great literature. What it can do — and is preparing to do — is automate the industrial production of a standardized product, a commodity, and thereby reveal how much of what we have published all along was itself industrial product — a commodity on the shelf of the cultural supermarket, and not art as such.Writing against the grainIf the flattening is structural, the defense has to be structural too, and it comes in two parts depending on what you are fighting.Against mere repetitiveness — the sense that everything a model gives you is the same beige draft — the fix is partly mechanical, and Verbalized Sampling is the most useful trick I know. Instead of “write me an opening,” ask: “Give me ten openings that differ in voice and structure, each with the probability you’d otherwise have chosen it, and rank them.”Then work from the bottom of the list, not the top. (The probabilities, as noted earlier, are the model’s own approximations rather than direct readings of its internals — but the request still pries it off its defaults.) You are reaching past the model’s trained preference for the obvious and pulling up the range it would otherwise hide. Pair this with the one thing AI systematically refuses to do: name things. Real places, real books, real brands, specific dates, names.Against flattening proper — the deeper deadness of structure and time — there is no prompt that will save you, and it would be dishonest to pretend otherwise. Part of this deadness is alignment, and you cannot prompt a model out of its own alignment. But the temporal part is not alignment at all; it is architecture. A stateless system has no clock to bend, and no instruction installs one.Either way, what is missing has to be supplied by hand — everything the training and the architecture between them strip away: a timeline that refuses to run straight, a revelation placed so that it rewrites what came before, a theme left unstated, an ending that withholds the comfort of acceptance, a protagonist whose choices you cannot cleanly approve. The model will give you a competent, reassuring, well-made story. The difficult, unsettling, time-folded one is still yours to write.What is all this about? That distinctiveness, originality, and the unexpected are still valued — perhaps even more than before. Tokarczuk said the long, demanding novel was becoming uneconomic. She is probably right, but not because a machine can write one.It is because the machine can flood the market that the long novel used to compete in, with an endless supply of the average. The work that survives that flood will be the work that the average cannot reach. The human job, more than ever, is to be rare. So it has always been — and, it seems, it always will be. At least in any foreseeable future.Sources and further readingStoryScope: Investigating Idiosyncrasies in AI Fiction, Russell et al., arXiv:2604.03136Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction, arXiv:2605.27878Temporal Flattening in LLM-Generated Text, Cao, Go, Hu & Sushmita, Findings of ACL 2026, arXiv:2604.12097Do Large Language Models Always Tell the Same Stories?, arXiv:2606.17350Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity, Zhang et al., arXiv:2510.01171On detection collapse after fine-tuning (the 97%→3% figure): Chakrabarty, Ginsburg & Dhillon, Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers, arXiv:2510.13939Olga Tokarczuk on her last long novel: My Company Polska; on the AI controversy and her response: Literary Hub (1), Literary Hub (2)Don’t Steal This Book protest: The Bookseller, EuronewsGranta / Commonwealth Short Story Prize: The Guardian (19 May 2026), The Guardian (20 June 2026)The case for acceptance: “Why Novelists Should Embrace Artificial Intelligence,” Literary Hub; The IndependentWriters’ AI adoption (61% overall, 42% of fiction authors): Gotham Ghostwriters / Josh Bernoff, AI and the Writing Profession, Nov 2025; contrasting figure (13%): Authors Guild AI surveyThis story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!Imagination, Aligned was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →