Can a Machine Write Like A Human?
Can a Machine Write Like a Human?I had a little research project for an educational institution I’m currently enrolled in, and I might have overdone it. After spending 30 hours analyzing 63 pieces of text with 819 numerical values, which I’m still slightly confused about, I realized that I really…
Can a Machine Write Like a Human?I had a little research project for an educational institution I’m currently enrolled in, and I might have overdone it. After spending 30 hours analyzing 63 pieces of text with 819 numerical values, which I’m still slightly confused about, I realized that I really didn’t need to go this far to get an A on a project that’s worth 5% of my grade.(I often realise this far too late)The question it aims to answer is simple: Can a Machine Write Like a Human?This essay is going to be an attempt at repurposing my 30-minute-long slideshow notes, which I pumped out the weekend before the deadline, because I now know that I am only allowed to yap for 10 minutes. Let’s get into it.Can a Machine Write Like A Human?If you’re in the writing world, you might have heard of the Commonwealth Short Story Prize. It selects five regional winners annually, who then compete for the overall prize.One of the stories winning the prize, “The Serpent in the Grove” by Jamir Nazir, went absolutely viral after being named a regional winner. Critics on X and Bluesky claim it shows obvious markers of AI use.Image from EksentrikaSupposedly, the story has multiple “not x, but y” constructions and lists of three, which can be a telltale sign of AI usage. The Guardian also says that lines like “She had the kind of walking that made benches become men” and “Sun on galvanise is a cruel instrument” are just weird chunky metaphors.(Read another article I wrote discussing telltale signs of AI writing)At the same time, unlikely metaphors are also a sign of good original writing, and em dashes used to be the hallmarks of refined society. This raises the irrelevant question: Is anything that feels like AI necessarily slop?Foundation director-general Razmi Farook says, “A young writer in Kingston or Kolkata, in Kuala Lumpur or Kigali, must now prove not only her talent but her very humanity.”(Why do they have to pay because tech giants have stolen art from artists and are using it to create machines that generate more art?)I’m not here to argue whether Jamir used AI or not. AI detectors are frequently wrong. In fact, about 61% of human-written essays by non-native English speakers are flagged as AI-generated.While the ethicality and nuance behind all of this are important and should be discussed, I’d like to focus on metrics. My little research project just focuses on quantifying how well AI imitates the style of human writers.Questions, Procedures and Results:Here’s the research question:To what extent do stylometry metrics quantify the success of a Large Language Model in imitating the writing style of prominent literary authors: George Orwell, Mary Shelley, and Jane Austen?These authors are contemporary geniuses, and absolutely no one can say their work is inspired by a machine. Except for an AI detector:AI Detector flagging the first paragraph of Mary Shelley as 100% AI-generated.Stylometry is just a statistical way to compare and assess the styles of a few texts. It’s based on the idea that every author has a subconscious linguistic fingerprint. They may use certain function words more often or punctuate in certain ways.Stylometry was used back in 2013, when J.K Rowling published a book called Cuckoo’s Calling. Except no one really knew she published it. She wrote it under the pen name Robert Galbraith. The Sunday Times was able to confirm that Cuckoo’s Calling was actually written by J.K Rowling.Metrics Analyzed:Sentence Length: no. of words per sentenceSentence Length Standard Deviation: how much the sentence length strays from the text’s average (This is also called burstiness. High burstiness means a rhythmic mix of short and long sentences)Lexical diversity: the number of unique words per hundred words in a text *unique words refe to those used only once in the textNouns, Verbs, Adjectives, AdverbsMost Frequently Used Words: which are function words like “for”, “she”, “he”, etc.I chose three human excerpts from each author and built neutral prompts based on the storyline of the excerpt.I tried to change the names and nouns, or not mention them altogether, so that the AI model doesn’t recognize and pull lines directly from the chosen text.However, I found that even though I changed all proper nouns, the AI model was able to identify which book I was pulling the text from. In the prompts where I didn’t mention names, it often took the names of the characters in the book. That was the only time I was annoyed to see Elizabeth Bennet and Fitzwilliam Darcy in the same sentence.(If you don’t get that reference and haven’t watched or read the Oscar-nominated Pride and Prejudice on screen, please put this article down. Nothing is quite as educational as Elizabeth and Darcy nailing banter.)Since this could be a confounding variable, I also decided to make three neutral prompts, which are based on modern-day nuisances like traffic, phones, and space travel (not a nuisance)To ensure that an LLM’s inherently stochastic response doesn’t interfere with general patterns, I generated 3 AI texts for each prompt.Let’s analyze the results:Sentence Architecture:If you look at the sentence length, you can see that AI writes in shorter sentences compared to human writing.While the average of Mary Shelley’s human texts has a sentence length of about 33 words, AI-generated texts (neutral and story-based) only hit 20–25 words per sentence. This average is actually closer to academic articles, not creative writing.We can also see that, while across human texts there’s a large variance (about 7+/- the average) in average sentence length, AI texts have a standard deviation of only 2.7. It seems to be consistent and mechanical, very much unlike human creativity.Read the following excerpts:Mr. Knightley might quarrel with her, but Emma could not quarrel with herself. He was so much displeased that it was longer than usual before he came to Hartfield again; and when they did meet, his grave looks showed that she was not forgiven. She was sorry, but could not repent. On the contrary, her plans and proceedings were more and more justified and endeared to her by the general appearances of the next few days.Mr. Knightley looked upon the proceedings at Hartfield with a displeasure he found increasingly difficult to conceal. To his discerning eye, the engineering of young minds was a grave responsibility, not a pastime for a fertile imagination; yet Emerald remained entirely unrepentant. When taxed with the folly of her interference, she merely elevated her chin, secure in the conviction that her motives were of the loftiest order.The first one is Human Written, and it’s from Emma by Jane Austen. Even though the average sentence length might be higher, you can see that the author diversifies the sentence length. Lines like “She was sorry, but could not repent” are short and punchy, while there are also longer, more detailed lines.This tendency to change up sentence length is called burstiness. It helps create rhythm, and it’s one of the most human things about how we write. We can see that AI is significantly lower when it comes to burstiness than humans.The lexical diversity of human writers is also lower than that of their AI counterparts. (Lexical diversity is the number of words that are only mentioned once, i.e., the unique words, divided by the total number of words.)This is something that has actually changed as models progressed. Earlier AI models like GPT-3.5 had lower lexical diversity, but GPT-4, GPT-5, and Llama 3 often have a more sophisticated vocabulary and vary a lot.While this can be rich and create interesting texts, it can also feel like purple prose. If a novice writer were trying to impress someone, they’d open up the dictionary and try to crowd their sentences with complex five-syllable words that feel unnatural. That’s purple prose and perhaps it best describes how AI writes.What is Purple Prose from ReedsyOn the other hand, humans tend to repeat words a little more, which can actually create rhythm or a musical quality in prose. It can reinforce key themes throughout a story.Word ClassNow that we’ve considered the sentences themselves, let’s talk about what those sentences are composed of: words.Across every piece of text, the AI-generated version uses more adjectives and nouns than the original human text. AI overdescribes the setting, and it floods the work with nouns (high nominal density).This might be because during reinforcement learning from human feedback, human raters tend to rate responses higher when they feel they are more thorough and detailed. The model, therefore, gets rewarded for describing the setting and the objects in a room. This is called “reward over-optimization”.This works fine in academic articles or essays, but in creative writing, that same habit means the story doesn’t move forward. Most seasoned writers can agree that creative writing is all about action.Adjectives and NounsThis might be why human speech and writing are more verb-centric, emphasizing relationships, movement, and actions rather than just labeling objects. Researchers like Leonard Talmy, who focus on cognitive linguistics, say that human language is fundamentally organized around events.If a human writer is trying to describe how the room is cold and grey, they won’t explicitly use nouns to describe that. They’d show someone shivering or cuddled in front of a bonfire, rather than use adjectives to describe the wind.I tried to combine what these individual trends show us into a text. This is the Descriptive Ratio, or the number of adjectives to the number of verbs. It shows us whether a piece of writing is fundamentally static or dynamic. A higher ratio means the writing is more static and descriptive, while a lower ratio means the writing is more dynamic.This is also why I’m optimistic about the future of writing. Human writing is so much better than AI because of this distinction between static and dynamic.“Show, Don’t Tell”We’ve all heard of the quintessential piece of writing advice plastered on every third-grade classroom: “Show, Don’t Tell”. Seasoned writers might also focus on making readers active participants in the story. Instead of stating facts directly, they provide hints that encourage readers to “read between the lines”.The latter feels more immersive, because it shows you what’s happening and let’s you infer from it. It also has four verbs: braced, shivered, rubbing, feel. The fist text has only two, which is the linking verb “was” repeared twice.DiversityI also calculated the adverb-to-verb ratio and the adjective-to-noun ratio and put them on a box and whisker plot. What I could see was that the differences in each AI text were minuscule; all the points were packed closely together. On the other hand, HW is spread apart.The technical name for this is model collapse. It’s when a model starts converging its responses towards a narrow set of safe, high-scoring patterns instead of producing variation.When you think about how several human writers rely on AI for writing, it concerns me. This is such a major loss of diversity in art and text.3. Most Frequently Used WordsNow, we come to a confusing but powerful way stylometry recognizes an author’s fingerprint.This is called Most Frequently Used Words. In a corpus, i.e., a set of texts, you can zoom in on the words that get repeated the most.fWhen you look at that list, it’s words like “the”, “of”, “she”, or “but”. They aren’t words that contain content matter but are the glue of a sentence.The reason this works better than other metrics is that an author can consciously control their plot, vocabulary, and subject matter. However, there’s no conscious decision to repeat a function word a certain number of times.MFW is the foundation for authorship attribution, going back to a statistician named John Burrows in the 1980s. It’s also exactly the technique that confirmed J.K. Rowling as ‘Robert Galbraith’.Burrows Delta helps turn MFW into a comparison number that tells us whether two texts are different or similar.It works in four steps.The algorithm isolates those 100 most frequent words across every single text in the whole dataset. This is for each and every text.2. For each of those words, in each text, it calculates something called a Z-score, which shows us whether that function word is used more or less than the average across the entire dataset.3. For any pair of texts, it compares their Z-scores word by word, and averages the difference. A smaller difference means the texts are similar.With 63 texts in my dataset, that means comparing every single text to every other text — 63 times 63 — which is 3,969 individual delta numbers.How can you see a pattern in 4,000 numbers?One way is through cluster analysis, which takes those distances and organizes them into a family tree.This is called a dendrogram — that branching tree structure you see on the slide.Cluster AnalsysisIt works from the bottom up: the two most similar texts get joined first, at the shortest possible branch length. Then the next most similar pair or group gets joined. This repeats, branch by branch, until every single text is connected into one tree, with branch length representing how different two groups are.Another way is MDS. Multi-dimensional scaling takes all of those pairwise distance numbers and places them on a space. Texts that are stylistically similar get placed close together. Texts that are stylistically different get pushed apart. The X and Y axis here means nothing in particular, they just creates 2D space to place the texts at.The dendrogram flattens some of that distance information into pure hierarchy, while MDS tries to preserve it spatially.Here are some of the most interesting results I’ve gotten:Look at the top three branches, all three human-written Shelley texts group together first, before joining anything else. That’s exactly what we’d hope to see: real Shelley writing recognized as most similar to other real Shelley writing.However, it joins up next with the Frankenstein-imitation texts, more closely than you’d expect. Shelley’s actual style and the AI’s attempt at her style aren’t as far apart as we might hope.When you look at neutral texts, and this is same across the baord, every single one of those AI ‘neutral’ texts clusters into one big group, completely separate from the three real human Austen texts, which form their own clean, distinct clade.However when you look a the right, Jane Austens excerpt from her book Sense and Sensibility doesn’t sit with her other texts, it nestles directly into the AI cluster. In this specific case, the imitation was good enough that the algorithm couldn’t tell it apart.ThoughtsSo can a machine write like a human?If we look at the data, the answer leans towards no.AI’s sentences are shorter like academic text, and it has low burstiness.LLM’s tend to over-describe settings with high nominal and adjectival density. They tell us about a static world rather than show us an active one.AI also tends to cluster tightly. Human writers express a vast and unpredictable spread of stylistic diversity, while AI models lean towards a safe narrow average.Our machines fail us. AI detectors are barely accurate, and half the time I suspect a well-worded email is probably just ChatGPT.I don’t think it’s pointless to create and write art. AI can’t see the world the way humans do, and human work will always be valued.This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!Can a Machine Write Like A Human? was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.Source: Generative AI Pub — Published — Category: Image AI