How Pangram Identifies AI Writing
It’s been decent progress on the book this week. I’m excited to be sharing some fascinating and diverse stories on framing (along with my overarching framework for the craft) with all of you soon.
While AI has been super useful in suggesting examples for inclusion, most of the case studies I have included are my own. Of course, the framework and writing are all 100% mine.
Thanks for reading The Story Rules Newsletter! Subscribe for free to receive new posts and support my work.
Welcome to the one hundred eighty-sixth edition of ‘3-to-1 by Story Rules‘.
A newsletter recommending good examples of storytelling across:
- 3 tweets, and
- 1 long-form content piece
Let’s dive in.
𝕏 3 Tweets of the week

Haha, nice use of contrast (technically, chiasmus).

Interesting observation by framing guru, Lulu Cheng Meservey of the post by Dario Amodei

As I posted, I’m happy to be taking some of the responsibility for this transition.
🎧 1 long-form listen of the week
If you’ve been online, it’s likely that you’ve come across Pangram, the AI detection software that identifies whether a piece of text is human or AI-generated and estimates the share of each.
Frequently, many polished and well-articulated pieces by famous people turn out to be 100% AI-generated (although there are minor chances of false positives).
In this illuminating conversation, Max Spero, Pangram’s young founder, speaks with Derek Thompson about AI writing, why Pangram is better than other tools at detecting it, and how we can cope in the age of AI writing.
Thompson’s opening monologue on why people have a negative perception of AI is superb:
Thompson: I’ve been doing shows recently about why so many Americans say they hate AI. It’s not just that AI chief executives have promised the technology will displace tens of millions of jobs. It’s not just that AI optimists variously claim it will destroy the world and/or cure cancer, like some kind of cosmic coin being flipped to determine the future of existence. It’s not just that its physical manifestation, the data centre, is a big ugly box. It’s not just that a lot of Americans don’t trust big corporations or government. And it’s not just that they think AI videos are corny or annoying or all four at once.
The reason I think a lot of people just don’t like AI is so obvious that it is literally the first word – the A, for artificial. People don’t like AI because it is so often used to make things that are fake. Fake inspirational monologues on LinkedIn, fake posts on Twitter, fake photos on Instagram, fake videos on TikTok.
Spero shares that already 50% of the internet is bots, and soon it might be 99%!
Spero: Obviously humanity is not going to go extinct if we can’t tell AI writing from human writing. But what we are doing is trying to save the internet. We’re trying to save the written word and communication. If you look at the scale of AI and bot traffic on the internet, it’s gone from a small fraction of human traffic to recently passing 50/50. So about 50% of internet traffic is bots. And I think it’s not that long before it’s going to be 99% of internet traffic is bots. And not just readers either, but writers, communicators, agents that are trying to get things done on behalf of their owners.
They call it the ‘Dead Internet theory’ and mull over whether there is any value when someone says that ‘Person X wrote something’:
Spero: We’re very much at risk of this idea of dead internet theory, where the internet is just this echo chamber of bots talking to bots, and it’s going to be impossible to parse out the signal from the noise when there’s so much AI chatter.
Thompson: I’m worried about the relationship between humans and humans. Something happens to the way that you can speak to another person on the internet authentically when you plausibly believe that their output might just be ChatGPT or Claude. We’ve historically had this association whereby if it’s written, then a human wrote it. If it’s written and signed by Derek Thompson, therefore Derek has in his head an understanding of what he wrote. But this idea that everyone suddenly has an AI ghostwriter means we don’t know who we’re dealing with when we read certain posts – and that scrambles and weakens the relationship between people on the internet, which further deadens the dead internet theory.
The New York Times used a strong line to describe AI slop:
Thompson: Just this morning, the New York Times reported that Spotify, LinkedIn and others are ‘trying to dig out of a digital sewage heap full of low-quality content made by artificial intelligence.’
Spero then confirms the high share of online content which is now AI-generated:
Spero: Our study found that 29% of long form content – so 250 words or more – on Twitter was AI generated. And then of short form content, 50 to 250 words, it was about 9%.
If you look at the internet at large, we did this study with Stanford and the Internet Archive which found that in May of last year it was already 40% of the internet that was AI generated. And I think this number is just going to keep expanding. There are people trying to pollute the internet for different reasons – whether they want their own narratives in the LLM training data, or they’re trying to win at SEO.
Spero: It’s not just the social platforms – they obviously have this big problem.
LinkedIn is of course the worst offender, and Spero theorises it’s because of the incentives:
Spero: LinkedIn had 41% of their long form content AI generated…
Why is LinkedIn more full of AI content than Twitter? I think it’s because there’s this incentive that people feel – posting on LinkedIn regularly is going to improve my career prospects, it’s going to make me more likely to get a good job because I’m going to be visible to people who are relevant in my career. And obviously if posting more has an advantage to your career, then people are going to find the easiest way to post more, which ends up being AI.
Pangram’s detection is reckoned to be the best:
Spero: There’s a really good study by the University of Chicago where they tested a whole bunch of AI detectors, some open source, some commercial, and they found that Pangram was the best by a large margin. They tested almost 2,000 human texts and found zero false positives.
A fascinating insight: about how telltale AI signs keep migrating “up a level”—from the word to the sentence to the paragraph:
Spero: Early ChatGPT and AI models were fairly biased. They were trained on this instruction-tuned data set and they would way overuse some words and phrases – delve, tapestry, intricate. They really loved these individual words. So if you saw the word delve in an out-of-place setting, that was an immediate red flag. And then maybe a year and a half later they were able to hammer out these word-level inconsistencies, but there were still patterns showing up. For example the ‘it’s not just X but Y’ pattern – it’s called negative parallelism, and LLMs love it because it has a lot of impact on the reader. Or sentence constructions using em dashes, because these were also associated with good writing.
Now I think they’ve even hammered some of these out. And so how I tell is I’m relying on longer context signals – at the paragraph or sentence level. I can look at the shape of the text and tell you, oh, that looks like it’s ChatGPT or Claude.
Thompson: … Something I’ve noticed is that long pieces of AI writing often try to summarise and resummarise and resummarise. So now the tell is at the scale of the paragraph. It’s like – this is genuinely important, this is the main course, the bottom line is. Every sentence is trying to outdo the previous one in terms of summarising what it’s trying to say better.
The reason for some of those dramatic AI tells (‘Here’s the most important part’, ‘Here’s the part no one tells you’) is that AI is like a keen-to-impress puppy, trying to make the master happy:
Spero: AI wants to try and make every sentence it says the most impressive sentence. It’s just really trying to impress the reader.
Thompson: There’s a tall poppy syndrome thing going on with AI writing, where every sentence is trying to stand out taller than the one that came before it. And it drives me crazy as a lover of language, because that’s not how good writing works at any appropriate length.
The underlying reason for that is RLHF (the puppy getting treats for getting stuff right):
Spero: AI models are typically trained in two stages. The first is pre-training, when it’s predicting the next token – looking at internet data and saying, I think this word comes next. The second stage is reinforcement learning, where it’s already good at predicting the next token and instead we’re giving it a new reward function – give it a reward if it’s judged as producing good writing or the correct answer.
This reinforcement learning basically turns up a positive signal to 11. People, when they read an answer and feel that it was a little bit impressive, they’re like – oh, this is meaningful, and then it summarised it. You take that reward signal and then you ratchet it up. And eventually what it’s doing is summarising itself over and over, because that’s how it knows it can make its answer the most legible to the human. It’s making these sentences really impressive and trying to make them meaningful, because it knows it’ll get rated as a better writer if it does this.
Thompson relates a story from his early Atlantic days on the need to reduce signposts:
Thompson: When I was writing my first long features for The Atlantic, my first cover stories for the magazine, I was working with an editor named Don Peck. On my first drafts he said, ‘You’re signposting too much early in the essay – trying to bang people over the head with this is why this essay is important, this is why this anecdote is important, these are the six things I’m going to explain to you.’ He said that might make sense at the scale of a thousand words, but in a 7,000-word essay, readers want to go on a bit of a journey. They want to get a little bit lost. They want you to tell them a story and leave the candy at the end of that story – and then there’s a grand reveal about, oh no, this is what the story means. It’s not what you thought it meant.
Writing going from signpost, signpost, signpost – this is what I’m saying, this is what I’m saying – to: this is a story, you’re walking down a path and you don’t know where that bend in the road is headed. That’s long form writing. And that to me is what AI is terrible at. There’s no sense of that talented novelist or creative non-fiction writer’s sense of – I’m going to take you on a story here, and you’ll only understand its import by the time we get to the end.
Spero’s explains that AI’s training set would have had very few such 7,000-word posts:
Spero: It’s simply not trained for this sort of task. And maybe even if it is, it doesn’t have enough training data. The number of 7,000-word articles is very small. The number of novels is also relatively vanishingly small compared to internet posts.
A dark thought akin to a pincer movement—even as AI tries to sound more human, humans sound more like AI:
Thompson: The less obvious and maybe just as plausible way this could happen is that humans get worse at writing like humans. We write more and more like AI. And you have this pincer movement where the convergence of AI and human writing ultimately closes the gap between what is identifiably AI writing and what is identifiably human writing.
Spero: I think human language does shift. I don’t know if you’ve ever felt yourself saying something that feels like an AI-ism. I just said ‘AI-fy’, which is not even a word – so maybe even that is an AI-ism. Obviously human language is changing, but it changes at such a long scale. The scale at which language changes for humans happens on a much longer time horizon than AI text changing. So we should probably only be looking at the AI text, because that’s what’s changing really rapidly.
With AI getting better and better, will we care if something is human-authored? Both feel we will:
Spero: My argument is we still are going to care a ton. The internet’s going to be 99% bots and 1% humans. I’m going to care that I can get information I couldn’t get from Claude or ChatGPT. And to do this I need to be hearing from a real human, not a bot.
Thompson: In general we are, in very ordinary ways, impressed by activities that we know robots can do better than humans – but we’re only impressed when the human can do it. Take pitching in baseball. Can a robot throw a ball faster than 105 mph? Well, a robot can launch a rocket at 100,000 mph, so it could certainly throw a ball faster than 105. But it’s not interesting. No one wants to watch a robot pitch for the Milwaukee Brewers or the Pittsburgh Pirates. They want to watch Paul Skenes. They want to watch the human throw the fastest fastball ever thrown by a human. That’s the only thing that has athletic or artistic value.
Spero: I look at the art world a lot. The camera came out, photography was invented, and previously the closest thing you could have to a faithful reproduction was a painting – and then suddenly a photograph can do better along basically every axis. And what happened? Painting was not completely supplanted by photography. Instead, artists found different ways to create.
That’s all from this week’s edition.
Photo by Diana Polekhina on Unsplash