Graphite 研究显示 Opus 5.5 高频使用 this matters 等写作痕迹
Opus 5.5 loves to tell you ‘this matters’ (and other AI writing tells)
Graphite 的研究发现,Claude Opus 5.5 使用 this matters 的频率是人类写作的 116 倍,why X matters 为 92 倍,单词 dependable 则高出 23 倍。
Now that LLM-generated prose is everywhere, human beings are eager for ways to sniff it out. While early tells like em-dashes and “delve” are long gone, researchers say there are still plenty of telltale habits that AI models fall back on when writing prose.
A new study from the marketing firm Graphite looked at the writing habits of frontier models, sussing out each model’s favorite words and phrases. While old tells like em-dash use have been stamped out, models still fall back on contrast-heavy constructions, with each model version showing its own unique quirks. The biggest surprise is how broad the scope of tells turns out to be. Graphite found 13,000 phrases that were at least twice as common in the AI content as human content — their definition of a “tell.”
“It turns out that Claude models are actually getting closer to the human word distribution over time,” Graphite’s chief AI officer Greg Druck told TechCrunch. “And for the GPT models, it’s getting further away.”
Studying AI-generated writing at scale required a careful study design. Graphite started with a corpus of 10,000 articles published before the release of ChatGPT, serving as the human-generated control group. Then researchers had different AI models rewrite the articles from summaries, hoping to eliminate as much source bias as possible. With matching samples from both humans and each model, they could compare how often certain words and phrases appeared in AI writing, as well as broader patterns in sentence construction.
According to Graphite’s results, Claude Opus 5.5’s biggest tell is the word “dependable,” which pops up 23 times more often than in human samples. While Opus 5.5 now avoids the “it’s not X, it’s Y” sentence construction, it still tends to say something “is more than an X, it’s a Y.”
Above all, Opus loves to tell you why things matter, using the phrase “this matters” 116 times more often than human writing, while “why X matters” occurs 92 times more often.
OpenAI’s Astra has a different set of tip-offs. This model loves to describe “another dimension” of whatever it’s talking about, and tends to hedge claims by saying an action “may provide” or “can provide” a particular benefit. Its biggest tell is what Graphite calls the “corrective framing,” where a topic is defined as “not simply X” or offered as an alternative, “rather than relying on X.” According to graphite’s research, those constructions were more than 100 times more common in Astra-generated prose than in human writing.
Notably, all the frontier labs seem to have responded to the idea that models overuse em-dashes. In Graphite’s samples, Opus 5.5 used the punctuation mark 99% less often than Opus 5. Astra now uses it 88% less than human samples, whereas Gemini 3.1 Pro has almost completely eliminated the em-dash from its writing.
But while individual tells change, Graphite says the overall number is mostly holding steady. “It’s not like the tells are decreasing,” Druck told TechCrunch. “They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own.”
It’s surprising that tells are so persistent, given the labs’ focus on human-like writing styles. In the Opus 5.5 release, Anthropic boasted that the model “communicates more naturally than prior models,” saying early users “found its writing clearer and easier to follow.”
OpenAI made similar claims when releasing the GPT-6 versions of Sol and Luna, saying users could “expect to see more clarity, less jargon, [and] fewer odd turns of phrase.”
But Druck is skeptical about how much the labs can do to completely eliminate telltale construction or phrases.
“A general hypothesis I have is that the labs are less able to control some of these things than you might expect,” Druck says. “These are giant models with billions of parameters. They have some finite number of tests they can run, and things slip through.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
来源:TechCrunch · AI · techcrunch.com