The Claude Opus 5.5 writing style has moved measurably closer to how humans write, with new analysis showing a near-elimination of em dashes, shorter sentences, and a sharper drop in semicolon use compared with its predecessor. The catch, if you want to call it one, is that the model has become more verbose in the process.
Arena, an AI benchmarking tool, analysed high-reasoning Text Arena responses from August and September 2026 and found that 10 of 12 writing measures moved in what it considers a better direction for Opus 5.5 versus Opus 5. The headline figure is the em dash: Opus 5 used 15.2 per 1,000 words; Opus 5.5 brings that down to just 0.8, a reduction of roughly 95%. Anyone who has spent time trying to spot AI-generated text will recognise that particular tic immediately, and it has been one of the more reliable tells for a while now.
What the Claude Opus 5.5 Writing Style Data Actually Shows
The em dash collapse is the attention-grabbing number, but the semicolon data is almost as dramatic. According to HackerNoon, semicolon usage fell from 6.10 to 1.64 per 1,000 words, a drop of about 73%. Semicolons are not quite the AI cliché that em dashes became, but they do cluster in outputs that are trying to sound considered and formal. Fewer of them, combined with shorter sentences, pushes the prose toward something more conversational.
On sentence length, Opus 5.5 averages 10.03 words per sentence against Opus 5’s 12.14. That 17% reduction matters in practice: shorter sentences tend to carry less of the nested-clause structure that makes AI prose feel laboured. Simpler wording follows from the same tendency.
The word-distribution comparison goes further still. VentureBeat reports that Opus 5.5 scores 0.052 against a human writing sample on a word-distribution divergence measure, down from 0.064 for Opus 5, a 19% reduction. In plain terms, the model’s vocabulary choices are getting closer to the distribution a human writer would produce. That is a different kind of metric from surface-level punctuation counting, and arguably a more meaningful one: you can train a model to drop em dashes deliberately, but pulling the overall word distribution toward human norms suggests something deeper is shifting.
More Words Per Answer, Fewer Obvious AI Patterns
There is a tradeoff, though. Average response length increased from 453 to 481 words, making Opus 5.5 the longest-writing model in Arena’s comparison. Shorter sentences, longer answers: the model is packing more of them in. Whether that counts as a problem depends on the use case. For someone drafting copy or trying to avoid AI-slop content, more natural-sounding sentences probably matter more than word count. For a task with hard output constraints, it is worth knowing the model now runs a little long by default.
Separately, Anthropic notes that at its lowest effort setting, Opus 5.5 beat Opus 5 running at high effort on BigFinance Bench, while using about 60% fewer output tokens. That efficiency gain is on the reasoning side rather than the prose side, but it sits alongside the writing changes as evidence that Opus 5.5 is not simply Opus 5 with a few prompting nudges applied. The underlying behaviour has been retrained, not patched.
It is worth remembering that most of the AI industry’s model-development attention over the past year or so has pointed squarely at coding benchmarks. Watching Anthropic invest this much in how Claude actually writes (at a measurable, statistically tracked level) is a slightly different priority signal. The internet’s em dash surplus has a new adversary, and the data suggests it is losing.

