18 frontier models scored by Human Score, how much each writes like a person and not a chatbot. It blends the live arena vote (40%) with four mechanical axes against pre-AI writing. Higher is better.
loading the crowd vote…Human Score, 0–100, most human first. Switch domains to see who slips where - the model that writes clean email is not always the one that writes clean essays. Every board blends the same way: 40% the arena votes cast on that domain’s pairs, 60% the machine measurement against the pre-AI baseline.
Ranked by Human Score. Each axis is on a 0–100 scale where 100 = writes like a human, so you can see exactly where a model wins or loses. Toward the middle the scores run close.
| # | model | human score | arena 40% | concise | templating | rhythm | tells | $/M |
|---|
Arena = the crowd vote (live Elo, 40%) · Conciseness = resisting length inflation vs a human on the same task · Templating = avoiding reused openers & skeletons across prompts · Rhythm = natural paragraph-length variance · Tells = avoiding over-used AI words. Each 0–100; 100 = human-like.
Blended price per million tokens against the Human Score. If cost bought a more human writer, this would trend up and to the right. It doesn't.
The single biggest input is people. This is the live Elo from the arena, where players flag the sloppier of two blind samples. Least-flagged rises; it updates the board above as votes land.
No LLM judges scoring other LLMs. Every mechanical number is measured against writing that provably predates ChatGPT. The full harness, scenarios and all 19,928 raw generations are on GitHub →
Slop is measured as distance from genuine human writing collected before generative AI existed: Enron email, a blog-authorship corpus and student essays, archived tweets, and Discord chat.
The arena 40% plus conciseness, templating, rhythm
and tells at 15% each. Every axis is normalised to a 0–100 human-likeness scale before blending.
Every model got the identical 112 scenarios across four domains, five samples each, for 19,928 real, unedited generations. Default settings, no cherry-picking.
The mechanical index measures slop by rule; the arena measures it by vote. Where they agree, the signal is strong. Full methodology →
The board is the machine's verdict. The arena collects the crowd's, blind, one pair at a time. Play a round and see if humans agree.
Play “Spot the Slop” → View the code & data →