The board · live

Every model, ranked

18 frontier models scored by Human Score, how much each writes like a person and not a chatbot. It blends the live arena vote (40%) with four mechanical axes against pre-AI writing. Higher is better.

loading the crowd vote…
The board

Who writes the least slop?

Human Score, 0–100, most human first. Switch domains to see who slips where - the model that writes clean email is not always the one that writes clean essays. Every board blends the same way: 40% the arena votes cast on that domain’s pairs, 60% the machine measurement against the pre-AI baseline.

more humanmore slop

Full table

Every model, every axis

Ranked by Human Score. Each axis is on a 0–100 scale where 100 = writes like a human, so you can see exactly where a model wins or loses. Toward the middle the scores run close.

#model human score arena 40%concisetemplatingrhythmtells $/M

Arena = the crowd vote (live Elo, 40%) · Conciseness = resisting length inflation vs a human on the same task · Templating = avoiding reused openers & skeletons across prompts · Rhythm = natural paragraph-length variance · Tells = avoiding over-used AI words. Each 0–100; 100 = human-like.


Price vs quality

Does paying more buy a more human writer?

Blended price per million tokens against the Human Score. If cost bought a more human writer, this would trend up and to the right. It doesn't.


The human axis · 40% of the score

What the crowd thinks

The single biggest input is people. This is the live Elo from the arena, where players flag the sloppier of two blind samples. Least-flagged rises; it updates the board above as votes land.

loading the crowd vote…

How it works

Open method, open data

No LLM judges scoring other LLMs. Every mechanical number is measured against writing that provably predates ChatGPT. The full harness, scenarios and all 19,928 raw generations are on GitHub →

Human baselines, pre-Oct-2022

Slop is measured as distance from genuine human writing collected before generative AI existed: Enron email, a blog-authorship corpus and student essays, archived tweets, and Discord chat.

Five axes, one score

The arena 40% plus conciseness, templating, rhythm and tells at 15% each. Every axis is normalised to a 0–100 human-likeness scale before blending.

Same task, real output

Every model got the identical 112 scenarios across four domains, five samples each, for 19,928 real, unedited generations. Default settings, no cherry-picking.

Two independent lenses

The mechanical index measures slop by rule; the arena measures it by vote. Where they agree, the signal is strong. Full methodology →

Think you can spot the slop?

The board is the machine's verdict. The arena collects the crowd's, blind, one pair at a time. Play a round and see if humans agree.

Play “Spot the Slop” View the code & data