Creative Writing Benchmark Finds Frontier AI Models Near Amateur Human Writers but Behind Professionals

news
Creative Writing Benchmark Finds Frontier AI Models Near Amateur Human Writers but Behind Professionals

A new creative writing benchmark comparing human writers with 24 large language models suggests that the strongest AI systems can compete with amateur writers, but professional human writers still maintain a clear advantage.

The test used 475 long form writing prompts and evaluated the results with a reward model trained on human preference data. The benchmark measured predicted reader preference rather than objective writing quality, so the scores should be treated as one evaluation framework rather than a universal measure of writing ability.

GPT 6 Astra recorded the highest overall predicted win rate at 87.8 percent, slightly ahead of the benchmark's combined human writer score of 86.6 percent.

However, the human category included both amateur and professional writers, and the professional group scored substantially higher when separated.

Frontier models performed far better than smaller systems

The strongest AI models dominated the upper part of the benchmark.

GPT 5.6 Sol reached a 77.6 percent predicted win rate, while Claude Fable 5.1 recorded 70.3 percent.

Performance dropped significantly further down the leaderboard.

Writer or modelPredicted win rate
GPT 6 Astra87.8 percent
Human writers overall86.6 percent
GPT 5.6 Sol77.6 percent
Claude Fable 5.170.3 percent
Many other large modelsAround 50 to 65 percent
Qwen3.8 27B23.2 percent
DeepSeek V4.1 Flash19.8 percent
Gemma 4 26B10.9 percent

Models such as Claude Opus 5, Kimi K3 and Grok 4.6 reportedly landed in the roughly 50 to 65 percent range.

The results suggest that creative writing quality varies widely between model families and sizes, rather than improving uniformly across all modern LLMs.

Professional writers still hold an advantage

The benchmark's most important distinction is between amateur and professional human writers.

Although GPT 6 Astra slightly exceeded the combined human baseline, professional writers remained noticeably ahead when measured separately.

That matters because combining different skill levels into one human score can make the comparison appear closer than it is.

Professional writers are often better at maintaining narrative structure, character consistency, pacing and stylistic control across long passages.

Those skills become increasingly important in multi chapter writing, where weaknesses in coherence can become more visible over time.

Long form coherence remains difficult for smaller models

One recurring weakness identified in discussions around the benchmark was loss of coherence during longer prompts.

Smaller models were more likely to repeat phrases, lose track of earlier events or allow story structure to drift across multiple chapters.

Human writers generally have more flexibility to adjust tone, pacing and narrative direction during the writing process.

The benchmark also found a large difference in output length.

Human writers produced an average of 2,592 tokens per prompt.

Claude Fable 5.1 generated around 2,114 tokens, while GPT 6 Astra produced about 1,537.

Output length alone does not indicate quality, but it does show that human participants tended to develop their responses more extensively.

The benchmark measures predicted preference, not absolute skill

There are important limitations to the results.

The scoring system relied on a reward model trained to predict human preference rather than direct evaluation by a large group of readers for every response.

That means the percentages do not prove that one writer or model is universally better.

They show how the benchmark's evaluator expected readers to respond within this specific test.

Creative writing is also highly subjective. A model may perform well in one genre or prompt style while struggling with another.

Still, the benchmark provides useful evidence that the gap between the strongest AI systems and average human writers has narrowed substantially.

At the same time, it also shows that professional writing remains a much harder target, particularly when a task requires sustained coherence, style control and complex storytelling across long passages.

Discover: News

Discussion (0)

Be the first to comment.