Back

AI-Written Short Stories Often Rated Higher Than Human Works in Blind Tests

At a glance

  • Multiple studies assessed reader preferences for AI and human-written stories
  • AI-generated stories were often rated higher than human-authored ones
  • Readers generally struggled to distinguish between AI and human writing

Recent research has examined how readers perceive short stories created by artificial intelligence compared to those written by humans, using blind tests and controlled studies to evaluate preferences and identification accuracy.

One preprint study involved 1,498 adult participants who read either human-written or AI-generated short stories and rated their quality and how engaging they found them. In this study, AI-generated stories received higher ratings for both quality and absorption, regardless of whether readers knew the origin of the text.

Further experiments from the same research included 905 adults who read both types of stories and attempted to identify whether each was written by a human or an AI system. Results showed that participants were no better than random chance at correctly identifying the source of the stories.

Additional analysis from the preprint indicated that participants who reported greater familiarity with AI were more likely to accurately identify the origin of the stories, while expertise in fictional literature did not show a similar effect.

What the numbers show

  • 1,498 adults rated story quality and absorption in Study 1
  • 905 adults attempted to identify story origin in Studies 2 and 3
  • Hungarian study reported 66% overall identification accuracy for AI-authored texts

An independent study tested 20 participants with three short stories, two written by ChatGPT-4o and one by a well-known Italian author. In this blind comparison, the AI-generated stories received slightly higher average ratings and were chosen more often, though the difference was small.

A sociolinguistic study in Hungarian with 576 respondents found that participants tended to prefer texts they believed were written by humans. However, participants in this study were more likely to correctly identify AI-authored texts than human-authored ones, and the overall accuracy rate for identifying AI texts was above chance.

In a separate blind test organized by a fantasy author, readers evaluated flash fiction without knowing the authorship. According to the report, readers could not reliably distinguish between AI-written and human-written stories and tended to rate the AI-written stories higher than those by award-winning authors.

Across these studies, findings consistently showed that readers often rated AI-generated stories as highly as, or higher than, those written by humans, while also having difficulty reliably telling the difference between the two types of authorship.

* This article is based on publicly available information at the time of writing.

Sources and further reading

Note: This section is not provided in the feeds.

Related Articles

  1. A humanoid robot finished the Beijing E-Town Half-Marathon in 50 minutes and 26 seconds, breaking the men's human world record, reports say.

  2. Debate continues over human spaceflight's value, with ISS costs at $150 billion raising questions about scientific returns and future missions.

  3. A new AI Task Force was launched in June 2025 to address governance issues, according to the Council on Criminal Justice.

  4. Current AI lacks consciousness, generating text through statistical predictions, creating an illusion of sentience, according to expert consensus.

  5. The House passed a defense bill restricting humanoid robots from China, Russia, and Iran, reflecting national security concerns, according to reports.

More on Technology

  1. The Rubin Observatory's new COSMOS field image showcases over 500,000 galaxies and 50,000 stars, available through the Rubin Science Platform.

  2. Chinese AI labs have released open-weight models. Moonshot AI raised $2 billion, achieving a valuation of $20 billion, according to reports.

  3. Quantum computing remains complex, with practical systems still years away. Current hardware features 50-1000 noisy qubits, according to experts.

  4. The Friend pendant and OpenAI's screenless speaker highlight a growing trend in personal AI devices, emphasizing seamless, unobtrusive communication.

  5. A map details 116,084 ALPR cameras across the US, with Flock Safety devices comprising over 82%, according to community data.