AI-generated stories rated higher quality than human-written ones, study finds
A study finds readers rate AI-generated short stories as more absorbing and higher quality than human-written ones, although stories labeled as human-written received higher ratings. Additional experiments suggest people struggle to tell the difference between the two.

A new study suggests readers may find short stories produced by artificial intelligence more absorbing and of higher quality than those written by people. The research, published in the journal Judgment and Decision Making, involved 1,682 adults who each read one of six stories – three written by humans and three by ChatGPT. The AI stories were designed around themes similar to the human-authored pieces. Participants were told whether a story came from a person or an AI, but the labels were not always accurate.
Ratings showed that people who read AI-generated stories gave them higher marks for absorption and quality than those who read human-written stories. At the same time, stories labelled as human-written received higher ratings regardless of their true origin. The authors suggest attitudes toward AI play a role: participants with a more favourable view of AI rated ChatGPT-labelled stories higher than those with a less positive attitude.
Dr Deena Skolnick Weisberg, senior author from Villanova University, said AI systems can already produce short stories considered at least as good as human ones, and that views of AI abilities should be updated. However, as a creative writer herself, she argued this does not mean writing should be left to machines. AI-generated novels may require space alongside human and co-written works, but human creativity still matters. People write for self-expression, challenge and understanding, she said, and AI's ability to imitate human stories does not change that.
Two additional experiments gave 905 adults pairs of stories on the same theme, one human and one AI, and asked them to identify the human author. In one experiment about 40% guessed correctly, which is worse than chance; in the other, 52% did. The researchers concluded there was no general evidence that people could reliably distinguish between the two. Familiarity with fiction did not help, but greater experience with AI systems was associated with better guesses.
Weisberg noted AI writing may have a recognisable style that improves with practice. She added that the preference for human-labelled stories likely reflects a desire for authenticity, but readers actually favoured AI stories because they are easier to read and digest.
Luke Kennard, a professor at the University of Birmingham who was not involved, was sharply critical. He said AI is often imagined as a neutral robot, but it is really predictive code trained on a massive database of human writing, running in resource-hungry data centres with harmful effects on local communities. He rejected the idea that AI is just another tool, calling it an existential threat. While AI can produce coherent and even enjoyable poems, he said that is unsurprising given its training. The real question, he argued, is whether it is worth it – and his answer was a categorical no.


