Technology

Study reveals AI grades essays higher than humans do

A 2026 Cardiff and Melbourne study finds generative AI grading is inconsistent with human markers

Published August 27, 2026
Study reveals AI grades essays higher than humans do
Study reveals AI grades essays higher than humans do

One essay scored 40 points apart depending on who or what was marking it. That gap, on a 100-point scale, is the starkest finding from a new study by researchers at Cardiff University and the University of Melbourne, published in Assessment & Evaluation in Higher Education, testing whether ChatGPT can reliably grade student writing the way a human academic does.

The team uploaded 50 undergraduate bioscience essays to two versions of ChatGPT, grading each one against seven assessment criteria under four different prompting setups.

Every AI-generated score was then checked against the mark a human grader had already assigned the same essay. In nearly every case, the AI models returned higher averages than their human counterparts.

Lower-scoring essays saw the largest jump in AI-assigned marks, while stronger essays were sometimes marked down relative to human grading. Only in the middle of the scale, where essays humans considered average, did AI scores line up reasonably well with human judgement.

The models showed reasonable stability while marking the same essay twice using the same prompt, but not while marking a variety of essays of varying quality, which, according to the authors of the study, renders the scoring patterns of ChatGPT inconsistent and unable to reliably predict human markers' marks.

From Cardiff University, one of the study co-authors William Kay, stressed that the results proved the need to keep the responsibility of marking papers with humans rather than language models. He mentioned the use of vague descriptors for grading, such as "good" or "excellent", rather than specific and distinctive criteria, as one of the reasons why it is difficult for the models to mark essays consistently.

In addition to the mentioned challenges, there are ethical issues related to the use of AI tools to grade students' work that have not been addressed by universities yet.

Pareesa Afreen
Pareesa Afreen is a reporter and sub editor specialising in technology coverage, with 3 years of experience. She reports on digital innovation, gadgets, and emerging tech trends while ensuring clarity and accuracy through her editorial role, delivering accessible and engaging stories for a fast-evolving digital audience.