Study reveals AI grades essays higher than humans do
A 2026 Cardiff and Melbourne study finds generative AI grading is inconsistent with human markers
One essay scored 40 points apart depending on who or what was marking it. That gap, on a 100-point scale, is the starkest finding from a new study by researchers at Cardiff University and the University of Melbourne, published in Assessment & Evaluation in Higher Education, testing whether ChatGPT can reliably grade student writing the way a human academic does.
The team uploaded 50 undergraduate bioscience essays to two versions of ChatGPT, grading each one against seven assessment criteria under four different prompting setups.
Every AI-generated score was then checked against the mark a human grader had already assigned the same essay. In nearly every case, the AI models returned higher averages than their human counterparts.
Lower-scoring essays saw the largest jump in AI-assigned marks, while stronger essays were sometimes marked down relative to human grading. Only in the middle of the scale, where essays humans considered average, did AI scores line up reasonably well with human judgement.
The models showed reasonable stability while marking the same essay twice using the same prompt, but not while marking a variety of essays of varying quality, which, according to the authors of the study, renders the scoring patterns of ChatGPT inconsistent and unable to reliably predict human markers' marks.
From Cardiff University, one of the study co-authors William Kay, stressed that the results proved the need to keep the responsibility of marking papers with humans rather than language models. He mentioned the use of vague descriptors for grading, such as "good" or "excellent", rather than specific and distinctive criteria, as one of the reasons why it is difficult for the models to mark essays consistently.
In addition to the mentioned challenges, there are ethical issues related to the use of AI tools to grade students' work that have not been addressed by universities yet.
-
Gemini Live now runs tasks while you talk: Here’s what it can do
-
OpenAI engineer quite coding for filmmaking: Here's why
-
Experts split on Meta's $18bn teen safety settlement
-
What’s changing on Instagram, Facebook for teens after Meta’s $17B settlement?
-
Nvidia to buy Hugging Face for $12.9bn: Report
-
Brazil sues discord for $97M over child safety failures and lack of protection
-
Meta agrees to $18B settlement with 29 US States over youth social media addiction lawsuit
-
YouTube Music may be getting Apple's Liquid Glass look
-
Meta's secret plan to cut 60% of some teams with AI: Report
-
Study reveals 37% of German firms link hacks to spies
-
Anthropic IPO can break SpaceX's Wall Street record: Here's how
-
ChatGPT Work can log into websites to complete tasks for you: Here's how
-
Meta trial: Instagram CEO says few teens used ‘Take a Break’ safety feature
-
Meta, US States weigh mid-trial settlement in landmark teen addiction lawsuit
-
OpenAI’s Jalapeno chip beats Nvidia in efficiency tests: Here's how
-
WhatsApp rolls out stronger account security features: What users need to know
-
Brazil takes major action against TikTok owner ByteDance over teen data
-
Meta weighs hardware 'Kill Switch' for smart glasses amid covert filming backlash