This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: New research from Cardiff University and the University of Melbourne has investigated whether GenAI could mimic human marking when evaluating written student work. The results have been published in the journal Assessment & Evaluation in Higher Education.
William Kay of Cardiff University's School of Biosciences said, "The rapid rise of generative artificial intelligence—known as GenAI—has promoted interest in whether it can support the evaluation of student work in higher education. We wanted to understand whether large language models (LLMs) could mimic human assessment of extended written assignments well enough to guide students in judging the quality of their work." To test the capabilities of LLMs in evaluating student work, the researchers used two versions of a popular GenAI platform, ChatGPT, to mark 50 undergraduate bioscience essays. They asked ChatGPT to mark the essays against seven assessment criteria and applied four different prompting conditions.
LLM-assigned marks were compared with human marks, evaluating both average scores and mark variability. "Our findings indicate that marks awarded by LLMs varied considerably and were inadequate predictors of the human marks awarded to essays," Kay said. "We found significant discrepancies between GenAI-marked essays and those marked by humans.
Overall essay marks were relatively similar when evaluated by humans and GenAI, but when assessing student performance based on individual marking criteria and on an essay-by-essay basis, the differences between LLM- and human-assigned marks were substantial." In all but one case, LLMs typically returned higher average marks than humans. The largest difference in the average mark awarded by any LLM compared with a human-awarded mark was 16.1 marks, and 40 marks at the individual essay level. "We also observed that LLMs reduced the marks for high-scoring essays and inflated them for low-scoring work, resulting in systematic compression of marks toward the middle.
"The findings of this study highlight that, at present, GenAI is unable to reliably assign a mark to a subjective piece of written work comparable to that of humans—even with extensive training of the LLM," Kay said. "At present, the LLMs tested are not suitable alternatives to human tutors for providing individual students with predicted grades on their work. "While there is interest across the sector in whether the pattern-recognition capabilities of LLMs could facilitate objective grading of students' work, making marking more efficient and relieving pressure on staff, the findings of this research indicate that at present this is not advisable.
Aside from the ethical issues of submitting student work to GenAI tools without express consent, LLMs cannot and should not be relied upon to assign grades to students' extended written work. "As LLMs become more sophisticated, it is possible that their ability to mimic human judgment may improve in the future. But, as we find in this study, aligning marks between humans and GenAI may be hard to achieve." William P.
Extract — continue reading at the source.