Flawed Oversight: Teachers Don’t Catch AI’s Grading Mistakes When AI is Harsh

AuthorsRigissa Megalokonomou, Sofoklis Goulas, Panagiotis Sotirakopoulos
PublishedJune 2025
PublisherElsevier
DOIhttp://dx.doi.org/10.2139/ssrn.5294732
Number of Pages66

Using a randomised experiment, we test how educators revise grades labelled as AI- or human-generated.

Every teacher saw the same student work and an erroneous 5/10 mark; we varied (i) whether the correct answers caused the score to be harsh or lenient and (ii) the stated source of the mark.

The outcome, the grading fairness gap, is the distance between teachers’ revised marks and the objective grade.

Under a harsh recommendation, the gap was 22 per cent larger with an AI label; in the lenient case, the fairness gap under AI and human labels was statistically indistinguishable.

Mediation analysis shows that, in the harsh case, higher attributions of ability and responsibility to the algorithm transmit no less than half of the effect, whereas weaker attributions in the lenient case trigger stricter corrections.

Thus, acceptance of algorithmic advice hinges on error direction and inferred credibility, rather than on artificial intelligence alone.