The Comparative Effectiveness of AI Tools in Enhancing Feedback Quality and Efficiency in Undergraduate Mathematics Assessments

Authors

  • Shun Ling Chiang Department of Mathematics, Hong Kong Baptist University, Hong Kong

DOI:

https://doi.org/10.33422/worldte.v5i1.1999

Keywords:

AI grading, Mathematics Education, Hand-written Mathematics, Higher Education, Teaching Efficiency

Abstract

High-quality, timely feedback is essential for students’ learning in mathematics; yet large class sizes and heavy workloads for teachers often result in students receiving brief, generic comments that fail to address misconceptions effectively. While artificial intelligence tools show promise for improving grading efficiency and feedback quality in STEM education, empirical evidence specific to undergraduate mathematics remains limited. This study conducted a rigorous comparative evaluation of two AI platforms — Gradescope (AI answer grouping and batch feedback) and StarGrader (generative AI with custom prompts for personalised feedback) — against traditional manual grading. The study analysed 6,922 student-question responses from three core undergraduate courses in Multivariable Calculus, Linear Algebra, and Differential Equations. Results demonstrated exceptionally high inter-rater reliability. Gradescope achieved 99.91% exact agreement with manual grading and 99.95% of scores within ±5 marks, while StarGrader attained 95.57% exact agreement and 96.44% practical reliability. Time savings were modest (average for Gradescope 22.2% and StarGrader 19.8%), but reached 30% for another course in Mathematical Analysis, and up to 35.1% in computation-heavy worksheets with refined prompts. Prompt engineering turned out to be critical especially for StarGrader, with accuracy improving from 43.75% to 94.17% through systematic refinement. Both student surveys and instructor ratings indicated more positive perceptions of AI-generated feedback regarding clarity, personalisation, and usefulness. The findings indicate that AI tools, when supported by effective prompt strategies and human oversight, can substantially enhance feedback quality and efficiency in mathematics assessments. Two evidence-based user manuals were developed to support wider adoption of the tools.

Metrics

Metrics Loading ...

Downloads

Published

2026-07-25

How to Cite

Chiang, S. L. (2026). The Comparative Effectiveness of AI Tools in Enhancing Feedback Quality and Efficiency in Undergraduate Mathematics Assessments. Proceedings of The World Conference on Research in Teaching and Education, 5(1), 19–37. https://doi.org/10.33422/worldte.v5i1.1999