Speaker
Description
Grading open-ended physics work is tedious, time-consuming, and impractical for large courses. With the emergence of large language models (LLMs) and their ongoing improvements in problem-solving capabilities, a question arises: Can AI meaningfully evaluate students’ conceptual explanations in complex domains such as quantum mechanics? This research study examines AI-assisted grading by comparing trained human grading with AI-generated grading on conceptual essay responses from an online quantum mechanics course. Using a shared rubric, we analyze the level of agreement among graders and with Google Gemini, and investigate how each group interprets and applies evaluative criteria. The results highlight where AI evaluation aligns with human judgment, where it diverges, and what these differences reveal about grading validity, rubric interpretation, and student reasoning in conceptual quantum mechanics. Ultimately, the study raises a central question for physics education: should AI be trusted to grade students’ thinking, and if so, under what conditions?