Oct 10 – 11, 2026
William & Mary, Integrated Science Center 4
America/New_York timezone
Registration for PICUP workshop and Zoom attendance still open!

When AI Becomes the Grader: Comparing LLMs Versus Human Graders on Physics Essays

Oct 10, 2026, 12:00 PM
15m
William & Mary, Integrated Science Center 4

William & Mary, Integrated Science Center 4

540 Landrum Dr Williamsburg, VA 23185
Talk (15 min) auditorium

Speaker

Jason Tran (Georgetown University)

Description

Grading open-ended physics work is tedious, time-consuming, and impractical for large courses. With the emergence of large language models (LLMs) and their ongoing improvements in problem-solving capabilities, a question arises: Can AI meaningfully evaluate students’ conceptual explanations in complex domains such as quantum mechanics? This research study examines AI-assisted grading by comparing trained human grading with AI-generated grading on conceptual essay responses from an online quantum mechanics course. Using a shared rubric, we analyze the level of agreement among graders and with Google Gemini, and investigate how each group interprets and applies evaluative criteria. The results highlight where AI evaluation aligns with human judgment, where it diverges, and what these differences reveal about grading validity, rubric interpretation, and student reasoning in conceptual quantum mechanics. Ultimately, the study raises a central question for physics education: should AI be trusted to grade students’ thinking, and if so, under what conditions?

Primary author

Jason Tran (Georgetown University)

Co-authors

James Freericks (Georgetown University) Leanne Doughty (Georgetown University)

Presentation materials

There are no materials yet.