University of Wisconsin–Madison

WCER Community Learning Series — Evaluating Language Model Assessment of Students’ Writing

Shamya Karumbaiah

A headshot of Shamya Karumbaiah with text that says WCER Community Learning Series with Shamya Karumbaiah, Assistant Professor, Department of Educational Psychology

Educational Sciences, Room 259 or virtual

A headshot of Shamya Karumbaiah with text that says WCER Community Learning Series with Shamya Karumbaiah, Assistant Professor, Department of Educational Psychology

Join Shamya Karumbaiah for this WCER Community Learning Series event.

Participants will learn how to systematically evaluate errors and biases when using language models to analyze student writing and qualitative data more broadly. The session combines interactive hands-on exploration with discussion of empirical evidence. We will discuss findings of two connected studies from the TRAIL Lab.

The first study, Prompt vs. Supervise, compares prompting general-purpose language models against training supervised classifiers to assess students’ explanations of physics concepts from an inquiry-based learning unit, weighing the trade-offs between flexibility and accuracy. The second, Beyond Generalization, shows how high accuracy isn’t enough to trust a model in real-world classrooms and introduces Behavior Analysis to stress-test models against realistic variations in student writing (e.g., spelling errors, acronyms, multilingual), revealing failures that standard evaluation misses.

Participants will conduct behavior analysis on student writing or qualitative data from their context using the AI Behavior Analysis Tool (AIBAT), a tool developed by TRAIL Lab for stakeholder-driven contextual evaluation of language model failures. 


Shamya Karumbaiah is an assistant professor in the UW–Madison Department of Educational Psychology and director of The Responsible AI for Learning (TRAIL) Lab. She studies human-centered AI for teaching and learning with the aim of augmenting human intelligence.

Karumbaiah’s current research focuses on constructing a scientific and critical understanding of equitable and responsible use of AI in classrooms. After more than ten years as a computer scientist, she earned a Ph.D. in learning sciences from the University of Pennsylvania. Her dissertation empirically investigated sources of bias in AI-based learning systems. Before joining UW–Madison, she spent a year as a postdoc fellow at Carnegie Mellon University, where she studied AI-augmented teacher practices in human-AI partnered instruction.