Authors
Michael L. ChrzanMeghavarshini KrishnaswamyRobert GibboniKatie WetstoneWei AiJing Liu
Topics
Speech Recognition and SynthesisEmotion and Mood RecognitionIntelligent Tutoring Systems and Adaptive LearningMichael Leon Chrzan 1, Meghavarshini Krishnaswamy 1, Robert Gibboni 2, Katie Wetstone 2, Wei Ai 3, and Jing Liu 1 1Center for Educational Data Science and Innovation, University of Maryland 2DrivenData 3College of Information and the Institute for Advanced Computer Studies, University of Maryland Corresponding Authors: mlchrzan@umd.edu, jliu28@umd.eduAbstract—Automated analysis of K-12 classroom dynamics faces challenges due to background noise and variable child speech, often confounding acoustic-only models. This study evaluates a multimodal speaker identification framework an choring acoustic embeddings with LLM-derived semantic con text. Using a subset of the EDSI dataset (8 math classrooms, N = 2, 801 utterances), we found an acoustic baseline (ECAPA TDNN) achieved only 39.0% accuracy. By integrating transcript based ”contextual anchoring” into a gradient boosting classifier, our multimodal approach raised student identification to 50.3%. Performance also improved for utterances over 5 seconds, reach ing 76.9% accuracy (vs. 64.9% baseline) with a 90.9% Top 3 accuracy. Additionally, the model distinguished teacher vs. student roles with 99.3% accuracy. This approach advances the feasibility of automated feedback systems capable of considering individual student participation, a crucial step for supporting equitable instruction at scale. Index Terms—speaker identification, multimodal, K-12 class rooms
About
PublishedJun 10, 2026
TypePreprint
Citations0