Exa is a modern AI search engine with SERP API, website crawler tools, and deep research API. Power your app with web search AI and web crawling API.
Publication

Multimodal Speaker Identification in Classroom Environments

Jun 10, 2026 · 6 authors · 3 topics

Automated analysis of K-12 classroom dynamics faces challenges due to background noise and variable child speech, often confounding acoustic-only models. This study evaluates a multimodal speaker identification framework anchoring acoustic embeddings with LLM-derived semantic context. Using a subset of the EDSI dataset (8 math classrooms, N = 2,801 utterances), we found an acoustic baseline (ECAPA-TDNN) achieved only 39.0% accuracy. By integrating transcript-based "contextual anchoring" into a gradient boosting classifier, our multimodal approach raised student identification to 50.3%. Performance also improved for utterances over 5 seconds, reaching 76.9% accuracy (vs. 64.9% baseline) with a 90.9% Top-3 accuracy. Additionally, the model distinguished teacher vs. student roles with 99.3% accuracy. This approach advances the feasibility of automated feedback systems capable of considering individual student participation, a crucial step for supporting equitable instruction at scale.

Showing the abstract — retrieve the full paper via the Exa API.

Michael L. ChrzanMeghavarshini KrishnaswamyRobert GibboniKatie WetstoneWei AiJing Liu
Speech Recognition and SynthesisEmotion and Mood RecognitionIntelligent Tutoring Systems and Adaptive Learning
PublishedJun 10, 2026
TypeArticle
Citations0

Powered by the Exa API

Multimodal Speaker Identification in Classroom Environments | Exa