Abstract
Student behaviour detection is essential in intelligent education for real-time monitoring of classroom dynamics, personalised teaching, and assessment of learning engagement. Conventional approaches—video-based action recognition, pose estimation, and standard object detection—face major limitations: they require massive labelled datasets (e.g., 1.58 million frames in AVA), struggle with context dependence, fail in crowded/occluded scenes, and cannot reliably handle the extreme multi-scale variation (up to 25-fold pixel differences between front and back rows), severe student overlaps, and behavioural similarities typical of real classrooms. Here we show that feature distillation in a teacher– student framework substantially improves YOLOv7 for this task. Knowledge is transferred from a high-capacity YOLOv7-W6 teacher to a lightweight YOLOv7-tiny student, enabling model compression while boosting accuracy. On the SCB-Dataset5 (7,428 images, 106,830 annotations across 20 distinct student behaviours, partitioned 4:1 into training and validation sets), the optimised model increases Map(0.5) by 5.52 %, reduces size by 15 MB, runs at 54.5 FPS, and delivers a 23 % gain in small-object AP versus YOLOv5s, with marked improvements on complex behaviours such as reading. Feature-activation maps confirm clearer target localisation. This lightweight, real-time solution enables practical deployment in resource-constrained smart classrooms and opens pathways to multimodal educational AI.
Showing the abstract — retrieve the full paper via the Exa API.