Abstract
Scaling high-quality mathematics assessment requires balancing automated generation with pedagogical standards. While Large Language Models show promise for educational content creation, systematic evaluation of their output and integration with human expertise remain underexplored. We propose a Generate–Filter–Edit human–AI pipeline combining GPT-4's generative capacity with expert teacher judgment to produce middle-school mathematics items for large-scale deployment. Analyzing 5,634 GPT-4-generated questions, we found 68.74% ready for immediate use, 19.42% requiring revision, and 11.84% discarded. Through comprehensive statistical significance testing on 378 textual features, we identified key linguistic predictors distinguishing quality levels. Low-quality items exhibit significantly higher paragraph-level textual overlap (Cohen's d up to 0.403) and greater verbosity, while items requiring revision show higher rare vocabulary ratios. Teacher editing behaviors follow a light-touch approach, focusing on mathematical language formalization while preserving content fidelity. Post-deployment engagement analysis revealed significant correlations between textual complexity features and student–chatbot interaction patterns. Our results validate scalable human–AI collaboration for mathematics item development and provide evidence-based frameworks for automated quality assessment and editing support systems.
Showing the abstract — retrieve the full paper via the Exa API.