Publication

Generate-Filter-Edit: A Human-AI Collaborative Pipeline for Developing and Automatically Evaluating Middle School Mathematics Questions

Hai Li, Wanli Xing, Chenglu Li, Ran Gao

Jun 22, 2026
Abstract

Scaling high-quality mathematics assessment requires balancing automated generation with pedagogical standards. While Large Language Models show promise for educational content creation, systematic evaluation of their output and integration with human expertise remain underexplored. We propose a Generate–Filter–Edit human–AI pipeline combining GPT-4's generative capacity with expert teacher judgment to produce middle-school mathematics items for large-scale deployment. Analyzing 5,634 GPT-4-generated questions, we found 68.74% ready for immediate use, 19.42% requiring revision, and 11.84% discarded. Through comprehensive statistical significance testing on 378 textual features, we identified key linguistic predictors distinguishing quality levels. Low-quality items exhibit significantly higher paragraph-level textual overlap (Cohen's d up to 0.403) and greater verbosity, while items requiring revision show higher rare vocabulary ratios. Teacher editing behaviors follow a light-touch approach, focusing on mathematical language formalization while preserving content fidelity. Post-deployment engagement analysis revealed significant correlations between textual complexity features and student–chatbot interaction patterns. Our results validate scalable human–AI collaboration for mathematics item development and provide evidence-based frameworks for automated quality assessment and editing support systems.

1Authors

AuthorAffiliationh-indexCitations
Hai LiUniversity of Florida8192
Wanli XingUniversity of Miami374,679
Chenglu LiUniversity of Utah252,084
Ran GaoUniversity of Florida11

Showing the abstract — retrieve the full paper via the Exa API.

Powered by the Exa API