Model Releases
EduAgentQG: Multi-Agent Personalized Mathematics Question Generation with Explicit Diversity and Objective-Aware Evaluation
arXiv:2511.11635v2 Announce Type: replace-cross Abstract: In intelligent education, personalized mathematics question generation aims to produce mathematics questions that satisfy educational requirem
arXiv:2511.11635v2 Announce Type: replace-cross Abstract: In intelligent education, personalized mathematics question generation aims to produce mathematics questions that satisfy educational requirements while supporting adaptive assessment and learning. Existing LLM-based single-agent and multi-agent methods improve generation flexibility, but they still tend to rely on aggregated feedback or model randomness, making it difficult to jointly ensure dimension-wise objective alignment and controllable diversity. To address these challenges, we propose EduAgentQG, a multi-agent collaborative framework for personalized mathematics question generation with explicit diversity and objective-aware evaluation. EduAgentQG organizes question generation as a closed-loop process of planning, writing, evaluation, refinement, and checking: structured generation plans and multiple generation directions guide candidate generation, while fine-grained evaluation verifies logical correctness, solvability, and objective alignment in knowledge concepts, difficulty, grade level, and core competencies. We first construct a mathematics question generation benchmark containing 10,273 questions across Grades 1-9, covering 634 knowledge concepts, 16 core competencies, and three difficulty levels; for evaluation, it is organized into two subsets: MathChoice, with 489 educational objectives for multiple-choice question generation, and MathBlank, with 500 educational objectives for fill-in-the-blank question generation. Experiments show that EduAgentQG consistently outperforms COT, COT_N, ReAct, and EQPR in diversity, Objective Consistency, and Win Rate.
Related
- DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation
- AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing
- Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
- PeReGrINE: Evaluating Personalized Review Fidelity with User Item Graph Context
Source: arXiv cs.CL | 2026-08-27