
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
Short summary
CSTutorBench is a new benchmark for evaluating small language models (4B–120B parameters) as CS tutors in VEX VR, a block-based robotics environment. Across 11 models tested with 17 scenario-based questions and a human-in-the-loop LLM-as-judge pipeline, models excelled at surface-level criteria like tone and vocabulary but struggled with deeper pedagogical behaviors such as avoiding answer leakage and engaging with student debugging histories. Model family and instruction-tuning approach were better predictors of tutoring quality than parameter count, and targeted prompt revisions grounded in educational research improved scores for 10 of 11 models.
- •CSTutorBench evaluates SLMs as CS tutors for block-based programming using 17 scenario questions and a pedagogical rubric
- •Models perform well on surface-level criteria but fail at deeper pedagogical behaviors like avoiding answer leakage
- •Model family and instruction-tuning matter more than parameter count; prompt engineering improved 10 of 11 models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?