Back to feed
AR
arXiv CS.AI
7/8/2026
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming

CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming

Short summary

CSTutorBench is a new benchmark for evaluating small language models (4B–120B parameters) as CS tutors in VEX VR, a block-based robotics environment. Across 11 models tested with 17 scenario-based questions and a human-in-the-loop LLM-as-judge pipeline, models excelled at surface-level criteria like tone and vocabulary but struggled with deeper pedagogical behaviors such as avoiding answer leakage and engaging with student debugging histories. Model family and instruction-tuning approach were better predictors of tutoring quality than parameter count, and targeted prompt revisions grounded in educational research improved scores for 10 of 11 models.

  • CSTutorBench evaluates SLMs as CS tutors for block-based programming using 17 scenario questions and a pedagogical rubric
  • Models perform well on surface-level criteria but fail at deeper pedagogical behaviors like avoiding answer leakage
  • Model family and instruction-tuning matter more than parameter count; prompt engineering improved 10 of 11 models

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more