Back to feed
r/MachineLearning
r/MachineLearning
7/19/2026
ASCIITermDraw-Bench | Evaluating VLMs on ASCII Generation and Editing Tasks [P]

ASCIITermDraw-Bench | Evaluating VLMs on ASCII Generation and Editing Tasks [P]

Short summary

ASCIITermDraw-Bench is a new benchmark evaluating Vision Language Models on their ability to generate and edit ASCII-based diagrams across 80 tasks in four categories: basic layouts, network topologies, software architecture, and image-conditioned editing. Each response is scored structurally and semantically via an LLM judge with 95% confidence intervals. The current leaderboard is led by Gemma-4-31B-IT at 73.8%, with the benchmark and methodology publicly available on Hugging Face.

  • New benchmark with 80 tasks across 4 categories evaluates VLMs on ASCII diagram generation and editing
  • Dual scoring system uses structural verification plus LLM-judge semantic scoring with 95% confidence intervals
  • Gemma-4-31B-IT leads at 73.8%; full methodology and 12 example tasks available on Hugging Face

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more