r/MachineLearning
7/19/2026
![ASCIITermDraw-Bench | Evaluating VLMs on ASCII Generation and Editing Tasks [P]](https://preview.redd.it/p8w6ju0sk5eh1.png?width=140&height=84&auto=webp&s=dc9838126b956835f1fe68d2840d3e4adbb10dcc)
ASCIITermDraw-Bench | Evaluating VLMs on ASCII Generation and Editing Tasks [P]
Short summary
ASCIITermDraw-Bench is a new benchmark evaluating Vision Language Models on their ability to generate and edit ASCII-based diagrams across 80 tasks in four categories: basic layouts, network topologies, software architecture, and image-conditioned editing. Each response is scored structurally and semantically via an LLM judge with 95% confidence intervals. The current leaderboard is led by Gemma-4-31B-IT at 73.8%, with the benchmark and methodology publicly available on Hugging Face.
- •New benchmark with 80 tasks across 4 categories evaluates VLMs on ASCII diagram generation and editing
- •Dual scoring system uses structural verification plus LLM-judge semantic scoring with 95% confidence intervals
- •Gemma-4-31B-IT leads at 73.8%; full methodology and 12 example tasks available on Hugging Face
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


