Dev.to
7/15/2026

Model Showdown Round 9: Qwen 3.6 27B vs Qwen 3.6 35B-A3B vs Qwythos-9B vs GLM-4.7-Flash vs Nemotron-3-Nano
Short summary
Round 9 of a local model showdown tests five open-weight models (Qwen 3.6 27B, Qwen 3.6 35B-A3B, Qwythos-9B, GLM-4.7-Flash, Nemotron-3-Nano) on a real coding task using an RTX 5090 and llama.cpp. Two models failed before the task even started due to llama.cpp template bugs requiring live patching. The series provides rigorous, reproducible benchmarks of dense vs MoE architectures for agentic coding workflows.
- •Five local models benchmarked on identical coding task with standardized agent harness
- •Two models (Qwythos-9B, Nemotron-3-Nano) failed on first request due to llama.cpp template bugs
- •Tests dense vs MoE vs hybrid Transformer-Mamba architectures on real-world coding agent tasks
- •Run on RTX 5090 with llama.cpp, no API keys or cloud spend required
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



