Dev.to
8/1/2026

The original title is: "7.5B Gemma-4-e4b outscored 24B Devstral on a 56-task local coding benchmark"
Original: A 7.5B model beat a 24B on my coding benchmark.
Short summary
A developer built a 56-task coding benchmark with hidden tests and ran 16 local model configurations on a single 16 GB RTX 5060 Ti. The 7.5B gemma-4-e4b (5 GiB) scored 42/56, beating the 24B devstral-small-2-24b (40/56) and gpt-oss-20b (38/56) with non-overlapping ranges across three runs each. The ranking inverts public leaderboards, and deterministic test-based grading avoids LLM-judge inflation that plagues other benchmarks.
- •7.5B gemma-4-e4b outscored 24B devstral and 20B gpt-oss on a 56-task hidden-test coding benchmark on a 16 GB card
- •Rankings invert public leaderboards; deterministic compiler/test-runner grading avoids LLM-judge inflation
- •Run-to-run variance is model-specific: one model swings 8 points between identical consecutive runs
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



