Back to feed
Dev.to
Dev.to
8/1/2026
The original title is: "7.5B Gemma-4-e4b outscored 24B Devstral on a 56-task local coding benchmark"

The original title is: "7.5B Gemma-4-e4b outscored 24B Devstral on a 56-task local coding benchmark"

Original: A 7.5B model beat a 24B on my coding benchmark.

Short summary

A developer built a 56-task coding benchmark with hidden tests and ran 16 local model configurations on a single 16 GB RTX 5060 Ti. The 7.5B gemma-4-e4b (5 GiB) scored 42/56, beating the 24B devstral-small-2-24b (40/56) and gpt-oss-20b (38/56) with non-overlapping ranges across three runs each. The ranking inverts public leaderboards, and deterministic test-based grading avoids LLM-judge inflation that plagues other benchmarks.

  • 7.5B gemma-4-e4b outscored 24B devstral and 20B gpt-oss on a 56-task hidden-test coding benchmark on a 16 GB card
  • Rankings invert public leaderboards; deterministic compiler/test-runner grading avoids LLM-judge inflation
  • Run-to-run variance is model-specific: one model swings 8 points between identical consecutive runs

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more