Back to feed
MarkTechPost
MarkTechPost
7/9/2026
Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput

Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput

Short summary

NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant reducing Nemotron-3-Super from 120.7B to 75.3B parameters while delivering 2.03x throughput (100 tok/s per user on 8xB200). Iterative Puzzle technique alternates compression with knowledge distillation. H100 inference handles 8x more concurrent 1M-token requests (1→8), reducing inference costs.

  • Parameter reduction from 120.7B to 75.3B with 2.03x throughput gain
  • Iterative Puzzle alternates structural compression with knowledge distillation recovery
  • H100 concurrent requests scale from 1 to 8 at 1M-token context

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more