MarkTechPost
7/9/2026

NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput
Short summary
NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super using Iterative Puzzle, which alternates hardware-aware structural compression with knowledge distillation recovery. The model shrinks from 120.7B total / 12.8B active parameters to 75.3B / 9.3B while delivering 2.03x server throughput at matched per-user throughput on an 8xB200 node. On a single H100, 1M-token concurrency improves from 1 concurrent request to 8.
- •Nemotron-Labs-3-Puzzle-75B-A9B compresses Nemotron-3-Super from 120.7B/12.8B to 75.3B/9.3B parameters using Iterative Puzzle compression with distillation recovery
- •Delivers 2.03x server throughput at 100 tok/s per user on an 8xB200 node
- •H100 concurrency for 1M-token requests increases from 1 to 8 simultaneous requests
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


