Back to feed
MarkTechPost
MarkTechPost
7/9/2026
NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput

NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput

Short summary

NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super using Iterative Puzzle, which alternates hardware-aware structural compression with knowledge distillation recovery. The model shrinks from 120.7B total / 12.8B active parameters to 75.3B / 9.3B while delivering 2.03x server throughput at matched per-user throughput on an 8xB200 node. On a single H100, 1M-token concurrency improves from 1 concurrent request to 8.

  • Nemotron-Labs-3-Puzzle-75B-A9B compresses Nemotron-3-Super from 120.7B/12.8B to 75.3B/9.3B parameters using Iterative Puzzle compression with distillation recovery
  • Delivers 2.03x server throughput at 100 tok/s per user on an 8xB200 node
  • H100 concurrency for 1M-token requests increases from 1 to 8 simultaneous requests

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more