MarkTechPost
7/9/2026

Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput
Short summary
NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant reducing Nemotron-3-Super from 120.7B to 75.3B parameters while delivering 2.03x throughput (100 tok/s per user on 8xB200). Iterative Puzzle technique alternates compression with knowledge distillation. H100 inference handles 8x more concurrent 1M-token requests (1→8), reducing inference costs.
- •Parameter reduction from 120.7B to 75.3B with 2.03x throughput gain
- •Iterative Puzzle alternates structural compression with knowledge distillation recovery
- •H100 concurrent requests scale from 1 to 8 at 1M-token context
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


