DeepLearningAI
6/3/2026

Optimize, deploy, and benchmark an open-source LLM with vLLM
Short summary
DeepLearningAI and Red Hat's course teaches efficient open-source LLM serving using vLLM. You'll learn to shrink model weights through quantization, leverage vLLM's memory management techniques like PagedAttention and prefix caching to handle high-throughput, low-latency inference, and benchmark deployments under realistic traffic. The course covers the full optimize-deploy-benchmark workflow using real tools like LLM Compressor, GuideLLM, and lm-eval.
- •Covers vLLM-based optimization and deployment for open-source LLMs
- •Teaches memory management techniques (quantization, PagedAttention, prefix caching)
- •Includes hands-on benchmarking with real tools and models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



