arXiv cs.LG
7/9/2026

D2PO: Optimizing Diffusion Samplers via Dynamic Preference
Short summary
D2PO reformulates diffusion sampler optimization as a preference-based alignment problem using the DPO framework, addressing the limitation that low-NFE student samplers trained via regression sacrifice high-frequency texture fidelity. The method models the sampling policy as an energy-based model, derives energy from the pretrained score network, and introduces dynamic preferences that self-improve as sampling policies are learned. Experiments show D2PO consistently outperforms conventional regression-based schedulers under low-NFE constraints.
- •D2PO applies Direct Preference Optimization to diffusion samplers, treating timestep schedules and CFG weights as optimizable policy parameters
- •Energy-based model formulation enables preference comparisons in perturbed spaces capturing both structural consistency and fine-grained detail
- •Dynamic self-improving preferences replace static teacher supervision, outperforming regression-based schedulers under low-NFE constraints
Generated with AI, which can make mistakes.
Is this a good recommendation for you?