Back to feed
arXiv cs.LG
arXiv cs.LG
7/9/2026
D2PO: Optimizing Diffusion Samplers via Dynamic Preference

D2PO: Optimizing Diffusion Samplers via Dynamic Preference

Short summary

D2PO reformulates diffusion sampler optimization as a preference-based alignment problem using the DPO framework, addressing the limitation that low-NFE student samplers trained via regression sacrifice high-frequency texture fidelity. The method models the sampling policy as an energy-based model, derives energy from the pretrained score network, and introduces dynamic preferences that self-improve as sampling policies are learned. Experiments show D2PO consistently outperforms conventional regression-based schedulers under low-NFE constraints.

  • D2PO applies Direct Preference Optimization to diffusion samplers, treating timestep schedules and CFG weights as optimizable policy parameters
  • Energy-based model formulation enables preference comparisons in perturbed spaces capturing both structural consistency and fine-grained detail
  • Dynamic self-improving preferences replace static teacher supervision, outperforming regression-based schedulers under low-NFE constraints

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more