Dictionary entry Trending
direct preference optimization
also DPO
Definition of direct preference optimization
-
: a method for aligning a language model with pairs of preferred and rejected responses by directly optimizing the model on those preferences, without first training a separate explicit reward model and then running reinforcement learning
- The team used direct preference optimization on pairs in which reviewers chose the clearer of two answers.
- DPO simplified the post-training pipeline, but its result still depended on the quality and coverage of the preference data.
Origin & history
Rafael Rafailov and collaborators introduced Direct Preference Optimization in a 2023 paper titled Your Language Model is Secretly a Reward Model. The method rewrites the preference-learning objective so that a policy can be trained with a classification-style loss relative to a reference model.
Test yourself
Which of these is the meaning of direct preference optimization?
Cite this entry
"Direct preference optimization." AI Dictionary, Dadgogo, https://dadgogo.com/dictionary/direct-preference-optimization/. Accessed 8 Oct. 2026.
Dictionary entries near direct preference optimization
- deepfake
- diffusion model
- digital human
- digital twin
- direct preference optimization
- discrete diffusion
- distillation
- distribution shift
- doomer