DadgogoAI Dictionary

Dictionary entry Trending

direct preference optimization

noun \də-ˈrekt ˈpre-f(ə-)rən(t)s ˌäp-tə-mə-ˈzā-shən\

also DPO

Advertisement

Definition of direct preference optimization

  1. : a method for aligning a language model with pairs of preferred and rejected responses by directly optimizing the model on those preferences, without first training a separate explicit reward model and then running reinforcement learning

    • The team used direct preference optimization on pairs in which reviewers chose the clearer of two answers.
    • DPO simplified the post-training pipeline, but its result still depended on the quality and coverage of the preference data.

Origin & history

first known use2023

Rafael Rafailov and collaborators introduced Direct Preference Optimization in a 2023 paper titled Your Language Model is Secretly a Reward Model. The method rewrites the preference-learning objective so that a policy can be trained with a classification-style loss relative to a reference model.

Test yourself

Which of these is the meaning of direct preference optimization?

Cite this entry

"Direct preference optimization." AI Dictionary, Dadgogo, https://dadgogo.com/dictionary/direct-preference-optimization/. Accessed 8 Oct. 2026.

Dictionary entries near direct preference optimization

Advertisement