AI & Computingpreprint2026-08-09

Cross model Direct Preference Optimization

Open access0 citations

Abstract

Since direct preference optimization (DPO) was introduced in 2023, it has become a widely used preference optimization strategy for post-training large language models (LLMs). DPO's implicit reward formulation is coupled to its reference policy, and its theoretical guarantees assume the preference data distribution aligns with the model being trained. In practice, however, preference datasets are frequently reused across different models without this assumption being satisfied. Subsequent research has shown that distribution shifts between training and evaluation data degrade DPO's effectiveness, raising questions about how well preference data transfers across models; a common but understudied practice. In this paper we analyze more precisely the effectiveness of DPO using different open-weight large language models, where preference pairs are generated with one model and then used to optimize a second model. We evaluate cross-model DPO transfer across two model families, Meta's Llama and Alibaba Cloud's Qwen series, using prompts from the UltraFeedback dataset with Shakespearean style as the preference axis, measuring implicit-reward ranking accuracy on held-out preference pairs built from unseen test-split prompts. We find that the preference signal itself produces no detectable gap - (same-model ranking accuracy averages 0.900 vs 0.910 cross-model, with every cell in the 6×6 matrix significantly above the 0.5 chance level - while cross model evaluation produces an eight point distribution gap shift. However, a generation-level check shows that the DPO-trained models still produce modern English on plain prompts: at this training scale the transferred preference is expressed in the implicit reward, not in generation behaviour.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-09

Authors: Rafayel Latif

Institutions: Sterling Research Group