Institution

Sterling Research Group

USfacility

Recent research

  • AI & ComputingOpen access

    Cross model transfer of Direct Preference Optimization

    Since direct preference optimization (DPO) was introduced in 2023, it has become a widely used preference optimization strategy for post-training large language models (LLMs). DPO's implicit reward formulation is coupled to its reference policy, and its theoretical guarantees ass...

    Zenodo (CERN European Organization for Nuclear Research)2026-08-090 citationsDOI
  • AI & ComputingOpen access

    Cross model transfer of Direct Preference Optimization

    Since direct preference optimization (DPO) was introduced in 2023, it has become a widely used preference optimization strategy for post-training large language models (LLMs). DPO's implicit reward formulation is coupled to its reference policy, and its theoretical guarantees ass...

    Zenodo (CERN European Organization for Nuclear Research)2026-08-090 citationsDOI
  • AI & ComputingOpen access

    Cross model Direct Preference Optimization

    Since direct preference optimization (DPO) was introduced in 2023, it has become a widely used preference optimization strategy for post-training large language models (LLMs). DPO's implicit reward formulation is coupled to its reference policy, and its theoretical guarantees ass...

    Zenodo (CERN European Organization for Nuclear Research)2026-08-090 citationsDOI
  • Climate & EnvironmentOpen access

    Graded Branch Unpredictability Near Nonlinear Transitions: A Decision-Scale Framework and Held-Out Energy-Balance Demonstration

    Forecast laws near nonlinear transitions may be poorly concentrated even when they are not cleanly bimodal. For a conditional law P(t,h) on a metric state space, we study Uδ(t,h) = 1 − supₐ P(t,h){d(Y,a) ≤ δ}: the minimum probability that a point forecast misses the realization b...

    Zenodo (CERN European Organization for Nuclear Research)2026-08-070 citationsDOI