Institution
Sterling Research Group
Recent research
- AI & ComputingOpen access
Cross model transfer of Direct Preference Optimization
Since direct preference optimization (DPO) was introduced in 2023, it has become a widely used preference optimization strategy for post-training large language models (LLMs). DPO's implicit reward formulation is coupled to its reference policy, and its theoretical guarantees ass...
- AI & ComputingOpen access
Cross model transfer of Direct Preference Optimization
Since direct preference optimization (DPO) was introduced in 2023, it has become a widely used preference optimization strategy for post-training large language models (LLMs). DPO's implicit reward formulation is coupled to its reference policy, and its theoretical guarantees ass...
- AI & ComputingOpen access
Cross model Direct Preference Optimization
Since direct preference optimization (DPO) was introduced in 2023, it has become a widely used preference optimization strategy for post-training large language models (LLMs). DPO's implicit reward formulation is coupled to its reference policy, and its theoretical guarantees ass...
- Climate & EnvironmentOpen access
Forecast laws near nonlinear transitions may be poorly concentrated even when they are not cleanly bimodal. For a conditional law P(t,h) on a metric state space, we study Uδ(t,h) = 1 − supₐ P(t,h){d(Y,a) ≤ δ}: the minimum probability that a point forecast misses the realization b...