Author
Hao Xu
Recent research
- AI & ComputingOpen access
Enhancing large language model reasoning with reward models: an analytical survey
Abstract Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune LLMs during reinforcement learning (RL) and help select the best answer from multiple candidates during inference. In this...