Fetching the paper…
Reading the bibliography…
Best-of-N (BoN) sampling, a common strategy for test-time scaling of Large Language Models (LLMs), relies on reward models to select the best candidate solution from multiple generations.
The use of scoring rubrics: Reliability, validity and educational consequences
Anders Jonsson and Gunilla Svingby. 2007 · 2007
Earlier work this paper cites.
A study of reliability of marking and absolute grading in secondary schools
K Abdul Gafoor and P Jisha. 2014 · 2014
Earlier work this paper cites.
Measuring Mathematical Problem Solving With the MATH Dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack William Rae, and Laurent Sifre. 2022 · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Benchmarking foundation models with language-model-as-an-examiner
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. 2023 · 2023
Earlier work this paper cites.
Llm-blender: Ensembling large language models with pairwise comparison and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023 · 2023
Earlier work this paper cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023 · 2023
Earlier work this paper cites.
Math-shepherd: A label-free step-by-step verifier for llms in mathematical reasoning
Peiyi Wang, Lei Li, Zhihong Shao, RX Xu, Damai Dai, Yifei Li, Deli Chen, Y Wu, and Zhifang Sui. 2023 · 2023
Earlier work this paper cites.
Slic-hf: Sequence likelihood calibration with human feedback
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J Liu. 2023 · 2023
Earlier work this paper cites.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Earlier work this paper cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Earlier work this paper cites.
The neurips 2021 consistency experiment
Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Wortman Vaughan. 2021 · 2024
Cited alongside, same era.
Tournament design: A review from an operational research perspective
Karel Devriesere, László Csató, and Dries Goossens. 2024 · 2024
Cited alongside, same era.
RLHF workflow: From reward modeling to online RLHF
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang. 2024 · 2024
Cited alongside, same era.
OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Cited alongside, same era.
Evaluating robustness of reward models for mathematical reasoning
Sunghwan Kim, Dongjin Kang, Taeyoon Kwon, Hyungjoo Chae, Jungsoo Won, Dongha Lee, and Jinyoung Yeo. 2024 · 2024
Cited alongside, same era.
Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems
Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng, Yiwen Hu, Yiru Tang, Jiapeng Wang, Xiaoxue Cheng, Huatong Song, et al. 2024 · 2024
Later among the works it cites.
Introducing openai o1 preview
OpenAI. 2024 · 2024
Later among the works it cites.
Alibaba Qwen Team. 2024 · 2024
Later among the works it cites.
QwQ-32B Preview
Qwen Team, Alibaba. 2023 · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rewardbench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. 2024 · 2024
Cited alongside, same era.
Aligning large language models by on-policy self-judgment
Sangkyu Lee, Sungdong Kim, Ashkan Yousefpour, Minjoon Seo, Kang Min Yoo, and Youngjae Yu. 2024 · 2024
Cited alongside, same era.
From generation to judgment: Opportunities and challenges of llm-as-a-judge
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2024 · 2024
Cited alongside, same era.
Numinamath
Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu. 2024 · 2024
Cited alongside, same era.
Skywork reward model series
Chris Yuhao Liu and Liang Zeng. 2024 · 2024
Cited alongside, same era.
Rm-bench: Benchmarking reward models of language models with subtlety and style
Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, and Juanzi Li. 2024 · 2024
Cited alongside, same era.
Llm critics help catch llm bugs
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. 2024 · 2024
Cited alongside, same era.
Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang. 2024 · 2024
Later among the works it cites.
Online iterative reinforcement learning from human feedback with general preference model
Chenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong, Nan Jiang, and Tong Zhang. 2024 · 2024
Later among the works it cites.
Introducing llama 3.1: Our most capable models to date
Meta AI. 2025 · 2025
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2025
Closest in time.
Deepspeed: Extreme-scale model training for everyone
Jeffrey Rasley, Zhizhou He, Jianwen Song, and Mikhail Smelyanskiy. 2020 · 2025
Closest in time.
Hugging face’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clément Delangue, and Alexander Maisonneuve. 2020 · 2025
Closest in time.