Fetching the paper…
Reading the bibliography…
Inference-time computation offers a powerful axis for scaling the performance of language models.
Various techniques used in connection with random digits
John Von Neumann · 1963
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Nemirovski, David Borisovich Yudin, and Edgar Ronald Dawson · 1983
Earlier work this paper cites.
Information-based complexity
Joseph F Traub, Grzegorz W Wasilkowski, and Henryk Woźniakowski · 1988
Earlier work this paper cites.
Weakly learning DNF and characterizing statistical query learning using Fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Channel coding: Non-asymptotic fundamental limits
Yury Polyanskiy · 2010
Earlier work this paper cites.
Information-based complexity, feedback and dynamics in convex programming
Maxim Raginsky and Alexander Rakhlin · 2011
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization
Alekh Agarwal, Peter L Bartlett, Pradeep Ravikumar, and Martin J Wainwright · 2012
Earlier work this paper cites.
A complete characterization of statistical query learning with applications to evolvability
Vitaly Feldman · 2012
Earlier work this paper cites.
Online local learning via semidefinite programming
P. Christiano · 2014
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
A general characterization of the statistical query complexity
Vitaly Feldman · 2017
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Dueling posterior sampling for preference-based reinforcement learning
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue, and Joel Burdick · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Preference-based reinforcement learning with finite-time guarantees
Yichong Xu, Ruosong Wang, Lin Yang, Aarti Singh, and Artur Dubrawski · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Is pessimism provably efficient for offline RL?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Dueling RL: Reinforcement learning with trajectory preferences
Aldo Pacchiano, Aadirupa Saha, and Jonathan Lee · 2021
Earlier work this paper cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Earlier work this paper cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, and Liwei Wang · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
The sample complexity of approximate rejection sampling with applications to smoothed online learning
Adam Block and Yury Polyanskiy · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Earlier work this paper cites.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alex D’Amour, DJ Dvijotham, Adam Fisch, Katherine Heller, Stephen Pfohl, Deepak Ramachandran, et al · 2023
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
Xidong Feng, Ziyu Wan, Muning Wen, Ying Wen, Weinan Zhang, and Jun Wang · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
REBEL: Reinforcement learning via regressing relative rewards
Zhaolin Gao, Jonathan D Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, J Andrew Bagnell, Jason D Lee, and Wen Sun · 2024
Later among the works it cites.
BoNBoN alignment for large language models and the sweetness of best-of-n sampling
Lin Gui, Cristina Gârbacea, and Victor Veitch · 2024
Later among the works it cites.
Self-play with adversarial critic: Provable and scalable offline alignment for language models
Xiang Ji, Sanjeev Kulkarni, Mengdi Wang, and Tengyang Xie · 2024
Later among the works it cites.
Regularized best-of-n sampling to mitigate reward hacking for language model alignment
Yuu Jinnai, Tetsuro Morimura, Kaito Ariu, and Kenshi Abe · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Cited alongside, same era.
Is RLHF more difficult than standard RL?
Yuanhao Wang, Qinghua Liu, and Chi Jin · 2023
Cited alongside, same era.
Making RL with preference-based feedback efficient via randomization
Runzhe Wu and Wen Sun · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Cited alongside, same era.
Saeed Khaki, JinJin Li, Lan Ma, Liu Yang, and Prathap Ramachandra · 2024
Later among the works it cites.
Args: Alignment as reward-guided search
Maxim Khanov, Jirayu Burapacheep, and Yixuan Li · 2024
Later among the works it cites.
Openassistant conversations-democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al · 2024
Later among the works it cites.
Information theoretic guarantees for policy alignment in large language models
Youssef Mroueh · 2024
Later among the works it cites.
Controlled decoding from language models
Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanping Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, et al · 2024
Later among the works it cites.
West-of-n: Synthetic preference generation for improved reward modeling
Alizée Pace, Jonathan Mallinson, Eric Malmi, Sebastian Krause, and Aliaksei Severyn · 2024
Later among the works it cites.
Treebon: Enhancing inference-time alignment with speculative tree-search and best-of-n sampling
Jiahao Qiu, Yifu Lu, Yifan Zeng, Jiacheng Guo, Jiayi Geng, Huazheng Wang, Kaixuan Huang, Yue Wu, and Mengdi Wang · 2024
Later among the works it cites.
Sail into the headwind: Alignment via robust rewards and dynamic labels against reward hacking
Paria Rashidinejad and Yuandong Tian · 2024
Later among the works it cites.
Bond: Aligning LLMs with Best-of-N distillation
Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Ramé, Bobak Shariari, Sarah Perrin, Abe Friesen, Geoffrey Cideron, et al · 2024
Later among the works it cites.
Decoding-time language model alignment with multiple objectives
Ruizhe Shi, Yifang Chen, Yushi Hu, ALisa Liu, Noah Smith, Hannaneh Hajishirzi, and Simon Du · 2024
Later among the works it cites.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Understanding preference fine-tuning through the lens of coverage
Yuda Song, Gokul Swamy, Aarti Singh, J Andrew Bagnell, and Wen Sun · 2024
Later among the works it cites.
Inference scaling fLaws: The limits of llm resampling with imperfect verifiers
Benedikt Stroebl, Sayash Kapoor, and Arvind Narayanan · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
From decoding to meta-generation: Inference-time algorithms for large language models
Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui · 2024
Later among the works it cites.
Exploratory preference optimization: Harnessing implicit Q*-approximation for sample-efficient rlhf
Tengyang Xie, Dylan J Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah, and Alexander Rakhlin · 2024
Later among the works it cites.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang · 2024
Later among the works it cites.
Genarm: Reward guided generation with autoregressive reward model for test-time alignment
Yuancheng Xu, Udari Madhushani Sehwag, Alec Koppel, Sicheng Zhu, Bang An, Furong Huang, and Sumitra Ganesh · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
A theoretical analysis of Nash learning from human feedback under general KL-regularized preference
Chenlu Ye, Wei Xiong, Yuheng Zhang, Nan Jiang, and Tong Zhang · 2024
Later among the works it cites.
Rest-mcts*: Llm self-training via process reward guided tree search
Dan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue, Yuxiao Dong, and Jie Tang · 2024
Later among the works it cites.
Probabilistic inference in language models via twisted sequential monte carlo
Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Grosse · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · 2025
Closest in time.
Gpt-4o-mini: A scaled-down variant of gpt-4, 2024a
OpenAI · 2025
Closest in time.