Fetching the paper…
Reading the bibliography…
Current reinforcement learning from human feedback (RLHF) pipelines for large language model (LLM) alignment typically assign scalar rewards to sequences, using the final token as a surrogate indicator for the quality of the entire sequence.
Attention is not explanation, 2019
Sarthak Jain and Byron C. Wallace · 1902
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Noisy-input entropy search for efficient robust bayesian optimization, 2020
Lukas P. Fröhlich, Edgar D. Klenske, Julia Vinogradska, Christian Daniel, and Melanie N. Zeilinger · 2002
Earlier work this paper cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo, 2020
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2005
Earlier work this paper cites.
On the relationship between shapley and owen values
S. López and Martha Saboyá · 2009
Earlier work this paper cites.
Choosing the sample size of a computer experiment: A practical guide
Jason L Loeppky, Jerome Sacks, and William J Welch · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms, 2012
Jasper Snoek, Hugo Larochelle, and Ryan P. Adams · 2012
Earlier work this paper cites.
”why should i trust you?”: Explaining the predictions of any classifier, 2016
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Ae: A domain-agnostic platform for adaptive experimentation
Eytan Bakshy, Lili Dworkin, Brian Karrer, Konstantin Kashin, Ben Letham, Ashwin Murthy, and Shaun Singh · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Earlier work this paper cites.
Learning to utilize shaping rewards: a new approach of reward shaping
Yujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang, Yingfeng Chen, Jianye Hao, Feng Wu, and Changjie Fan · 2020
Earlier work this paper cites.
Efficient shapley explanation for features importance estimation under uncertainty
Xiaoxiao Li, Yuan Zhou, Nicha C. Dvornek, Yufeng Gu, Pamela Ventola, and James S. Duncan · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec · 2020
Earlier work this paper cites.
Does bert learn as humans perceive? understanding linguistic styles through lexica
Shirley Anugrah Hayati, Dongyeop Kang, and Lyle Ungar · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Robust multi-objective bayesian optimization under input noise, 2022
Samuel Daulton, Sait Cakmak, Maximilian Balandat, Michael A. Osborne, Enlu Zhou, and Eytan Bakshy · 2022
Earlier work this paper cites.
Karin De Langis and Dongyeop Kang · 2022
Earlier work this paper cites.
Scaling laws for reward model overoptimization, 2022
Leo Gao, John Schulman, and Jacob Hilton · 2022
Cited alongside, same era.
Abhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade, and Sergey Levine · 2022
Cited alongside, same era.
Solving math word problems with process- and outcome-based feedback, 2022
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation model alignment, 2023
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Cited alongside, same era.
Checkpoint merging via bayesian optimization in llm pretraining
Deyuan Liu, Zecheng Wang, Bingning Wang, Weipeng Chen, Chunshan Li, Zhiying Tu, Dianhui Chu, Bo Li, and Dianbo Sui · 2024
Later among the works it cites.
Optimizing instructions and demonstrations for multi-stage language model programs
Krista Opsahl-Ong, Michael J Ryan, Josh Purtell, David Broman, Christopher Potts, Matei Zaharia, and Omar Khattab · 2024
Later among the works it cites.
Token-level proximal policy optimization for query generation
Yichen Ouyang, Lu Wang, Fangkai Yang, Pu Zhao, Chenghua Huang, Jianfeng Liu, Bochen Pang, Yaming Yang, Yuefeng Zhan, Hao Sun, et al · 2024
Later among the works it cites.
From r r to q ∗ q^{*} : Your language model is secretly a q-function, 2024
Rafael Rafailov, Joey Hejna, Ryan Park, and Chelsea Finn · 2024
Later among the works it cites.
Vanishing gradients in reinforcement finetuning of language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Geyang Guo, Ranchi Zhao, Tianyi Tang, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Cited alongside, same era.
Let’s verify step by step, 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
Fine-grained human feedback gives better rewards for language model training, 2023
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Yihua Zhang, Prashant Khanduri, Ioannis Tsaknakis, Yuguang Yao, Mingyi Hong, and Sijia Liu · 2023
Cited alongside, same era.
Bayesian optimization with llm-based acquisition functions for natural language preference elicitation
David Austin, Anton Korikov, Armin Toroghi, and Scott Sanner · 2024
Cited alongside, same era.
Mechanistic interpretability for ai safety–a review
Leonard Bereska and Efstratios Gavves · 2024
Cited alongside, same era.
Enhancing reinforcement learning with dense rewards from language model critic
Meng Cao, Lei Shu, Lei Yu, Yun Zhu, Nevan Wichers, Yinxiao Liu, and Lei Meng · 2024
Cited alongside, same era.
Enhancing reinforcement learning with dense rewards from language model critic
Meng Cao, Lei Shu, Lei Yu, Yun Zhu, Nevan Wichers, Yinxiao Liu, and Lei Meng · 2024
Cited alongside, same era.
Noam Razin, Hattie Zhou, Omid Saremi, Vimal Thilak, Arwen Bradley, Preetum Nakkiran, Joshua Susskind, and Etai Littwin · 2024
Later among the works it cites.
Principled penalty-based methods for bilevel reinforcement learning and rlhf, 2024
Han Shen, Zhuoran Yang, and Tianyi Chen · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Meta Llama Team · 2024
Later among the works it cites.
Text2reward: Reward shaping with language models for reinforcement learning, 2024
Tianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu, Qian Luo, Victor Zhong, Yanchao Yang, and Tao Yu · 2024
Later among the works it cites.
Bayesian reward models for llm alignment
Adam X Yang, Maxime Robeyns, Thomas Coste, Zhengyan Shi, Jun Wang, Haitham Bou Ammar, and Laurence Aitchison · 2024
Later among the works it cites.
Tlcr: Token-level continuous reward for fine-grained reinforcement learning from human feedback
Eunseop Yoon, Hee Suk Yoon, SooHwan Eom, Gunsoo Han, Daniel Nam, Daejin Jo, Kyoung-Woon On, Mark Hasegawa-Johnson, Sungwoong Kim, and Chang Yoo · 2024
Later among the works it cites.
Token-level direct preference optimization
Yongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang, Haifeng Zhang, and Jun Wang · 2024
Later among the works it cites.
Searching for optimal solutions with LLMs via bayesian optimization
Dhruv Agarwal, Manoj Ghuhan Arivazhagan, Rajarshi Das, Sandesh Swamy, Sopan Khosla, and Rashmi Gangadharaiah · 2025
Closest in time.
Unexpected improvements to expected improvement for bayesian optimization, 2025
Sebastian Ament, Samuel Daulton, David Eriksson, Maximilian Balandat, and Eytan Bakshy · 2025
Closest in time.
Length-controlled alpacaeval: A simple way to debias automatic evaluators, 2025
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B. Hashimoto · 2025
Closest in time.
Reward shaping to mitigate reward hacking in rlhf, 2025
Jiayi Fu, Xuandong Zhao, Chengyuan Yao, Heng Wang, Qi Han, and Yanghua Xiao · 2025
Closest in time.
Alphapo – reward shape matters for llm alignment, 2025
Aman Gupta, Shao Tang, Qingquan Song, Sirou Zhu, Jiwoo Hong, Ankan Saha, Viral Gupta, Noah Lee, Eunki Kim, Siyu Zhu, Parag Agrawal, Natesh Pillai, and S. Sathiya Keerthi · 2025
Closest in time.
Align to structure: Aligning large language models with structural information, 2025
Zae Myung Kim, Anand Ramachandran, Farideh Tavazoee, Joo-Kyung Kim, Oleg Rokhlenko, and Dongyeop Kang · 2025
Closest in time.
Dpo meets ppo: Reinforced token optimization for rlhf, 2025
Han Zhong, Zikang Shan, Guhao Feng, Wei Xiong, Xinle Cheng, Li Zhao, Di He, Jiang Bian, and Liwei Wang · 2025
Closest in time.