Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) is a widely used technique for aligning Large Language Models (LLMs) with human preferences, yet it often suffers from sparse reward signals, making effective credit assignment challenging.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Cooperative games with coalition structures
Robert J Aumann and Jacques H Dreze · 1974
Earlier work this paper cites.
Values of games with a priori unions
Guilliermo Owen · 1977
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard Stuart Sutton · 1984
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 2004
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Framework of automatic text summarization using reinforcement learning
Seonggi Ryang and Takeshi Abekawa · 2012
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao · 2016
Earlier work this paper cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Earlier work this paper cites.
TL;DR: Mining Reddit to learn automatic summarization
Michael V"olske, Martin Potthast, Shahbaz Syed, and Benno Stein · 2017
Earlier work this paper cites.
Ask the right questions: Active question reformulation with reinforcement learning
Christian Buck, Jannis Bulian, Massimiliano Ciaramita, Wojciech Gajewski, Andrea Gesmundo, Neil Houlsby, and Wei Wang · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Towards sample efficient reinforcement learning
Yang Yu · 2018
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Cited alongside, same era.
Dealing with sparse rewards in reinforcement learning
Joshua Hare · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace · 2019
Cited alongside, same era.
Learning to solve the credit assignment problem
Benjamin James Lansdell, Prashanth Ravi Prakash, and Konrad Paul Kording · 2019
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi · 2023
Later among the works it cites.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Later among the works it cites.
Prompt valuation based on shapley values
Hanxi Liu, Xiaokai Mao, Haocheng Xia, Jian Lou, and Jinfei Liu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
A duality approach for regret minimization in average-award ergodic markov decision processes
Hao Gong and Mengdi Wang · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Enhancing reinforcement learning with dense rewards from language model critic
Meng Cao, Lei Shu, Lei Yu, Yun Zhu, Nevan Wichers, Yinxiao Liu, and Lei Meng · 2024
Later among the works it cites.
Dense reward for free in reinforcement learning from human feedback
Alex James Chan, Hao Sun, Samuel Holt, and Mihaela van der Schaar · 2024
Later among the works it cites.
Tokenshap: Interpreting large language models with monte carlo shapley value estimation
Roni Goldshmidt and Miriam Horovicz · 2024
Later among the works it cites.
Shed: Shapley-based automated dataset refinement for instruction fine-tuning
Yexiao He, Ziyao Wang, Zheyu Shen, Guoheng Sun, Yucong Dai, Yongkai Wu, Hongyi Wang, and Ang Li · 2024
Later among the works it cites.
The n+ implementation details of RLHF with PPO: A case study on TL;DR summarization
Shengyi Huang, Michael Noukhovitch, Arian Hosseini, Kashif Rasul, Weixun Wang, and Lewis Tunstall · 2024
Later among the works it cites.
Motif: Intrinsic motivation from artificial intelligence feedback
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, and Mikael Henaff · 2024
Later among the works it cites.
Explaining large language models decisions using shapley values
Behnam Mohammadi · 2024
Later among the works it cites.
Vanishing gradients in reinforcement finetuning of language models
Noam Razin, Hattie Zhou, Omid Saremi, Vimal Thilak, Arwen Bradley, Preetum Nakkiran, Joshua M. Susskind, and Etai Littwin · 2024
Later among the works it cites.
TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback
Eunseop Yoon, Hee Suk Yoon, SooHwan Eom, Gunsoo Han, Daniel Nam, Daejin Jo, Kyoung-Woon On, Mark Hasegawa-Johnson, Sungwoong Kim, and Chang Yoo · 2024
Later among the works it cites.
Learning explainable dense reward shapes via bayesian optimization
Ryan Koo, Ian Yang, Vipul Raheja, Mingyi Hong, Kwang-Sung Jun, and Dongyeop Kang · 2025
Closest in time.
Efficient shapley value-based non-uniform pruning of large language models
Chuan Sun, Han Yu, and Lizhen Cui · 2025
Closest in time.