Fetching the paper…
Reading the bibliography…
Reward design is a fundamental, yet challenging aspect of reinforcement learning (RL).
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Automatic shaping and decomposition of reward functions
Bhaskara Marthi · 2007
Earlier work this paper cites.
Learning potential for reward shaping in reinforcement learning with tile coding
Marek Grzes and Daniel Kudenko · 2008
Earlier work this paper cites.
Dynamic reward shaping: training a robot by voice
Ana C Tenorio-Gonzalez, Eduardo F Morales, and Luis Villasenor-Pineda · 2010
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
Abdeslam Boularias, Jens Kober, and Jan Peters · 2011
Earlier work this paper cites.
Dynamic potential-based reward shaping
Sam Michael Devlin and Daniel Kudenko · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Kristof Van Moffaert and Ann Nowé · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Deep direct reinforcement learning for financial signal representation and trading
Yue Deng, Feng Bao, Youyong Kong, Zhiquan Ren, and Qionghai Dai · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Coordinated deep reinforcement learners for traffic light control
Elise Van der Pol and Frans A Oliehoek · 2016
Earlier work this paper cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Erwin Coumans and Yunfei Bai · 2016
Earlier work this paper cites.
Multi-objective deep reinforcement learning
Hossam Mossalam, Yannis M Assael, Diederik M Roijers, and Shimon Whiteson · 2016
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Cited alongside, same era.
Flow: A modular learning framework for autonomy in traffic
Cathy Wu, Aboudy Kreidieh, Kanaad Parvate, Eugene Vinitsky, and Alexandre M Bayen · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Cited alongside, same era.
Learning dense rewards for contact-rich manipulation tasks
Zheng Wu, Wenzhao Lian, Vaibhav Unhelkar, Masayoshi Tomizuka, and Stefan Schaal · 2021
Later among the works it cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Later among the works it cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Later among the works it cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Multi-objectivization and ensembles of shapings in reinforcement learning
Tim Brys, Anna Harutyunyan, Peter Vrancx, Ann Nowé, and Matthew E Taylor · 2017
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Intellilight: A reinforcement learning approach for intelligent traffic light control
Hua Wei, Guanjie Zheng, Huaxiu Yao, and Zhenhui Li · 2018
Cited alongside, same era.
Zhi Zhang, Jiachen Yang, and Hongyuan Zha · 2019
Cited alongside, same era.
Automatic successive reinforcement learning with multiple auxiliary rewards
Zhao-Yang Fu, De-Chuan Zhan, Xin-Chun Li, and Yi-Xing Lu · 2019
Cited alongside, same era.
Evolving rewards to automate reinforcement learning
Aleksandra Faust, Anthony Francis, and Dar Mehta · 2019
Cited alongside, same era.
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Later among the works it cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu Hong Hoi · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Automated reinforcement learning: An overview
Reza Refaei Afshar, Yingqian Zhang, Joaquin Vanschoren, and Uzay Kaymak · 2022
Later among the works it cites.
Automated reinforcement learning (autorl): A survey and open problems
Jack Parker-Holder, Raghu Rajan, Xingyou Song, André Biedenkapp, Yingjie Miao, Theresa Eimer, Baohe Zhang, Vu Nguyen, Roberto Calandra, Aleksandra Faust, et al · 2022
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Later among the works it cites.
The perils of trial-and-error reward design: misdesign through overfitting and invalid task specifications
Serena Booth, Bradley W Knox, Julie Shah, Scott Niekum, Peter Stone, and Alessandro Allievi · 2023
Closest in time.
Execution-based code generation using deep reinforcement learning
Parshin Shojaee, Aneesh Jain, Sindhu Tipirneni, and Chandan K Reddy · 2023
Closest in time.
Phi-2: The surprising power of small language models, 2023
Mojan Javaheripi, Sébastien Bubeck, et al · 2023
Closest in time.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Polak Scowcroft, Neel Kant, Aidan Swope, et al · 2023
Closest in time.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Rlhf workflow: From reward modeling to online rlhf
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees, 2024
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao Peng, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.