Fetching the paper…
Reading the bibliography…
The reward signal plays a central role in defining the desired behaviors of agents in reinforcement learning (RL).
Ohio Supercomputer Center
Center, O. S. 1987 · 1987
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y.; Harada, D.; and Russell, S. 1999 · 1999
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Hutter, M. 2005 · 2005
Earlier work this paper cites.
Noisy reinforcements in reinforcement learning: some case studies based on gridworlds
Moreno, A.; Martín, J. D.; Soria, E.; Magdalena, R.; and Martínez, M. 2006 · 2006
Earlier work this paper cites.
Delusion, survival, and intelligent agents
Ring, M.; and Orseau, L. 2011 · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. 2014 · 2014
Earlier work this paper cites.
Robust bayesian inverse reinforcement learning with sparse behavior noise
Zheng, J.; Liu, S.; and Ni, L. M. 2014 · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Mané, D. 2016 · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G.; Dabney, W.; and Munos, R. 2017 · 2017
Cited alongside, same era.
Reinforcement learning with a corrupted reward channel
Everitt, T.; Krakovna, V.; Orseau, L.; Hutter, M.; and Legg, S. 2017 · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D.; Milli, S.; Abbeel, P.; Russell, S. J.; and Dragan, A. 2017 · 2017
Cited alongside, same era.
Robust deep reinforcement learning with adversarial attacks
Pattanaik, A.; Tang, Z.; Liu, S.; Bommannan, G.; and Chowdhary, G. 2017 · 2017
Cited alongside, same era.
Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning
Rakhsha, A.; Radanovic, G.; Devidze, R.; Zhu, X.; and Singla, A. 2020 · 2020
Later among the works it cites.
Reinforcement learning with perturbed rewards
Wang, J.; Liu, Y.; and Li, B. 2020 · 2020
Later among the works it cites.
Reward is enough
Silver, D.; Singh, S.; Precup, D.; and Sutton, R. S. 2021 · 2021
Later among the works it cites.
No-regret reinforcement learning with heavy-tailed rewards
Zhuang, V.; and Sui, Y. 2021 · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Later among the works it cites.
Reinforcement learning with stochastic reward machines
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust adversarial reinforcement learning
Pinto, L.; Davidson, J.; Sukthankar, R.; and Gupta, A. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Reward estimation for variance reduction in deep reinforcement learning
Romoff, J.; Henderson, P.; Piché, A.; Francois-Lavet, V.; and Pineau, J. 2018 · 2018
Cited alongside, same era.
An analysis of categorical distributional reinforcement learning
Rowland, M.; Bellemare, M.; Dabney, W.; Munos, R.; and Teh, Y. W. 2018 · 2018
Cited alongside, same era.
Provably robust blackbox optimization for reinforcement learning
Choromanski, K.; Pacchiano, A.; Parker-Holder, J.; Tang, Y.; Jain, D.; Yang, Y.; Iscen, A.; Hsu, J.; and Sindhwani, V. 2020 · 2020
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
Dabney, W.; Ostrovski, G.; Silver, D.; and Munos, R. 2018a
Cited in the paper.
Distributional reinforcement learning with quantile regression
Dabney, W.; Rowland, M.; Bellemare, M.; and Munos, R. 2018b
Cited in the paper.
Corazza, J.; Gavran, I.; and Neider, D. 2022 · 2022
Later among the works it cites.
Distributional Reward Estimation for Effective Multi-Agent Deep Reinforcement Learning
Hu, J.; Sun, Y.; Chen, H.; Huang, S.; Chang, Y.; Sun, L.; et al. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Later among the works it cites.
Learning Rewards to Optimize Global Performance Metrics in Deep Reinforcement Learning
Qian, J.; Weng, P.; and Tan, C. 2023 · 2023
Later among the works it cites.
A Long N-step Surrogate Stage Reward for Deep Reinforcement Learning
Zhong, J.; Wu, R.; and Si, J. 2023 · 2023
Later among the works it cites.