Fetching the paper…
Reading the bibliography…
Many advances that have improved the robustness and efficiency of deep reinforcement learning (RL) algorithms can, in one way or another, be understood as introducing additional objectives or constraints in the policy optimization step.
A closer look at drawbacks of minimizing weighted sums of objectives for pareto set generation in multicriteria optimization problems
I. Das and J. E. Dennis · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin A. Riedmiller · 2002
Earlier work this paper cites.
Q-decomposition for reinforcement learning agents
Stuart Russell and Andrew L. Zimdars · 2003
Earlier work this paper cites.
On test functions for evolutionary multi-objective optimization
Tatsuya Okabe, Yaochu Jin, Markus Olhofer, and Bernhard Sendhoff · 2004
Earlier work this paper cites.
An operator view of policy gradient methods
Dibya Ghosh, Marlos C. Machado, and Nicolas Le Roux · 2006
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altün · 2010
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Diederik M. Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Earlier work this paper cites.
Scalarized multi-objective reinforcement learning: Novel design techniques
Kristof Van Moffaert, Madalina M Drugan, and Ann Nowé · 2013
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of Pareto dominating policies
Kristof Van Moffaert and Ann Nowé · 2014
Earlier work this paper cites.
Linear support for multi-objective coordination graphs
Diederik M. Roijers, Shimon Whiteson, and Frans A. Oliehoek · 2014
Earlier work this paper cites.
Reinforcement learning from demonstration through shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E. Taylor, and Ann Nowé · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Multi-objective deep reinforcement learning
Hossam Mossalam, Yannis M. Assael, Diederik M. Roijers, and Shimon Whiteson · 2016
Cited alongside, same era.
Multi-objective reinforcement learning through continuous pareto manifold approximation
Simone Parisi, Matteo Pirotta, and Marcello Restelli · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
ϵ \epsilon -pal: An active learning approach to the multi-objective optimization problem
Marcela Zuluaga, Andreas Krause, and Markus Püschel · 2016
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
Pareto-DQN: Approximating the Pareto front in complex multi-objective decision problems
Mathieu Reymond and Ann Nowé · 2019
Later among the works it cites.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan · 2019
Later among the works it cites.
A distributional view on multi-objective policy optimization
Abbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert, Francis Song, Martina Zambelli, Murilo F. Martins, Nicolas Heess, Raia Hadsell, and Martin Reidmiller · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
Nicolas Heess, Dhruva TB, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, S. M. Ali Eslami, Martin Riedmiller, and David Silver · 2017
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Manifold-based multi-objective policy search with sample reuse
Simone Parisi, Matteo Pirotta, and Jan Peters · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Distributional policy gradients
Gabriel Barth-Maron, Matthew W. Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva TB, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Non-adversarial imitation learning and its connections to adversarial methods, 2020
Oleg Arenz and Gerhard Neumann · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Random hypervolume scalarizations for provable multi-objective black box optimization
Daniel Golovin and Qiuyi Zhang · 2020
Later among the works it cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, Jerry Li, Mohammad Norouzi, Matthew Hoffman, Nicolas Heess, and Nando de Freitas · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Offline reinforcement learning hands-on, 2020
Louis Monier, Jakub Kmec, Alexandre Laterre, Thomas Pierrot, Valentin Courgeau, Olivier Sigaud, and Karim Beguir · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2020
Later among the works it cites.
Critic regularized regression
Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh S Merel, Jost Tobias Springenberg, Scott E Reed, Bobak Shahriari, Noah Siegel, Caglar Gulcehre, Nicolas Heess, and Nando de Freitas · 2020
Later among the works it cites.
Prediction-guided multi-objective reinforcement learning for continuous robot control
Jie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus, Shinjiro Sueda, and Wojciech Matusik · 2020
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Closest in time.
Online and offline reinforcement learning by planning with a learned model
Julian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain, Ioannis Antonoglou, and David Silver · 2021
Closest in time.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Tengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong, and Yu Bai · 2021
Closest in time.