Fetching the paper…
Reading the bibliography…
In many real world applications, reinforcement learning agents have to optimize multiple objectives while following certain rules or satisfying a list of constraints.
Q-learning
Watkins, C. J. C. H.; and Dayan, P. 1992 · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. 1994 · 1994
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E. 1999 · 1999
Earlier work this paper cites.
Nonlinear Programming
Bertsekas, D. 1999 · 1999
Earlier work this paper cites.
Congested traffic states in empirical observations and microscopic simulations
Treiber, M.; Hennecke, A.; and Helbing, D. 2000 · 2000
Earlier work this paper cites.
An actor-critic algorithm for constrained Markov decision processes
Borkar, V. 2005 · 2004
Earlier work this paper cites.
A Geometric Approach to Multi-Criterion Reinforcement Learning
Mannor, S.; Shimkin, N.; and Mahadevan, S. 2004 · 2004
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2012 · 2012
Earlier work this paper cites.
An Online Actor–Critic Algorithm with Function Approximation for Constrained Markov Decision Processes
Bhatnagar, S.; and Lakshmanan, K. 2012 · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Cited alongside, same era.
Multi-objectivization of reinforcement learning problems by reward shaping
Brys, T.; Harutyunyan, A.; Vrancx, P.; Taylor, M. E.; Kudenko, D.; and Nowe, A. 2014 · 2014
Cited alongside, same era.
Multi-objective Reinforcement Learning with Continuous Pareto Frontier Approximation
Pirotta, M.; Parisi, S.; and Restelli, M. 2014 · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M. A.; Fidjeland, A.; Ostrovski, G.; Petersen, S.; Beattie, C.; Sadik, A.; Antonoglou, I.; King, H.; Kumaran, D.; Wierstra, D.; Legg, S.; and Hassabis, D. 2015 · 2015
Cited alongside, same era.
Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
Watter, M.; Springenberg, J. T.; Boedecker, J.; and Riedmiller, M. A. 2015 · 2015
Cited alongside, same era.
Safe Exploration in Continuous Action Spaces
Dalal, G.; Dvijotham, K.; Vecerík, M.; Hester, T.; Paduraru, C.; and Tassa, Y. 2018 · 2018
Later among the works it cites.
The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems
Krajewski, R.; Bock, J.; Kloeker, L.; and Eckstein, L. 2018 · 2018
Later among the works it cites.
High-level Decision Making for Safe and Reasonable Autonomous Lane Changing using Reinforcement Learning
Mirchevska, B.; Pek, C.; Werling, M.; Althoff, M.; and Boedecker, J. 2018 · 2018
Later among the works it cites.
Reward Constrained Policy Optimization
Tessler, C.; Mankowitz, D. J.; and Mannor, S. 2018 · 2018
Later among the works it cites.
A Scalable Framework For Real-Time Multi-Robot, Multi-Human Collision Avoidance
Bajcsy, A.; Herbert, S. L.; Fridovich-Keil, D.; Fisac, J. F.; Deglurkar, S.; Dragan, A. D.; and Tomlin, C. J. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
End-to-End Training of Deep Visuomotor Policies
Levine, S.; Finn, C.; Darrell, T.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; van den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; Dieleman, S.; Grewe, D.; Nham, J.; Kalchbrenner, N.; Sutskever, I.; Lillicrap, T. P.; Leach, M.; Kavukcuoglu, K.; Graepel, T.; and Hassabis, D. 2016 · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J.; Held, D.; Tamar, A.; and Abbeel, P. 2017 · 2017
Cited alongside, same era.
Tactical decision making for lane changing with deep reinforcement learning
Mukadam, M.; Cosgun, A.; Nakhaei, A.; and Fujimura, K. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks
Cheng, R.; Orosz, G.; Murray, R. M.; and Burdick, J. W. 2019 · 2019
Later among the works it cites.
Safely Probabilistically Complete Real-Time Planning and Exploration in Unknown Environments
Fridovich-Keil, D.; Fisac, J. F.; and Tomlin, C. J. 2019 · 2019
Later among the works it cites.
Dynamic Input for Deep Reinforcement Learning in Autonomous Driving
Huegle, M.; Kalweit, G.; Mirchevska, B.; Werling, M.; and Boedecker, J. 2019 · 2019
Later among the works it cites.
Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization
Kalweit, G.; Huegle, M.; and Boedecker, J. 2019 · 2019
Later among the works it cites.