Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) has achieved tremendous success in many complex decision-making tasks.
Über die lage der integralkurven gewöhnlicher differentialgleichungen
Mitio Nagumo · 1942
Earlier work this paper cites.
Lagrange multipliers revisited: a contribution to non-linear programming
Morton Slater · 1950
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
David Blackwell · 1956
Earlier work this paper cites.
Control system analysis and design via the second method of lyapunov:(i) continuous-time systems (ii) discrete time systems
R Kalman and J Bertram · 1959
Earlier work this paper cites.
The problem of abortion and the doctrine of the double effect
Philippa Foot · 1967
Earlier work this paper cites.
Risk-sensitive markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
Killing, letting die, and the trolley problem
Judith Jarvis Thomson · 1976
Earlier work this paper cites.
The temporal logic of programs
Amir Pnueli · 1977
Earlier work this paper cites.
A behavioural car-following model for computer simulation
Peter G Gipps · 1981
Earlier work this paper cites.
Linear programming and finite Markovian control problems
Lodewijk Kallenberg · 1983
Earlier work this paper cites.
A theory of the learnable
G. Valiant Leslie · 1984
Earlier work this paper cites.
Optimal policies for controlled markov chains with a constraint
Frederick J Beutler and Keith W Ross · 1985
Earlier work this paper cites.
Time-average optimal constrained semi-markov decision processes
Frederick J Beutler and Keith W Ross · 1986
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Randomized and past-dependent policies for markov decision processes with multiple constraints
Keith W Ross · 1989
Earlier work this paper cites.
Markov decision processes with sample path constraints: the communicating case
Keith W Ross and Ravi Varadarajan · 1989
Earlier work this paper cites.
Adaptive control: Stability, convergence, and robustness
Shankar Sastry, Marc Bodson, and James F Bartram · 1990
Earlier work this paper cites.
Multichain markov decision processes with a sample path constraint: A decomposition approach
Keith W Ross and Ravi Varadarajan · 1991
Earlier work this paper cites.
The effect of delayed feedback information on network performance
Andreas D Bovopoulos and Aurel A Lazar · 1992
Earlier work this paper cites.
A teaching method for reinforcement learning
Jeffery A Clouse and Paul E Utgoff · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Packet routing in dynamically changing networks: A reinforcement learning approach
Justin Boyan and Michael Littman · 1993
Earlier work this paper cites.
Consideration of risk in reinforcement learning
Matthias Heger · 1994
Earlier work this paper cites.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman · 1995
Earlier work this paper cites.
Analysis of adaptive step-size sa algorithms for parameter tracking
H. J. Kushner and J. Yang · 1995
Earlier work this paper cites.
Microeconomic theory
Andreu Mas-Colell, Michael Dennis Whinston, Jerry R Green, et al · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro et al · 1995
Earlier work this paper cites.
Robust and optimal control
John Doyle · 1996
Earlier work this paper cites.
Elevator group control using multiple reinforcement learning agents
Robert H Crites and Andrew G Barto · 1998
Earlier work this paper cites.
Multi-criteria reinforcement learning
Zoltán Gábor, Zsolt Kalmár, and Csaba Szepesvári · 1998
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Distributed value functions
Jeff Schneider, Weng-Keen Wong, Andrew Moore, and Martin Riedmiller · 1999
Earlier work this paper cites.
Optimal stopping of markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
John N Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
Constrained discounted markov decision processes and hamiltonian cycles
Eugene A Feinberg · 2000
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin Riedmiller · 2000
Earlier work this paper cites.
Constrained model predictive control: Stability and optimality
David Q Mayne, James B Rawlings, Christopher V Rao, and Pierre OM Scokaert · 2000
Earlier work this paper cites.
Asymptopia: an exposition of statistical asymptotic theory. 2000
David Pollard · 2000
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
J Andrew Bagnell and Jeff G Schneider · 2001
Earlier work this paper cites.
Lyapunov-constrained action sets for reinforcement learning
Theodore J Perkins and Andrew G Barto · 2001
Earlier work this paper cites.
Td algorithm for the variance of return and mean-variance reinforcement learning
Makoto Sato, Hajime Kimura, and Shibenobu Kobayashi · 2001
Earlier work this paper cites.
Self learning control of constrained markov chains-a gradient approach
F Vazquez Abad, Vikram Krishnamurthy, Katerine Martin, and Irina Baltcheva · 2002
Earlier work this paper cites.
Q-learning for risk-sensitive control
Vivek S Borkar · 2002
Earlier work this paper cites.
Portfolio optimization with conditional value-at-risk objective and constraints
Pavlo Krokhmal, Jonas Palmquist, and Stanislav Uryasev · 2002
Earlier work this paper cites.
Envelope theorems for arbitrary choice sets
Paul Milgrom and Ilya Segal · 2002
Earlier work this paper cites.
Lyapunov design for safe reinforcement learning
Theodore J Perkins and Andrew G Barto · 2002
Earlier work this paper cites.
Self learning control of constrained markov decision processes–a gradient approach
Felisa J Vázquez Abad and Vikram Krishnamurthy · 2003
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2003
Earlier work this paper cites.
An introduction to numerical analysis
Endre Süli and David F Mayers · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
Nate Kohl and Peter Stone · 2004
Earlier work this paper cites.
Cyberbotics ltd. webots™: professional mobile robot simulation
Olivier Michel · 2004
Earlier work this paper cites.
Convergence properties of policy iteration
Manuel S Santos and John Rust · 2004
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
Robust model predictive control of constrained linear systems with bounded disturbances
David Q Mayne, María M Seron, and SV Raković · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
An mdp-based recommender system
Guy Shani, David Heckerman, Ronen I Brafman, and Craig Boutilier · 2005
Earlier work this paper cites.
Stochastic iterative dynamic programming: A monte carlo approach to dual control
A. M. Thompson and W. R. Cluett · 2005
Earlier work this paper cites.
Extremely randomized trees
Pierre Geurts, Damien Ernst, and Louis Wehenkel · 2006
Earlier work this paper cites.
Discounted markov decision processes with utility constraints
Yoshinobu Kadota, Masami Kurano, and Masami Yasuda · 2006
Earlier work this paper cites.
Reinforcement learning with human teachers: evidence of feedback and guidance with implications for learning performance
Andrea L Thomaz and Cynthia Breazeal · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Christopher K Williams and Carl Edward Rasmussen · 2006
Earlier work this paper cites.
Percentile optimization in uncertain markov decision processes with application to efficient exploration
Erick Delage and Shie Mannor · 2007
Earlier work this paper cites.
A learning algorithm for risk-sensitive cost
Arnab Basu, Tirthankar Bhattacharyya, and Vivek S Borkar · 2008
Earlier work this paper cites.
Safe exploration for reinforcement learning
Alexander Hans, Daniel Schneegaß, Anton Maximilian Schäfer, and Steffen Udluft · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Computing var and cvar using stochastic approximation and adaptive unconstrained importance sampling
Olivier Bardou, Noufel Frikha, and Gilles Pages · 2009
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint
Vivek S Borkar · 2009
Earlier work this paper cites.
Mechatronic design of nao humanoid
David Gouaillier, Vincent Hugel, Pierre Blazevic, Chris Kilner, Jérôme Monceaux, Pascal Lafourcade, Brice Marnier, Julien Serre, and Bruno Maisonnier · 2009
Earlier work this paper cites.
Policy search for motor primitives in robotics
J Kober and J Peters · 2009
Earlier work this paper cites.
Regret-based reward elicitation for markov decision processes
Kevin Regan and Craig Boutilier · 2009
Earlier work this paper cites.
Commercial mobile robot simulation software
CL Webots · 2009
Earlier work this paper cites.
Formal methods: Practice and experience
Jim Woodcock, Peter Gorm Larsen, Juan Bicarregui, and John Fitzgerald · 2009
Earlier work this paper cites.
Autonomous helicopter aerobatics through apprenticeship learning
Pieter Abbeel, Adam Coates, and Andrew Y Ng · 2010
Earlier work this paper cites.
Optimizing debt collections using constrained reinforcement learning
Naoki Abe, Prem Melville, Cezar Pendus, Chandan K Reddy, David L Jensen, Vince P Thomas, James J Bennett, Gary F Anderson, Brent R Cooley, Melissa Kowalczyk, et al · 2010
Earlier work this paper cites.
Learning multi-agent state space representations
Yann-Michaël De Hauwere, Peter Vrancx, and Ann Nowé · 2010
Earlier work this paper cites.
Biped walk learning through playback and corrective demonstration
Çetin Meriçli and Manuela Veloso · 2010
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Earlier work this paper cites.
Oxford dictionary of English
Angus Stevenson · 2010
Earlier work this paper cites.
Parameterized maneuver learning for autonomous helicopter flight
Jie Tang, Arjun Singh, Nimbus Goehausen, and Pieter Abbeel · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
L1 adaptive control theory: Guaranteed robustness with fast adaptation
NAIRA HOVAKIMYAN and CHENGYU CAO · 2011
Earlier work this paper cites.
Neuroevolutionary reinforcement learning for generalized control of simulated helicopters
Rogier Koppejan and Shimon Whiteson · 2011
Earlier work this paper cites.
Control and optimization meet the smart power grid: Scheduling of power demands for optimal energy management
Iordanis Koutsopoulos and Leandros Tassiulas · 2011
Earlier work this paper cites.
Mean-variance optimization in markov decision processes
Shie Mannor and John N Tsitsiklis · 2011
Earlier work this paper cites.
Fast reinforcement learning for energy-efficient wireless communication
Nicholas Mastronarde and Mihaela van der Schaar · 2011
Earlier work this paper cites.
Model learning for robot control: a survey
Duy Nguyen-Tuong and Jan Peters · 2011
Earlier work this paper cites.
Biped robot walking using particle swarm optimization
Behzad Nikbin, Mohammad Reza Ranjbar, B Shafiee Sarjaz, and Hamed Shah-Hosseini · 2011
Earlier work this paper cites.
Autonomous climbing of spiral staircases with humanoids
Stefan Oßwald, Attila Görög, Armin Hornung, and Maren Bennewitz · 2011
Earlier work this paper cites.
Automated generation of cpg-based locomotion for robot nao
Ernesto Torres and Leonardo Garrido · 2011
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
Safe exploration of state and action spaces in reinforcement learning
Javier Garcia and Fernando Fernández · 2012
Earlier work this paper cites.
Pac bounds for discounted mdps
Tor Lattimore and Marcus Hutter · 2012
Earlier work this paper cites.
Nao walking down a ramp autonomously
Christian Lutz, Felix Atmanspacher, Armin Hornung, and Maren Bennewitz · 2012
Earlier work this paper cites.
Risk aversion in markov decision processes via near optimal chernoff bounds
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Safe exploration in markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Constructive nonlinear control
Rodolphe Sepulchre, Mrdjan Jankovic, and Petar V Kokotovic · 2012
Earlier work this paper cites.
Policy gradients with variance related risk criteria
Aviv Tamar, Dotan Di Castro, and Shie Mannor · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Hardware experiments of humanoid robot safe fall using aldebaran nao
Seung-Kook Yun and Ambarish Goswami · 2012
Earlier work this paper cites.
Provably safe and robust learning-based model predictive control
Anil Aswani, Humberto Gonzalez, S Shankar Sastry, and Claire Tomlin · 2013
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Humanoid robots learning to walk faster: From the real world to simulation and back
Alon Farchy, Samuel Barrett, Patrick MacAlpine, and Peter Stone · 2013
Earlier work this paper cites.
Intelligent cooperative control architecture: a framework for performance improvement using safe learning
Alborz Geramifard, Joshua Redding, and Jonathan P How · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Neural network reinforcement learning for visual control of robot manipulators
Zoran Miljković, Marko Mitić, Mihailo Lazarević, and Bojan Babić · 2013
Earlier work this paper cites.
Reachability-based safe learning with gaussian processes
Anayo K Akametalu, Jaime F Fisac, Jeremy H Gillula, Shahab Kaynama, Melanie N Zeilinger, and Claire J Tomlin · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Policy gradients beyond expectations: Conditional value-at-risk
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2014
Earlier work this paper cites.
Safe and robust learning control with gaussian processes
Felix Berkenkamp and Angela P Schoellig · 2015
Earlier work this paper cites.
Convex optimization algorithms
Dimitri Bertsekas · 2015
Earlier work this paper cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Chance-constrained dynamic programming with application to risk-aware robotic space exploration
Masahiro Ono, Marco Pavone, Yoshiaki Kuwata, and J Balaram · 2015
Earlier work this paper cites.
Distributed coordination control for multi-robot networks using lyapunov-like barrier functions
Dimitra Panagou, Dušan M Stipanović, and Petros G Voulgaris · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Safe exploration for optimization with gaussian processes
Yanan Sui, Alkis Gotovos, Joel Burdick, and Andreas Krause · 2015
Earlier work this paper cites.
Optimizing the cvar via sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2015
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Erwin Coumans and Yunfei Bai · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Multi-Principal Assistance Games
Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell, and Stuart Russell · 2020
Later among the works it cites.
L1-gp: L1 adaptive control with bayesian learning
Aditya Gahlawat, Pan Zhao, Andrew Patterson, Naira Hovakimyan, and Evangelos Theodorou · 2020
Later among the works it cites.
Teaching a humanoid robot to walk faster through safe reinforcement learning
Javier García and Diogo Shafie · 2020
Later among the works it cites.
A motion planning method for unmanned surface vehicle in restricted waters
Shangding Gu, Chunhui Zhou, Yuanqiao Wen, Xi Zhong, Man Zhu, Changshi Xiao, and Zhe Du · 2020
Later among the works it cites.
Minghao Han, Lixian Tian, Yuanand Zhang, Jun Wang, and Wei Pan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convex synthesis of randomized policies for controlled markov chains with density safety upper bound constraints
Mahmoud El Chamie, Yue Yu, and Behçet Açıkmeşe · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Cited alongside, same era.
Safety-constrained reinforcement learning for mdps
Sebastian Junges, Nils Jansen, Christian Dehnert, Ufuk Topcu, and Joost-Pieter Katoen · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Robust constrained learning-based nmpc enabling reliable mobile robot path tracking
Chris J Ostafew, Angela P Schoellig, and Timothy D Barfoot · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
Limit-deterministic büchi automata for linear temporal logic
Salomon Sickert, Javier Esparza, Stefan Jaax, and Jan Křetínskỳ · 2016
Cited alongside, same era.
Cautious reinforcement learning with logical constraints
Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening · 2020
Later among the works it cites.
Controller design via experimental exploration with robustness guarantees
Tobias Holicki, Carsten W Scherer, and Sebastian Trimpe · 2020
Later among the works it cites.
Voronoi-based multi-robot autonomous exploration in unknown environments via deep reinforcement learning
Junyan Hu, Hanlin Niu, Joaquin Carrasco, Barry Lennox, and Farshad Arvin · 2020
Later among the works it cites.
Subin Huh and Insoon Yang · 2020
Later among the works it cites.
Safe reinforcement learning using probabilistic shields
Nils Jansen, Bettina Könighofer, Sebastian Junges, AC Serban, and Roderick Bloem · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Gabriel Kalweit, Maria Huegle, Moritz Werling, and Joschka Boedecker · 2020
Later among the works it cites.
Safe learning and optimization techniques: Towards a survey of the state of the art
Youngmin Kim, Richard Allmendinger, and Manuel López-Ibáñez · 2020
Later among the works it cites.
Safe reinforcement learning for autonomous lane changing using set-based prediction
Hanna Krasowski, Xiao Wang, and Matthias Althoff · 2020
Later among the works it cites.
A constrained reinforcement learning based approach for network slicing
Yongshuai Liu, Jiaxin Ding, and Xin Liu · 2020
Later among the works it cites.
Ipo: Interior-point policy optimization under constraints
Yongshuai Liu, Jiaxin Ding, and Xin Liu · 2020
Later among the works it cites.
Safe off-policy reinforcement learning using barrier functions
Zahra Marvi and Bahare Kiumarsi · 2020
Later among the works it cites.
Reinforcement learning via fenchel-rockafellar duality
Ofir Nachum and Bo Dai · 2020
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian A Schroeder de Witt, Pierre-Alexandre Kamienny, Philip HS Torr, Wendelin Böhmer, and Shimon Whiteson · 2020
Later among the works it cites.
Robust constrained-mdps: Soft-constrained robust policy optimization under model uncertainty
Reazul Hasan Russel, Mouhacine Benosman, and Jeroen Van Baar · 2020
Later among the works it cites.
Constrained markov decision processes via backward value functions
Harsh Satija, Philip Amortila, and Joelle Pineau · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Hierarchical multiagent reinforcement learning for maritime traffic management
Arambam James Singh, Akshat Kumar, and Hoong Chuin Lau · 2020
Later among the works it cites.
Building healthy recommendation sequences for everyone: A safe reinforcement learning approach
Ashudeep Singh, Yoni Halpern, Nithum Thain, Konstantina Christakopoulou, EH Chi, Jilin Chen, and Alex Beutel · 2020
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine · 2020
Later among the works it cites.
Safe reinforcement learning in constrained markov decision processes
Akifumi Wachi and Yanan Sui · 2020
Later among the works it cites.
Safe reinforcement learning for autonomous vehicles through parallel constrained policy optimization
Lu Wen, Jingliang Duan, Shengbo Eben Li, Shaobing Xu, and Huei Peng · 2020
Later among the works it cites.
Reinforcement learning-based mobile offloading for edge computing against jamming and interference
Liang Xiao, Xiaozhen Lu, Tangwei Xu, Xiaoyue Wan, Wen Ji, and Yanyong Zhang · 2020
Later among the works it cites.
A primal approach to constrained policy optimization: Global optimality and finite-time analysis
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2020
Later among the works it cites.
Projection-based constrained policy optimization
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J Ramadge · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Safe reinforcement learning using robust mpc
Mario Zanon and Sébastien Gros · 2020
Later among the works it cites.
Robust multi-agent reinforcement learning with model uncertainty
Kaiqing Zhang, Tao Sun, Yunzhe Tao, Sahika Genc, Sunil Mallya, and Tamer Basar · 2020
Later among the works it cites.
First order constrained optimization in policy space
Yiming Zhang, Quan Vuong, and Keith Ross · 2020
Later among the works it cites.
Motion planning for an unmanned surface vehicle based on topological position maps
Chunhui Zhou, Shangding Gu, Yuanqiao Wen, Zhe Du, Changshi Xiao, Liang Huang, and Man Zhu · 2020
Later among the works it cites.
The review unmanned surface vehicle path planning: Based on multi-modality constraint
Chunhui Zhou, Shangding Gu, Yuanqiao Wen, Zhe Du, Changshi Xiao, Liang Huang, and Man Zhu · 2020
Later among the works it cites.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, and Roberto Martín-Martín · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Later among the works it cites.
Shahin Atakishiyev, Mohammad Salameh, Hengshuai Yao, and Randy Goebel · 2021
Later among the works it cites.
Towards safe, explainable, and regulated autonomous driving
Shahin Atakishiyev, Mohammad Salameh, Hengshuai Yao, and Randy Goebel · 2021
Later among the works it cites.
Conservative safety critics for exploration
Homanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine, Florian Shkurti, and Animesh Garg · 2021
Later among the works it cites.
Safe learning in robotics: From learning-based control to safe reinforcement learning
Lukas Brunke, Melissa Greeff, Adam W Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P Schoellig · 2021
Later among the works it cites.
Mingyu Cai and Cristian-Ioan Vasile · 2021
Later among the works it cites.
State augmented constrained reinforcement learning: Overcoming the limitations of learning with rewards
Miguel Calvo-Fullana, Santiago Paternain, Luiz FO Chamon, and Alejandro Ribeiro · 2021
Later among the works it cites.
A primal-dual approach to constrained markov decision processes
Yi Chen, Jing Dong, and Zhaoran Wang · 2021
Later among the works it cites.
Constrained multiagent markov decision processes: A taxonomy of problems and algorithms
Frits De Nijs, Erwin Walraven, Mathijs De Weerdt, and Matthijs Spaan · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Later among the works it cites.
Safe multi-agent reinforcement learning via shielding
Ingy ElSayed-Aly, Suda Bharadwaj, Christopher Amato, Rüdiger Ehlers, Ufuk Topcu, and Lu Feng · 2021
Later among the works it cites.
Model-based reinforcement learning for infinite-horizon discounted constrained markov decision processes
Aria HasanzadeZonuzy, Dileep Kalathil, and Srinivas Shakkottai · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for discounted mdps
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Unsolved problems in ml safety
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2021
Later among the works it cites.
Verifiably safe exploration for end-to-end reinforcement learning
Nathan Hunt, Nathan Fulton, Sara Magliacane, Trong Nghia Hoang, Subhro Das, and Armando Solar-Lezama · 2021
Later among the works it cites.
Lyapunov-based uncertainty-aware safe reinforcement learning
Ashkan B Jeddi, Nariman L Dehghani, and Abdollah Shafieezadeh · 2021
Later among the works it cites.
A sample-efficient algorithm for episodic finite-horizon mdp with constraints
Krishna C. Kalagarla, Rahul Jain, and Pierluigi Nuzzo · 2021
Later among the works it cites.
Minimizing safety interference for safe and comfortable automated driving with distributional reinforcement learning
Danial Kamran, Tizian Engelgeh, Marvin Busch, Johannes Fischer, and Christoph Stiller · 2021
Later among the works it cites.
Trust region policy optimisation in multi-agent reinforcement learning
Jakub Grudzien Kuba, Ruiqing Chen, Munning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang · 2021
Later among the works it cites.
Settling the variance of multi-agent policy gradients
Jakub Grudzien Kuba, Muning Wen, Yaodong Yang, Linghui Meng, Shangding Gu, Haifeng Zhang, David Henry Mguni, and Jun Wang · 2021
Later among the works it cites.
Cmix: Deep multi-agent reinforcement learning with peak and average constraints
Chenyi Liu, Nan Geng, Vaneet Aggarwal, Tian Lan, Yuan Yang, and Mingwei Xu · 2021
Later among the works it cites.
Learning policies with zero or bounded constraint violation for constrained mdps
Tao Liu, Ruida Zhou, Dileep Kalathil, P. R. Kumar, and Chao Tian · 2021
Later among the works it cites.
Resource allocation method for network slicing using constrained reinforcement learning
Yongshuai Liu, Jiaxin Ding, and Xin Liu · 2021
Later among the works it cites.
Policy learning with constraints in model-free reinforcement learning: A survey
Yongshuai Liu, Avishai Halev, and Xin Liu · 2021
Later among the works it cites.
Decentralized policy gradient descent ascent for safe multi-agent reinforcement learning
Songtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar, and Lior Horesh · 2021
Later among the works it cites.
Model-based constrained reinforcement learning using generalized control barrier function
Haitong Ma, Jianyu Chen, Shengbo Eben, Ziyu Lin, Yang Guan, Yangang Ren, and Sifa Zheng · 2021
Later among the works it cites.
Reinforcement learning for autonomous driving with latent state inference and spatial-temporal relationships
Xiaobai Ma, Jiachen Li, Mykel J Kochenderfer, David Isele, and Kikuo Fujimura · 2021
Later among the works it cites.
Isaac gym: High performance gpu based physics simulation for robot learning
Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, et al · 2021
Later among the works it cites.
Safe reinforcement learning: A control barrier function optimization approach
Zahra Marvi and Bahare Kiumarsi · 2021
Later among the works it cites.
A simple reward-free approach to constrained reinforcement learning
Sobhan Miryoosefi and Chi Jin · 2021
Later among the works it cites.
Robust reinforcement learning: A case study in linear quadratic regulation
Bo Pang and Zhong-Ping Jiang · 2021
Later among the works it cites.
Density constrained reinforcement learning
Zengyi Qin, Yuxiao Chen, and Chuchu Fan · 2021
Later among the works it cites.
Reinforcement learning with quantitative verification for assured multi-agent policies
Joshua Riley, Radu Calinescu, Colin Paterson, Daniel Kudenko, and Alec Banks · 2021
Later among the works it cites.
Model-free safe reinforcement learning for chemical processes using gaussian processes
Thomas Savage, Dongda Zhang, Max Mowbray, and Ehecatl Antonio Del Río Chanona · 2021
Later among the works it cites.
Safe deep reinforcement learning for multi-agent systems with continuous action spaces
Ziyad Sheebaelhamd, Konstantinos Zisis, Athina Nisioti, Dimitris Gkouletsos, Dario Pavllo, and Jonas Kohler · 2021
Later among the works it cites.
Shortest-path constrained reinforcement learning for sparse reward tasks
Sungryull Sohn, Sungtae Lee, Jongwook Choi, Harm van Seijen, Mehdi Fatemi, and Honglak Lee · 2021
Later among the works it cites.
Safe reinforcement learning by imagining the near future
Garrett Thomas, Yuping Luo, and Tengyu Ma · 2021
Later among the works it cites.
Probabilistic robust linear quadratic regulators with gaussian processes
Alexander von Rohr, Matthias Neumann-Brosig, and Sebastian Trimpe · 2021
Later among the works it cites.
Barrier function-based safe reinforcement learning for emergency control of power systems
Thanh Long Vu, Sayak Mukherjee, Renke Huang, and Qiuhua Huang · 2021
Later among the works it cites.
Safe reinforcement learning using advantage-based intervention
Nolan Wagener, Byron Boots, and Ching-An Cheng · 2021
Later among the works it cites.
Formal reachability analysis for multi-agent reinforcement learning systems
Xiaoyan Wang, Jun Peng, Shuqiu Li, and Bing Li · 2021
Later among the works it cites.
A provably-efficient model-free algorithm for constrained markov decision processes
Honghao Wei, Xin Liu, and Lei Ying · 2021
Later among the works it cites.
Uav anti-jamming video transmissions with qoe guarantee: A reinforcement learning-based approach
Liang Xiao, Yuzhen Ding, Jinhao Huang, Sicong Liu, Yuliang Tang, and Huaiyu Dai · 2021
Later among the works it cites.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Later among the works it cites.
Sample complexity of policy gradient finding second-order stationary points
Long Yang, Qian Zheng, and Gang Pan · 2021
Later among the works it cites.
Diverse auto-curriculum is critical for successful real-world multiagent learning systems
Yaodong Yang, Jun Luo, Ying Wen, Oliver Slumbers, Daniel Graves, Haitham Bou Ammar, Jun Wang, and Matthew E Taylor · 2021
Later among the works it cites.
Zhaocong Yuan, Adam W Hall, Siqi Zhou, Lukas Brunke, Melissa Greeff, Jacopo Panerati, and Angela P Schoellig · 2021
Later among the works it cites.
Safe continuous control with constrained model-based policy optimization, 2021
Moritz A. Zanger, Karam Daaboul, and J. Marius Zöllner · 2021
Later among the works it cites.
Sihan Zeng, Thinh T Doan, and Justin Romberg · 2021
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2021
Later among the works it cites.
Dear: Deep reinforcement learning for online advertising impression in recommender systems
Xiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang, Xiaobing Liu, Hui Liu, and Jiliang Tang · 2021
Later among the works it cites.
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2022
Closest in time.
Dynamic stochastic electric vehicle routing with safe reinforcement learning
Rafael Basso, Balázs Kulcsár, Ivan Sanchez-Diaz, and Xiaobo Qu · 2022
Closest in time.
Safety verification of autonomous systems: A multi-fidelity reinforcement learning approach
Jared J Beard and Ali Baheri · 2022
Closest in time.
Trustworthy safety improvement for autonomous driving using reinforcement learning
Zhong Cao, Shaobing Xu, Xinyu Jiao, Huei Peng, and Diange Yang · 2022
Closest in time.
Samba: Safe model-based & active reinforcement learning
Alexander I Cowen-Rivers, Daniel Palenicek, Vincent Moens, Mohammed Amin Abdullah, Aivar Sootla, Jun Wang, and Haitham Bou-Ammar · 2022
Closest in time.
Policy-based reinforcement learning for assortative matching in human behavior modeling
Ou Deng and Qun Jin · 2022
Closest in time.
Run time assured reinforcement learning for safe satellite docking
Kyle Dunlap, Mark Mote, Kai Delsing, and Kerianne L Hobbs · 2022
Closest in time.
Safe reinforcement learning using robust control barrier functions
Yousef Emam, Gennaro Notomista, Paul Glotfelter, Zsolt Kira, and Magnus Egerstedt · 2022
Closest in time.
Resilient reinforcement learning and robust output regulation under denial-of-service attacks
Weinan Gao, Chao Deng, Yi Jiang, and Zhong-Ping Jiang · 2022
Closest in time.
Provably efficient model-free constrained rl with linear function approximation
Arnob Ghosh, Xingyu Zhou, and Ness Shroff · 2022
Closest in time.
Constrained reinforcement learning for vehicle motion planning with topological reachability analysis
Shangding Gu, Guang Chen, Lijun Zhang, Jing Hou, Yingbai Hu, and Alois Knoll · 2022
Closest in time.
Motion planning for an unmanned surface vehicle with wind and current effects
Shangding Gu, Chunhui Zhou, Yuanqiao Wen, Changshi Xiao, and Alois Knoll · 2022
Closest in time.
A cmdp-within-online framework for meta-safe reinforcement learning
Vanshaj Khattar, Yuhao Ding, Bilgehan Sel, Javad Lavaei, and Ming Jin · 2022
Closest in time.
Robot reinforcement learning on the constraint manifold
Puze Liu, Davide Tateo, Haitham Bou Ammar, and Jan Peters · 2022
Closest in time.
Safe exploration in wireless security: A safe reinforcement learning algorithm with hierarchical structure
Xiaozhen Lu, Liang Xiao, Guohang Niu, Xiangyang Ji, and Qian Wang · 2022
Closest in time.
Muzero with self-competition for rate control in vp9 video compression
Amol Mandhane, Anton Zhernov, Maribeth Rauh, Chenjie Gu, Miaosen Wang, Flora Xue, Wendy Shang, Derek Pang, Rene Claus, Ching-Han Chiang, et al · 2022
Closest in time.
Safer: Data-efficient and safe reinforcement learning via skill acquisition
Dylan Slack, Yinlam Chow, Bo Dai, and Nevan Wichers · 2022
Closest in time.
Near-optimal sample complexity bounds for constrained mdps
Sharan Vaswani, Lin Yang, and Csaba Szepesvári · 2022
Closest in time.
Ensuring safety of learning-based motion planners using control barrier functions
Xiao Wang · 2022
Closest in time.
Cup: A conservative update policy algorithm for safe reinforcement learning
Long Yang, Jiaming Ji, Juntao Dai, Yu Zhang, Pengfei Li, and Gang Pan · 2022
Closest in time.
Constrained reinforcement learning via dissipative saddle flow dynamics
Tianqi Zheng, Pengcheng You, and Enrique Mallada · 2022
Closest in time.
Rita: Boost autonomous driving simulators with realistic interactive traffic flow
Zhengbang Zhu, Shenyu Zhang, Yuzheng Zhuang, Yuecheng Liu, Minghuan Liu, Liyuan Mao, Ziqing Gong, Weinan Zhang, Shixiong Kai, Qiang Gu, et al · 2022
Closest in time.
Safe exploration in model-based reinforcement learning using control barrier functions
Max H Cohen and Calin Belta · 2023
Closest in time.
Algorithm for constrained markov decision process with linear convergence
Egor Gladin, Maksim Lavrik-Karmazin, Karina Zainullina, Varvara Rudenko, Alexander Gasnikov, and Martin Takac · 2023
Closest in time.
Safe multi-agent reinforcement learning for multi-robot control
Shangding Gu, Jakub Grudzien Kuba, Yuanpei Chen, Yali Du, Long Yang, Alois Knoll, and Yaodong Yang · 2023
Closest in time.
Reload: Reinforcement learning with optimistic ascent-descent for last-iterate convergence in constrained mdps
Ted Moskovitz, Brendan O’Donoghue, Vivek Veeriah, Sebastian Flennerhag, Satinder Singh, and Tom Zahavy · 2023
Closest in time.
Constrained MDPs and the reward hypothesis
Csaba Szepesvári · 2023
Closest in time.
Model-free safe reinforcement learning through neural barrier certificate
Yujie Yang, Yuxuan Jiang, Yichen Liu, Jianyu Chen, and Shengbo Eben Li · 2023
Closest in time.
Last-iterate convergent policy gradient primal-dual methods for constrained mdps
Dongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, and Alejandro Ribeiro · 2024
Closest in time.
Safe multiagent learning with soft constrained policy optimization in real robot control
Shangding Gu, Dianye Huang, Muning Wen, Guang Chen, and Alois Knoll · 2024
Closest in time.
Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation
Shangding Gu, Bilgehan Sel, Yuhao Ding, Lu Wang, Qingwei Lin, Ming Jin, and Alois Knoll · 2024
Closest in time.
Faster algorithm and sharper analysis for constrained markov decision process
Tianjiao Li, Ziwei Guan, Shaofeng Zou, Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2024
Closest in time.
Provably safe reinforcement learning with step-wise violation constraints
Nuoya Xiong, Yihan Du, and Longbo Huang · 2024
Closest in time.
Constrained reinforcement learning with smoothed log barrier function
Baohe Zhang, Yuan Zhang, Lilli Frison, Thomas Brox, and Joschka Bödecker · 2024
Closest in time.