Fetching the paper…
Reading the bibliography…
Reinforcement learning provides an appealing framework for robotic control due to its ability to learn expressive policies purely through real-world interaction.
Risk-sensitive linear/quadratic/gaussian control
P. Whittle · 1981
Earlier work this paper cites.
An image synthesizer
Ken Perlin · 1985
Earlier work this paper cites.
Constrained Markov Decision Processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Coherent measures of risk
Philippe Artzner, Freddy Delbaen, Jean-Marc Eber, and David Heath · 1999
Earlier work this paper cites.
Some remarks on the value-at-risk and the conditional value-at-risk
Georg Ch Pflug · 2000
Earlier work this paper cites.
A comparison of var and cvar constraints on portfolio selection with the mean-variance model
Gordon J Alexander and Alexandre M Baptista · 2004
Earlier work this paper cites.
Time consistent dynamic risk measures
Kang Boda and Jerzy A Filar · 2006
Earlier work this paper cites.
Risk, var, cvar and their associated portfolio optimizations when asset returns have a multivariate student t distribution, 2011
William T. Shaw · 2011
Earlier work this paper cites.
Parametric return density estimation for reinforcement learning
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka · 2012
Earlier work this paper cites.
Scaling up robust mdps by reinforcement learning
Aviv Tamar, Huan Xu, and Shie Mannor · 2013
Earlier work this paper cites.
Algorithms for cvar optimization in mdps
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Earlier work this paper cites.
Policy gradients for cvar-constrained mdps
LA Prashanth · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva Tb, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Cited alongside, same era.
Anatomy of a rollover, 2018
Marina Bruce · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc Bellemare, and Rémi Munos · 2018
Cited alongside, same era.
Opening new dimensions: Vehicle motion planning and control using brakes while drifting
Tushar Goel, Jonathan Y. Goh, and J. Christian Gerdes · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
F1tenth: An open-source evaluation environment for continuous control and reinforcement learning
Matthew O’Kelly, Hongrui Zheng, Dhruv Karthik, and Rahul Mangharam · 2020
Later among the works it cites.
Learning to be safe: Deep rl with a safety critic
Krishnan Srinivasan, Benjamin Eysenbach, Sehoon Ha, Jie Tan, and Chelsea Finn · 2020
Later among the works it cites.
Safe reinforcement learning for autonomous vehicles through parallel constrained policy optimization
Lu Wen, Jingliang Duan, Shengbo Eben Li, Shaobing Xu, and Huei Peng · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Model predictive control of vehicle roll-over with experimental verification
Milad Jalali, Ehsan Hashemi, Amir Khajepour, Shih-ken Chen, and Bakhtiar Litkouhi · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Information-theoretic model predictive control: Theory and applications to autonomous driving
Grady Williams, Paul Drews, Brian Goldfain, James M. Rehg, and Evangelos A. Theodorou · 2018
Cited alongside, same era.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
Richard Cheng, Gábor Orosz, Richard M Murray, and Joel W Burdick · 2019
Cited alongside, same era.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Wenshuai Zhao, Jorge Peña Queralta, and Tomi Westerlund · 2020
Later among the works it cites.
Risk-sensitive safety analysis using conditional value-at-risk
Margaret P Chapman, Riccardo Bonalli, Kevin M Smith, Insoon Yang, Marco Pavone, and Claire J Tomlin · 2021
Later among the works it cites.
Randomized ensembled double q-learning: Learning fast without a model
Xinyue Chen, Che Wang, Zijian Zhou, and Keith Ross · 2021
Later among the works it cites.
Uncertainty quantification and deep ensembles
Rahul Rahaman et al · 2021
Later among the works it cites.
Autonomous reinforcement learning: Formalism and benchmarking
Archit Sharma, Kelvin Xu, Nikhil Sardana, Abhishek Gupta, Karol Hausman, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
Constrained policy optimization via bayesian world models
Yarden As, Ilnura Usmanova, Sebastian Curi, and Andreas Krause · 2022
Later among the works it cites.
Deep drifting: Autonomous drifting of arbitrary trajectories using deep reinforcement learning
Fabian Domberg, Carlos Castelar Wembers, Hiren Patel, and Georg Schildbach · 2022
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
Pierluca D’Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon, Marc G Bellemare, and Aaron Courville · 2022
Later among the works it cites.
Ensemble deep learning: A review
Mudasir A Ganaie, Minghui Hu, AK Malik, M Tanveer, and PN Suganthan · 2022
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
Max Schwarzer, Johan Samir Obando Ceron, Aaron Courville, Marc G Bellemare, Rishabh Agarwal, and Pablo Samuel Castro · 2023
Later among the works it cites.
Solving stabilize-avoid optimal control via epigraph form and deep reinforcement learning, 2023
Oswin So and Chuchu Fan · 2023
Later among the works it cites.
Fastrlap: A system for learning high-speed driving via deep rl and autonomous practicing, 2023
Kyle Stachowicz, Dhruv Shah, Arjun Bhorkar, Ilya Kostrikov, and Sergey Levine · 2023
Later among the works it cites.
Ensemble-based out-of-distribution detection
Donghun Yang, Kien Mai Ngoc, Iksoo Shin, Kyong-Ha Lee, and Myunggwon Hwang · 2079
Closest in time.