Fetching the paper…
Reading the bibliography…
In safety-critical applications, autonomous agents may need to learn in an environment where mistakes can be very costly.
Lyapunov-based safe policy optimization for continuous control
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M. (2019) · 1901
Earlier work this paper cites.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O. (2019) · 1901
Earlier work this paper cites.
Batch policy learning under constraints
Le, H. M., Voloshin, C., and Yue, Y. (2019) · 1903
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2019) · 1908
Earlier work this paper cites.
The application of bayesian methods for seeking the extremum
Mockus, J., Tiesis, V., and Zilinskas, A. (1978) · 1978
Earlier work this paper cites.
ALVINN: An autonomous land vehicle in a neural network
Pomerleau, D. (1989) · 1989
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
Kivinen, J. and Warmuth, M. K. (1997) · 1997
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. (2004) · 2004
Earlier work this paper cites.
Behavior transfer for value-function-based reinforcement learning
Taylor, M. E. and Stone, P. (2005) · 2005
Earlier work this paper cites.
A survey of robot learning from demonstration
D.Argall, B., Chernova, S., Veloso, M., and Browning, B. (2009) · 2009
Earlier work this paper cites.
Probabilistic policy reuse for inter-task transfer learning
Fernández, F., García, J., and Veloso, M. (2010) · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. (2010) · 2010
Earlier work this paper cites.
Transfer from multiple mdps
Lazaric, A. and Restelli, M. (2011) · 2011
Earlier work this paper cites.
Transferring task models in reinforcement learning agents
Fachantidis, A., Partalas, I., Tsoumakas, G., and Vlahavas, I. (2013) · 2013
Earlier work this paper cites.
Reachability-based safe learning with gaussian processes
Akametalu, A. K., Fisac, J. F., Gillula, J. H., Kaynama, S., Zeilinger, M. N., and Tomlin, C. J. (2014) · 2014
Earlier work this paper cites.
Approximate policy iteration schemes: a comparison
Scherrer, B. (2014) · 2014
Earlier work this paper cites.
Multi-armed bandits for intelligent tutoring systems
Clement, B., Roy, D., Oudeyer, P.-Y., and Lopes, M. (2015) · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F. (2015) · 2015
Cited alongside, same era.
GPyOpt: A bayesian optimization framework in python
Authors, T. G. (2016) · 2016
Cited alongside, same era.
Safe controller optimization for quadrotors with Gaussian processes
Berkenkamp, F., Schoellig, A. P., and Krause, A. (2016) · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Convex synthesis of randomized policies for controlled markov chains with density safety upper bound constraints
El Chamie, M., Yu, Y., and Açıkmeşe, B. (2016) · 2016
Cited alongside, same era.
Source task creation for curriculum learning
Stable baselines
Hill, A., Raffin, A., Ernestus, M., Gleave, A., Kanervisto, A., Traore, R., Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y. (2018) · 2018
Later among the works it cites.
Learning-based model predictive control for safe exploration
Koller, T., Berkenkamp, F., Turchetta, M., and Krause, A. (2018) · 2018
Later among the works it cites.
An algorithmic perspective on imitation learning
Osa, T., Pajarinen, J., Neumann, G., Bagnell, J. A., Abbeel, P., and Peters, J. (2018) · 2018
Later among the works it cites.
Agile autonomous driving using end-to-end deep imitation learning
Pan, Y., Cheng, C.-A., Saigol, K., Lee, K., Yan, X., Theodorou, E., and Boots, B. (2018) · 2018
Later among the works it cites.
Learning by playing – solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Narvekar, S., Sinapov, J., Leonetti, M., and Stone, P. (2016) · 2016
Cited alongside, same era.
Safe exploration in finite markov decision processes with gaussian processes
Turchetta, M., Berkenkamp, F., and Krause, A. (2016) · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A. P., and Krause, A. (2017) · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Automated curriculum learning for neural networks
Graves, A., Bellemare, M. G., Menick, J., Munos, R., and Kavukcuoglu, K. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Vanschoren, J. (2018) · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N. (2019) · 2019
Later among the works it cites.
Challenges of real-world reinforcement learning
Dulac-Arnold, G., Mankowitz, D., and Hester, T. (2019) · 2019
Later among the works it cites.
Learning to drive in a day
Kendall, A., Hawke, J., Janz, D., Mazur, P., Reda, D., Allen, J.-M., Lam, V.-D., Bewley, A., and Shah, A. (2019) · 2019
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2019) · 2019
Later among the works it cites.
Learning curriculum policies for reinforcement learning
Narvekar, S. and Stone, P. (2019) · 2019
Later among the works it cites.
Teacher algorithms for curriculum learning of deep RL in continuously parameterized environments
Portelas, R., Colas, C., Hofmann, K., and Oudeyer, P.-Y. (2019) · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D. (2019) · 2019
Later among the works it cites.
Safe exploration for interactive machine learning
Turchetta, M., Berkenkamp, F., and Krause, A. (2019) · 2019
Later among the works it cites.
Automatic curriculum learning for deep rl: A short survey
Portelas, R., Colas, C., Weng, L., Hofmann, K., and Oudeyer, P.-Y. (2020) · 2020
Closest in time.
Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks
Thananjeyan, B., Balakrishna, A., Rosolia, U., Li, F., McAllister, R., Gonzalez, J. E., Levine, S., Borrelli, F., and Goldberg, K. (2020) · 2020
Closest in time.