Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms, though successful, tend to over-fit to training environments hampering their application to the real-world.
Compatible natural gradient policy search
Pajarinen, J., Thai, H. L., Akrour, R., Peters, J., and Neumann, G. (2019) · 1902
Earlier work this paper cites.
Beyond confidence regions: Tight bayesian ambiguity sets for robust mdps
Petrik, M. and Russell, R. H. (2019) · 1902
Earlier work this paper cites.
Investigating generalisation in continuous deep reinforcement learning
Zhao, C., Sigaud, O., Stulp, F., and Hospedales, T. M. (2019) · 1902
Earlier work this paper cites.
Lecarpentier, E. and Rachelson, E. (2019) · 1904
Earlier work this paper cites.
Convergence conditions for ascent methods
Wolfe, P. (1969) · 1969
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Robust control and model uncertainty
Sargent, T. and Hansen, L. (2001) · 2001
Earlier work this paper cites.
Gaussian processes in machine learning
Rasmussen, C. E. (2003) · 2003
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N. (2005) · 2005
Earlier work this paper cites.
Robust reinforcement learning
Morimoto, J. and Doya, K. (2005) · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L. (2005) · 2005
Earlier work this paper cites.
Wasserstein geometry of gaussian measures
Takatsu, A. (2008) · 2008
Earlier work this paper cites.
Reinforcement Learning and Dynamic Programming Using Function Approximators
Busoniu, L., Babuska, R., Schutter, B. D., and Ernst, D. (2010) · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E. (2011) · 2011
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Nesterov, Y. (2011) · 2011
Cited alongside, same era.
Reinforcement Learning in Robotics: A Survey
Kober, J. and Peters, J. (2012) · 2012
Cited alongside, same era.
A survey on policy search for robotics
Deisenroth, M., Peters, J., and Neumann, G. (2013) · 2013
Cited alongside, same era.
Feedback control theory
Doyle, J. C., Francis, B. A., and Tannenbaum, A. R. (2013) · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A. (2013) · 2013
Cited alongside, same era.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B. (2013) · 2013
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
A convex optimization approach to distributionally robust markov decision processes with wasserstein distance
Yang, I. (2017) · 2017
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Asadi, K., Misra, D., and Littman, M. (2018) · 2018
Later among the works it cites.
Neural ordinary differential equations
Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online multi-task learning for policy gradient methods
Bou-Ammar, H., Eaton, E., Ruvolo, P., and Taylor, M. E. (2014) · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Risk-sensitive and robust decision-making: a cvar optimization approach
Chow, Y., Tamar, A., Mannor, S., and Pavone, M. (2015) · 2015
Cited alongside, same era.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Stochastic gradient methods for distributionally robust optimization with f-divergences
Namkoong, H. and Duchi, J. C. (2016) · 2016
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, X. B., Andrychowicz, M., Zaremba, W., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
A machine learning perspective on personalized medicine: An automized, comprehensive knowledge base with ontology for pattern recognition
Emmert-Streib, F. and Dehmer, M. (2018) · 2018
Later among the works it cites.
Reinforcement learning in financial markets - a survey
Fischer, T. G. (2018) · 2018
Later among the works it cites.
Learning-based model predictive control for safe exploration and reinforcement learning
Koller, T., Berkenkamp, F., Turchetta, M., and Krause, A. (2018) · 2018
Later among the works it cites.
Assessing generalization in deep reinforcement learning
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D. (2018) · 2018
Later among the works it cites.
Policy-conditioned uncertainty sets for robust markov decision processes
Tirinzoni, A., Petrik, M., Chen, X., and Ziebart, B. (2018) · 2018
Later among the works it cites.
Deep lagrangian networks for end-to-end learning of energy-based control for under-actuated systems
Lutter, M. and Peters, J. (2019) · 2019
Closest in time.
Deep lagrangian networks: Using physics as model prior for deep learning
Lutter, M., Ritter, C., and Peters, J. (2019) · 2019
Closest in time.
Action robust reinforcement learning and applications in continuous control
Tessler, C., Efroni, Y., and Mannor, S. (2019) · 2019
Closest in time.