Fetching the paper…
Reading the bibliography…
Advances in Reinforcement Learning (RL) have demonstrated data efficiency and optimal control over large state spaces at the cost of scalable performance.
Robust estimation of a location parameter
Huber, P. J · 1964
Earlier work this paper cites.
Learning Representations by Back-propagating Errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Efficient reinforcement learning through symbiotic evolution
Moriarty, D. E. and Mikkulainen, R · 1996
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Stanley, K. O. and Miikkulainen, R · 2002
Earlier work this paper cites.
Double q-learning
Hasselt, H. V · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N. M. O., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K. R., Saeedi, A., and Tenenbaum, J. B · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., Chen, X., and Chen, X · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., and et. al · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning, 2017
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Cited alongside, same era.
Cem-rl: Combining evolutionary and gradient-based methods for policy search, 2018
Pourchot, A. and Sigaud, O · 2018
Later among the works it cites.
Transfer learning for semg-based hand gesture classification using deep learning in a master- slave architecture
Suri, K. and Gupta, R · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T. P., and Riedmiller, M. A · 2018
Later among the works it cites.
Scalable reinforcement-learning-based neural architecture search for cancer deep learning research
Balaprakash, P., Egele, R., Salim, M., Wild, S., Vishwanath, V., Xia, F., Brettin, T., and Stevens, R · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation, 2017
Wu, Y., Mansimov, E., Liao, S., Grosse, R., and Ba, J · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Evolved policy gradients
Houthooft, R., Chen, Y., Isola, P., Stadie, B., Wolski, F., Ho, J., and Abbeel, P · 2018
Cited alongside, same era.
Ppo-cma: Proximal policy optimization with covariance matrix adaptation
Hämäläinen, P., Babadi, A., Ma, X., and Lehtinen, J · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S · 2018
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination, 2019
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Later among the works it cites.
Efficacy of modern neuro-evolutionary strategies for continuous control optimization, 2019
Pagliuca, P., Milano, N., and Nolfi, S · 2019
Later among the works it cites.
Seed rl: Scalable and efficient deep-rl with accelerated central inference
Espeholt, L., Marinier, R., Stanczyk, P., Wang, K., and Michalski, M · 2020
Closest in time.
Maxmin q-learning: Controlling the estimation bias of q-learning
Lan, Q., Pan, Y., Fyshe, A., and White, M · 2020
Closest in time.
Mm-ktd: Multiple model kalman temporal differences for reinforcement learning, 2020
Malekzadeh, P., Salimibeni, M., Mohammadi, A., Assa, A., and Plataniotis, K. N · 2020
Closest in time.
Backpropamine: training self-modifying neural networks with differentiable neuromodulated plasticity, 2020
Miconi, T., Rawal, A., Clune, J., and Stanley, K. O · 2020
Closest in time.
Multi-level fitness critics for cooperative coevolution
Rockefeller, G., Khadka, S., and Tumer, K · 2020
Closest in time.
Curl: Contrastive unsupervised representations for reinforcement learning
Srinivas, A., Laskin, M., and Abbeel, P · 2020
Closest in time.