Fetching the paper…
Reading the bibliography…
Increasing the scale of reinforcement learning experiments has allowed researchers to achieve unprecedented results in both training sophisticated agents for video games, and in sim-to-real transfer for robotics.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
MapReduce for parallel reinforcement learning
Li, Y. and Schuurmans, D · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Ré, C., Wright, S. J., and Niu, F · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Q( λ \lambda ) with off-policy corrections
Harutyunyan, A., Bellemare, M. G., Stepleton, T., and Munos, R · 2016
Earlier work this paper cites.
Vizdoom: A Doom-based AI research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaskowski, W · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Reinforcement learning through asynchronous advantage actor-critic on a GPU
Babaeizadeh, M., Frosio, I., Tyree, S., Clemons, J., and Kautz, J · 2017
Earlier work this paper cites.
OpenAI baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P · 2017
Earlier work this paper cites.
Learning to act by predicting the future
Dosovitskiy, A. and Koltun, V · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Emergent complexity via multi-agent competition
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I · 2018
Cited alongside, same era.
IMPALA: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D · 2018
Cited alongside, same era.
RLlib: Abstractions for distributed reinforcement learning
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2019
Later among the works it cites.
TorchBeast: A PyTorch platform for distributed RL
Küttler, H., Nardelli, N., Lavril, T., Selvatici, M., Sivakumar, V., Rocktäschel, T., and Grefenstette, E · 2019
Later among the works it cites.
Sim-to-(multi)-real: Transfer of low-level robust control policies to multiple quadrotors
Molchanov, A., Chen, T., Hönig, W., Preiss, J. A., Ayanian, N., and Sukhatme, G. S · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Goldberg, K., Gonzalez, J., Jordan, M. I., and Stoica, I · 2018
Cited alongside, same era.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., et al · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging AI applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., and Stoica, I · 2018
Cited alongside, same era.
Driving policy transfer via modularity and abstraction
Müller, M., Dosovitskiy, A., Ghanem, B., and Koltun, V · 2018
Cited alongside, same era.
IMPALA: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Deep reinforcement learning on a budget: 3D control and reasoning without a supercomputer
Beeching, E., Wolf, C., Dibangoye, J., and Simonin, O · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Cited alongside, same era.
Schmitt, S., Hessel, M., and Simonyan, K · 2019
Later among the works it cites.
rlpyt: A research code base for deep reinforcement learning in PyTorch
Stooke, A. and Abbeel, P · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Vizdoom competitions: Playing Doom from pixels
Wydmuch, M., Kempka, M., and Jaskowski, W · 2019
Later among the works it cites.
Does computer vision matter for action?
Zhou, B., Krähenbühl, P., and Koltun, V · 2019
Later among the works it cites.
SEED RL: Scalable and efficient deep-rl with accelerated central inference
Espeholt, L., Marinier, R., Stanczyk, P., Wang, K., and Michalski, M · 2019
Later among the works it cites.
Multi-task deep reinforcement learning with popart
Hessel, M., Soyer, H., Espeholt, L., Czarnecki, W., Schmitt, S., and van Hasselt, H · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, M., Baker, B., Chociej, M., Józefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W · 2020
Closest in time.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., and Mordatch, I · 2020
Closest in time.
DD-PPO: Learning near-perfect PointGoal navigators from 2.5 billion frames
Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., and Batra, D · 2020
Closest in time.