Fetching the paper…
Reading the bibliography…
In computer vision and natural language processing, innovations in model architecture that increase model capacity have reliably translated into gains in performance.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 1912
Earlier work this paper cites.
Praktische verfahren der gleichungsauflösung
R. Mises and H. Pollaczek-Geiringer · 1929
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Learning from delayed rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2004
Earlier work this paper cites.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2004
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Generative adversarial networks
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning. corr abs/1509.06461 (2015)
H. van Hasselt, A. Guez, and D. Silver · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Neural photo editing with introspective adversarial networks
A. Brock, T. Lim, J. M. Ritchie, and N. Weston · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas · 2016
Earlier work this paper cites.
Wasserstein generative adversarial networks
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. Courville · 2017
Earlier work this paper cites.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
On the role of planning in model-based deep reinforcement learning
J. B. Hamrick, A. L. Friesen, F. Behbahani, A. Guez, F. Viola, S. Witherspoon, T. Anthony, L. Buesing, P. Veličković, and T. Weber · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2020
Later among the works it cites.
High-throughput synchronous deep rl
I.-J. Liu, R. Yeh, and A. Schwing · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
E. Parisotto, F. Song, J. Rae, R. Pascanu, C. Gulcehre, S. Jayakumar, M. Jaderberg, R. L. Kaufman, A. Clark, S. Noury, et al · 2020
Later among the works it cites.
Sample factory: Egocentric 3d control from pixels at 100000 fps with asynchronous reinforcement learning
A. Petrenko, Z. Huang, T. Kumar, G. Sukhatme, and V. Koltun · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generalization and regularization in dqn
J. Farebrother, M. C. Machado, and M. Bowling · 2018
Cited alongside, same era.
Generalizable adversarial training via spectral normalization
F. Farnia, J. M. Zhang, and D. Tse · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Quantifying generalization in reinforcement learning
K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman · 2019
Cited alongside, same era.
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2020
Later among the works it cites.
D2rl: Deep dense architectures in reinforcement learning
S. Sinha, H. Bharadhwaj, A. Srinivas, and A. Garg · 2020
Later among the works it cites.
S. A. Sontakke, A. Mehrjou, L. Itti, and B. Schölkopf · 2020
Later among the works it cites.
dm control: Software and tasks for continuous control, 2020
Y. Tassa, S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, and N. Heess · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Self-supervised visual reinforcement learning with object-centric representations
A. Zadaianchuk, M. Seitzer, and G. Martius · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
R. Agarwal, M. C. Machado, P. S. Castro, and M. G. Bellemare · 2021
Closest in time.
Low-precision reinforcement learning
J. Bjorck, X. Chen, C. De Sa, C. P. Gomes, and K. Q. Weinberger · 2021
Closest in time.
Offline model-based optimization via normalized maximum likelihood estimation
J. Fu and S. Levine · 2021
Closest in time.
Spectral normalisation for deep reinforcement learning: an optimisation perspective
F. Gogianu, T. Berariu, M. Rosca, C. Clopath, L. Busoniu, and R. Pascanu · 2021
Closest in time.
Regularisation of neural networks by enforcing lipschitz continuity
H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree · 2021
Closest in time.
Deep reinforcement learning for autonomous driving: A survey
B. R. Kiran, I. Sobh, V. Talpaert, P. Mannion, A. A. Al Sallab, S. Yogamani, and P. Pérez · 2021
Closest in time.
Representation balancing offline model-based reinforcement learning
B.-J. Lee, J. Lee, and K.-E. Kim · 2021
Closest in time.
Return-based contrastive representation learning for reinforcement learning
G. Liu, C. Zhang, L. Zhao, T. Qin, J. Zhu, J. Li, N. Yu, and T.-Y. Liu · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?
S. Narang, H. W. Chung, Y. Tay, W. Fedus, T. Fevry, M. Matena, K. Malkan, N. Fiedel, N. Shazeer, Z. Lan, et al · 2021
Closest in time.
Efficient transformers in reinforcement learning using actor-learner distillation
E. Parisotto and R. Salakhutdinov · 2021
Closest in time.
Mastering visual continuous control: Improved data-augmented reinforcement learning
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Closest in time.