Fetching the paper…
Reading the bibliography…
Data efficiency is a key challenge for deep reinforcement learning.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J · 1999
Earlier work this paper cites.
Scaling-up knowledge for a cognizant robot
Degris, T. and Modayil, J · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Best Practices for Fine-Tuning Visual Classifiers to New Domains , pp. 435–442
Chu, B., Madhavan, V., Beijbom, O., Hoffman, J., and Darrell, T · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Tarvainen, A. and Valpola, H · 2017
Earlier work this paper cites.
Human learning in atari
Tsividis, P., Pouncy, T., Xu, J. L., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S · 2017
Earlier work this paper cites.
Large-scale study of curiosity-driven learning, 2018
Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., and Efros, A. A · 2018
Earlier work this paper cites.
Dopamine: A Research Framework for Deep Reinforcement Learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Earlier work this paper cites.
Investigating human priors for playing video games
Dubey, R., Agrawal, P., Pathak, D., Griffiths, T., and Efros, A · 2018
Earlier work this paper cites.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Earlier work this paper cites.
Neural predictive belief representations
Guo, Z. D., Azar, M. G., Piot, B., Pires, B. A., and Munos, R · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Earlier work this paper cites.
State representation learning for control: An overview
Lesort, T., Díaz-Rodríguez, N., Goudou, J.-F., and Filliat, D · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A. C · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Embracing change: Continual learning in deep neural networks
Hadsell, R., Rao, D., Rusu, A. A., and Pascanu, R · 2020
Later among the works it cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Warde-Farley, D., de Wiele, T. V., and Mnih, V · 2020
Later among the works it cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del R’ıo, J. F., Wiebe, M., Peterson, P., G’erard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, C., Vinyals, O., Munos, R., and Bengio, S · 2018
Cited alongside, same era.
Unsupervised state representation learning in atari
Anand, A., Racah, E., Ozair, S., Bengio, Y., Côté, M.-A., and Hjelm, R. D · 2019
Cited alongside, same era.
Learning representations by maximizing mutual information across views
Bachman, P., Hjelm, R. D., and Buchwalter, W · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G · 2019
Cited alongside, same era.
Data-efficient image recognition with contrastive predictive coding
Hénaff, O. J., Srinivas, A., De Fauw, J., Razavi, A., Doersch, C., Eslami, S., and Oord, A. v. d · 2019
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y · 2019
Cited alongside, same era.
Later among the works it cites.
Do recent advancements in model-based deep reinforcement learning really improve data efficiency?, 2020
Kielak, K. P · 2020
Later among the works it cites.
Rethinking the hyperparameters for fine-tuning
Li, H., Chaudhari, P., Yang, H., Lam, M., Ravichandran, A., Bhotika, R., and Soatto, S · 2020
Later among the works it cites.
Deep reinforcement and infomax learning
Mazoure, B., Combes, R. T. d., Doan, T., Bachman, P., and Hjelm, R. D · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D · 2020
Later among the works it cites.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Sohn, K., Berthelot, D., Li, C.-L., Zhang, Z., Carlini, N., Cubuk, E. D., Kurakin, A., Zhang, H., and Raffel, C · 2020
Later among the works it cites.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere, 2020
Wang, T. and Isola, P · 2020
Later among the works it cites.
A framework for efficient robotic manipulation, 2020
Zhan, A., Zhao, P., Pinto, L., Abbeel, P., and Laskin, M · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice, 2021
Anonymous · 2021
Closest in time.
Coverage as a principle for discovering transferable behavior in reinforcement learning, 2021
Campos, V., Sprechmann, P., Hansen, S. S., Barreto, A., Blundell, C., Vitvitskyi, A., Kapturowski, S., and Badia, A. P · 2021
Closest in time.
The value-improvement path: Towards better representations for reinforcement learning
Dabney, W., Barreto, A., Rowland, M., Dadashi, R., Quan, J., Bellemare, M. G., and Silver, D · 2021
Closest in time.
Scaling laws for transfer, 2021
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Closest in time.
Muesli: Combining improvements in policy optimization
Hessel, M., Danihelka, I., Viola, F., Guez, A., Schmitt, S., Sifre, L., Weber, T., Silver, D., and van Hasselt, H · 2021
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2021
Closest in time.
Unsupervised active pre-training for reinforcement learning, 2021
Liu, H. and Abbeel, P · 2021
Closest in time.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2021
Closest in time.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2021
Closest in time.
Decoupling representation learning from reinforcement learning, 2021
Stooke, A., Lee, K., Abbeel, P., and Laskin, M · 2021
Closest in time.