Fetching the paper…
Reading the bibliography…
Scaling issues are mundane yet irritating for practitioners of reinforcement learning.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Learning values across many orders of magnitude
van Hasselt, H. P., Guez, A., Hessel, M., Mnih, V., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Cited alongside, same era.
Learning from demonstrations for real world reinforcement learning
Hester, T., Vecerík, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Sendonaris, A., Dulac-Arnold, G., Osband, I., Agapiou, J. P., Leibo, J. Z., and Gruslys, A · 2017
Cited alongside, same era.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al · 2017
Cited alongside, same era.
Natural value approximators: Learning when to trust past estimates
Xu, Z., Modayil, J., van Hasselt, H. P., Barreto, A., Silver, D., and Schaul, T · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Curious: intrinsically motivated modular multi-goal reinforcement learning
Colas, C., Fournier, P., Chetouani, M., Sigaud, O., and Oudeyer, P.-Y · 2019
Later among the works it cites.
Learning anytime predictions in neural networks via adaptive loss balancing
Hu, H., Dey, D., Hebert, M., and Bagnell, J. A · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., and Munos, R · 2019
Later among the works it cites.
Adaptive auxiliary task weighting for reinforcement learning
Lin, X., Baweja, H. S., Kantor, G., and Held, D · 2019
Later among the works it cites.
Adapting behaviour for learning progress, 2019
Schaul, T., Borsa, D., Ding, D., Szepesvari, D., Ostrovski, G., Dabney, W., and Osindero, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Chen, Z., Badrinarayanan, V., Lee, C.-Y., and Rabinovich, A · 2018
Cited alongside, same era.
IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Cited alongside, same era.
Generalization and regularization in dqn
Farebrother, J., Machado, M. C., and Bowling, M · 2018
Cited alongside, same era.
Multi-task deep reinforcement learning with popart
Hessel, M., Soyer, H., Espeholt, L., Czarnecki, W., Schmitt, S., and van Hasselt, H · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Kendall, A., Gal, Y., and Cipolla, R · 2018
Cited alongside, same era.
Observe and look further: Achieving consistent performance on atari
Pohlen, T., Piot, B., Hester, T., Gheshlaghi, M., Horgan, D., Budden, D., Barth-Maron, G., Van Hasselt, H., Quan, J., Vecerik, M., Hessel, M., Munos, R., and Pietquin, O · 2018
Cited alongside, same era.
Learning by playing solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Cited alongside, same era.
van Hasselt, H., Quan, J., Hessel, M., Xu, Z., Borsa, D., and Barreto, A · 2019
Later among the works it cites.
A self-tuning actor-critic algorithm, 2020
Zahavy, T., Xu, Z., Veeriah, V., Hessel, M., Oh, J., van Hasselt, H., Silver, D., and Singh, S · 2019
Later among the works it cites.
Reverb: An efficient data storage and transport system for ml research, 2020
Cassirer, A., Barth-Maron, G., Sottiaux, T., Kroiss, M., and Brevdo, E · 2020
Later among the works it cites.
Temporally-extended ϵ \epsilon -greedy exploration, 2020
Dabney, W., Ostrovski, G., and Barreto, A · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Hennigan, T., Cai, T., Norman, T., and Babuschkin, I · 2020
Later among the works it cites.
Optax: Composable gradient transformation and optimisation, in JAX!, 2020
Hessel, M., Budden, D., Viola, F., Rosca, M., Sezener, E., and Hennigan, T · 2020
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Stooke, A., Lee, K., Abbeel, P., and Laskin, M · 2020
Later among the works it cites.
Adapting to reward progressivity via spectral reinforcement learning
Dann, M. and Thangarajah, J · 2021
Closest in time.