Fetching the paper…
Reading the bibliography…
We study the link between generalization and interference in temporal-difference (TD) learning.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S · 1996
Earlier work this paper cites.
Sparse spatial autoregressions
Pace, R. K. and Barry, R · 1997
Earlier work this paper cites.
Locally weighted projection regression: An o(n) algorithm for incremental real time learning in high dimensional space
Vijayakumar, S. and Schaal, S · 2000
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Whiteson, S., Tanner, B., Taylor, M. E., and Stone, P · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Statistical power analysis for the behavioral sciences
Cohen, J · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Maas, A. L., Hannun, A. Y., and Ng, A. Y · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: a method for stochastic optimization (2014)
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning, 2015
van Hasselt, H., Guez, A., and Silver, D · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Later among the works it cites.
On the importance of single directions for generalization
Morcos, A. S., Barrett, D. G., Rabinowitz, N. C., and Botvinick, M · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Later among the works it cites.
Assessing generalization in deep reinforcement learning
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D · 2018
Later among the works it cites.
Learning to learn without forgetting by maximizing transfer and minimizing interference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning, 2017
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M · 2017
Cited alongside, same era.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2018
Cited alongside, same era.
Generalization and regularization in dqn
Farebrother, J., Machado, M. C., and Bowling, M · 2018
Cited alongside, same era.
Riemer, M., Cases, I., Ajemian, R., Liu, M., Rish, I., Tu, Y., and Tesauro, G · 2018
Later among the works it cites.
Lipschitz regularity of deep neural networks: analysis and efficient estimation, 2018
Scaman, K. and Virmaux, A · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Temporal regularization for markov decision process
Thodoroff, P., Durand, A., Pineau, J., and Precup, D · 2018
Later among the works it cites.
Towards characterizing divergence in deep q-learning
Achiam, J., Knight, E., and Abbeel, P · 2019
Later among the works it cites.
Striving for simplicity in off-policy deep reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2019
Later among the works it cites.
Stiffness: A new perspective on generalization in neural networks
Fort, S., Nowak, P. K., and Narayanan, S · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., and Munos, R · 2019
Later among the works it cites.
Generalization guarantees for neural networks via harnessing the low-rank structure of the jacobian
Oymak, S., Fabian, Z., Li, M., and Soltanolkotabi, M · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
The machine learning reproducibility checklist v1.2
Pineau, J · 2019
Later among the works it cites.
Ray interference: a source of plateaus in deep reinforcement learning
Schaul, T., Borsa, D., Modayil, J., and Pascanu, R · 2019
Later among the works it cites.
Back{pack}: Packing more into backprop
Dangel, F., Kunstner, F., and Hennig, P · 2020
Closest in time.