Fetching the paper…
Reading the bibliography…
Reinforcement learning agents must generalize beyond their training experience.
J. Schmidhuber, “Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook,” Ph.D. dissertation, Technische Universität München, 1987
1987
Earlier work this paper cites.
J. Schmidhuber, “Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments,” Institut für Informatik, Technische Universität München. Technical Report FKI-126 , vol. 90, 1990
1990
Earlier work this paper cites.
——, “Curious model-building control systems,” in Proceedings of the International Joint Conference on Neural Networks, Singapore , vol. 2. IEEE press, 1991, pp. 1458–1463
1991
Earlier work this paper cites.
——, “Learning to control fast-weight memories: An alternative to recurrent nets,” Institut für Informatik, Technische Universität München, Tech. Rep. FKI-147-91, March 1991
1991
Earlier work this paper cites.
——, “Steps towards “self-referential” learning,” Dept. of Comp. Sci., University of Colorado at Boulder, Tech. Rep. CU-CS-627-92, November 1992
1992
Earlier work this paper cites.
——, “Learning to control fast-weight memories: An alternative to dynamic recurrent networks,” Neural Computation , vol. 4, no. 1, pp. 131–139, 1992
1992
Earlier work this paper cites.
——, “An introspective network that can learn to run its own weight change algorithm,” in Proc. IEE Int. Conf. on Artificial Neural Networks , Brighton, UK, May 1993, pp. 191–195
1993
Earlier work this paper cites.
——, “A self-referential weight matrix,” in Proc. Int. Conf. on Artificial Neural Networks (ICANN) , Amsterdam, Netherlands, Sep. 1993, pp. 446–451
1993
Earlier work this paper cites.
——, “A neural network that embeds its own meta-levels,” in Proc. IEEE Int. Conf. on Neural Networks (ICNN) , San Francisco, CA, USA, Mar. 1993
1993
Earlier work this paper cites.
——, “Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets,” in International Conference on Artificial Neural Networks (ICANN) , Amsterdam, Netherlands, Sep. 1993, pp. 460–463
1993
Earlier work this paper cites.
C. Von Der Malsburg, “The correlation theory of brain function,” in Models of neural networks . Springer, 1994, pp. 95–119
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. L. Roskies, “The binding problem,” Neuron , vol. 24, no. 1, pp. 7–9, 1999
1999
Earlier work this paper cites.
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner, “Torcs, the open racing car simulator,” Software available at http://torcs. sourceforge. net , vol. 4, no. 6, p. 2, 2000
2000
Earlier work this paper cites.
D. Wierstra, A. Foerster, J. Peters, and J. Schmidhuber, “Solving deep memory pomdps with recurrent policy gradients,” in International conference on artificial neural networks . Springer, 2007, pp. 697–706
2007
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” Journal of Artificial Intelligence Research , vol. 47, pp. 253–279, jun 2013
2013
Earlier work this paper cites.
J. Koutník, G. Cuccu, J. Schmidhuber, and F. Gomez, “Evolving large-scale neural networks for vision-based reinforcement learning,” in Proceedings of the 15th annual conference on Genetic and evolutionary computation , 2013, pp. 1061–1068
2013
Earlier work this paper cites.
——, “Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem,” Frontiers in psychology , vol. 4, p. 313, 2013
2013
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski, “Vizdoom: A doom-based ai research platform for visual reinforcement learning,” in 2016 IEEE conference on computational intelligence and games (CIG) . IEEE, 2016, pp. 1–8
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell, “The malmo platform for artificial intelligence experimentation.” in IJCAI . Citeseer, 2016, pp. 4246–4247
2016
Earlier work this paper cites.
K. Greff, A. Rasmus, M. Berglund, T. Hao, H. Valpola, and J. Schmidhuber, “Tagger: Deep unsupervised perceptual grouping,” in Advances in Neural Information Processing Systems , 2016, pp. 4484–4492
2016
Earlier work this paper cites.
S. A. Eslami, N. Heess, T. Weber, Y. Tassa, D. Szepesvari, G. E. Hinton et al. , “Attend, infer, repeat: Fast scene understanding with generative models,” in Advances in Neural Information Processing Systems , 2016, pp. 3225–3233
2016
Cited alongside, same era.
2017
Cited alongside, same era.
K. Greff, S. van Steenkiste, and J. Schmidhuber, “Neural expectation maximization,” in Advances in Neural Information Processing Systems , 2017, pp. 6691–6701
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
R. Veerapaneni, J. D. Co-Reyes, M. Chang, M. Janner, C. Finn, J. Wu, J. Tenenbaum, and S. Levine, “Entity abstraction in visual model-based reinforcement learning,” in Conference on Robot Learning . PMLR, 2020, pp. 1439–1456
2020
Later among the works it cites.
T. Kipf, E. van der Pol, and M. Welling, “Contrastive learning of structured world models,” in International Conference on Learning Representations , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, 2018
2018
Cited alongside, same era.
S. van Steenkiste, M. Chang, K. Greff, and J. Schmidhuber, “Relational neural expectation maximization: Unsupervised discovery of objects and their interactions,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
A. R. Kosiorek, H. Kim, I. Posner, and Y. W. Teh, “Sequential attend, infer, repeat: Generative modelling of moving objects,” in Proceedings of the 32Nd International Conference on Neural Information Processing Systems , ser. NIPS’18. USA: Curran Associates Inc., 2018, pp. 8615–8625
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision . Springer, 2020, pp. 213–229
2020
Later among the works it cites.
Y. Tang, D. Nguyen, and D. Ha, “Neuroevolution of self-interpretable agents,” in Proceedings of the 2020 Genetic and Evolutionary Computation Conference , 2020, pp. 414–424
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Milani, N. Topin, B. Houghton, W. H. Guss, S. P. Mohanty, K. Nakata, O. Vinyals, and N. S. Kuno, “Retrospective analysis of the 2019 minerl competition on sample efficient reinforcement learning,” in NeurIPS 2019 Competition and Demonstration Track . PMLR, 2020, pp. 203–214
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Later among the works it cites.
D. Hafner, “Benchmarking the spectrum of agent capabilities,” arXiv preprint arXiv:2109.06780 , 2021
2021
Later among the works it cites.
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-baselines3: Reliable reinforcement learning implementations,” Journal of Machine Learning Research , vol. 22, no. 268, pp. 1–8, 2021. [Online]. Available: http://jmlr.org/papers/v22/20-1364.html
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Tang and D. Ha, “The sensory neuron as a transformer: Permutation-invariant neural networks for reinforcement learning,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 4651–4664
2021
Later among the works it cites.
2021
Later among the works it cites.
I. Schlag, K. Irie, and J. Schmidhuber, “Linear transformers are secretly fast weight programmers,” in International Conference on Machine Learning . PMLR, 2021, pp. 9355–9366
2021
Later among the works it cites.
K. Irie, I. Schlag, R. Csordás, and J. Schmidhuber, “Going beyond linear transformers with recurrent fast weight programmers,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
N. Hansen and X. Wang, “Generalization in reinforcement learning by soft data augmentation,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 13 611–13 617
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Creswell, R. Kabra, C. Burgess, and M. Shanahan, “Unsupervised object-based transition models for 3d partially observable environments,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman, “Leveraging procedural generation to benchmark reinforcement learning,” in International conference on machine learning . PMLR, 2020, pp. 2048–2056
2056
Closest in time.