Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) agents make decisions using nothing but observations from the environment, and consequently, heavily rely on the representations of those observations.
Modular learning in neural networks
Dana H. Ballard · 1987
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard E. Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne E. Hubbard, and Lawrence D. Jackel · 1989
Earlier work this paper cites.
Design improvements in associative memories for cerebellar model articulation controllers (CMAC)
P. C. Edgar An, W. Thomas Miller III, and P. C. Parks · 1991
Earlier work this paper cites.
Theory and development of higher-order cmac neural networks
S.H. Lane, D.A. Handelman, and J.J. Gelfand · 1992
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
Christopher G. Atkeson and Juan Carlos Santamaría · 1997
Earlier work this paper cites.
On the role of tracking in stationary environments
Richard S. Sutton, Anna Koop, and David Silver · 2007
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S. Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael H. Bowling · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E. Taylor and Peter Stone · 2009
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Hassabis, Shane Legg, and Stig Petersen · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation
Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell · 2016
Cited alongside, same era.
State of the art control of Atari games using shallow reinforcement learning
Yitao Liang, Marlos C. Machado, Erik Talvitie, and Michael H. Bowling · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Self-correcting models for model-based reinforcement learning
Erik Talvitie · 2017
Cited alongside, same era.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Fuzzy tiling activations: A simple approach to learning sparse representations online
Yangchen Pan, Kirby Banman, and Martha White · 2021
Later among the works it cites.
Smaller world models for reinforcement learning
Jan Robine, Tobias Uelwer, and Stefan Harmeling · 2021
Later among the works it cites.
Planning in stochastic environments with a learned model
Ioannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert, and David Silver · 2022
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Danijar Hafner · 2022
Later among the works it cites.
Few-shot image generation using discrete content representation
Yan Hong, Li Niu, Jianfu Zhang, and Liqing Zhang · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abien Fred Agarap · 2018
Cited alongside, same era.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I. Jordan · 2018
Cited alongside, same era.
Revisiting the Arcade Learning Environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael Bowling · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2022
Later among the works it cites.
Feature generalization in deep reinforcement learning: An investigation into representation properties, 2022
Erfan Miahi · 2022
Later among the works it cites.
The Alberta plan for AI research
Richard S. Sutton, Michael H. Bowling, and Patrick M. Pilarski · 2022
Later among the works it cites.
Investigating the properties of neural network representations in reinforcement learning
Han Wang, Erfan Miahi, Martha White, Marlos C. Machado, Zaheer Abbas, Raksha Kumaraswamy, Vincent Liu, and Adam White · 2022
Later among the works it cites.
Loss of plasticity in continual deep reinforcement learning
Zaheer Abbas, Rosie Zhao, Joseph Modayil, Adam White, and Marlos C. Machado · 2023
Closest in time.
Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo de Lazcano, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan Terry · 2023
Closest in time.
Learning disentangled discrete representations
David Friede, Christian Reimers, Heiner Stuckenschmidt, and Mathias Niepert · 2023
Closest in time.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy P. Lillicrap · 2023
Closest in time.
Continual learning as computationally constrained reinforcement learning
Saurabh Kumar, Henrik Marklund, Ashish Rao, Yifan Zhu, Hong Jun Jeon, Yueyang Liu, and Benjamin Van Roy · 2023
Closest in time.
Transformers are sample-efficient world models
Vincent Micheli, Eloi Alonso, and François Fleuret · 2023
Closest in time.