Fetching the paper…
Reading the bibliography…
When agents interact with a complex environment, they must form and maintain beliefs about the relevant aspects of that environment.
Predictive coding–i
Peter Elias · 1955
Earlier work this paper cites.
Optimal control of markov processes with incomplete state information
Karl J Astrom · 1965
Earlier work this paper cites.
A dual back-propagation scheme for scalar reward learning
Paul Munro · 1987
Earlier work this paper cites.
Learning how the world works: Specifications for predictive networks in robots and brains
Paul J Werbos · 1987
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
Lawrence R Rabiner · 1989
Earlier work this paper cites.
The truck backer-upper: An example of self-learning in neural networks
Derrick Nguyen and Bernard Widrow · 1990
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Memory approaches to reinforcement learning in non-Markovian domains
Long-Ji Lin and Tom M Mitchell · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable markov decision problems
Tommi Jaakkola, Satinder P Singh, and Michael I Jordan · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Value-function approximations for partially observable markov decision processes
Milos Hauskrecht · 2000
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman · 2008
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Peter Deisenroth and Carl E. Rasmussen · 2011
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Recurrent reinforcement learning: a hybrid approach
Xiujun Li, Lihong Li, Jianfeng Gao, Xiaodong He, Jianshu Chen, Li Deng, and Ji He · 2015
Earlier work this paper cites.
Towards conceptual compression
Karol Gregor, Frederic Besse, Danilo Jimenez Rezende, Ivo Danihelka, and Daan Wierstra · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Cited alongside, same era.
Density estimation using real nvp
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2016
Cited alongside, same era.
Model-based reinforcement learning with parametrized physical models and optimism-driven exploration
Chris Xie, Sachin Patil, Teodor Moldovan, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Cited alongside, same era.
Neural belief states for partially observed domains
Pol Moreno, Jan Humplik, George Papamakarios, Bernardo Avila Pires, Lars Buesing, Nicolas Heess, and Theophane Weber · 2018
Later among the works it cites.
Generative temporal models with spatial memory for partially observed environments
Marco Fraccaro, Danilo Jimenez Rezende, Yori Zwols, Alexander Pritzel, SM Eslami, and Fabio Viola · 2018
Later among the works it cites.
Neural scene representation and rendering
SM Ali Eslami, Danilo Jimenez Rezende, Frederic Besse, Fabio Viola, Ari S Morcos, Marta Garnelo, Avraham Ruderman, Andrei A Rusu, Ivo Danihelka, Karol Gregor, et al · 2018
Later among the works it cites.
The kanerva machine: A generative distributed memory
Yan Wu, Greg Wayne, Alex Graves, and Timothy Lillicrap · 2018
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to act by predicting the future
Alexey Dosovitskiy and Vladlen Koltun · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Tejas D Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Cited alongside, same era.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Cited alongside, same era.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racanière, Daan Wierstra, and Shakir Mohamed · 2017
Cited alongside, same era.
Video pixel networks
Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Alexander A Alemi, Ben Poole, Ian Fischer, Joshua V Dillon, Rif A Saurous, and Kevin Murphy · 2017
Cited alongside, same era.
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P Kingma · 2017
Cited alongside, same era.
Chia-Chun Hung, Timothy Lillicrap, Josh Abramson, Yan Wu, Mehdi Mirza, Federico Carnevale, Arun Ahuja, and Greg Wayne · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Later among the works it cites.
Neural predictive belief representations
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo A. Pires, Toby Pohlen, and Rémi Munos · 2018
Later among the works it cites.
Generalized elbo with constrained optimization, geco
Danilo Jimenez Rezende and Fabio Viola · 2018
Later among the works it cites.
Learning and querying fast generative models for reinforcement learning
Lars Buesing, Theophane Weber, Sebastien Racaniere, SM Eslami, Danilo Rezende, David P Reichert, Fabio Viola, Frederic Besse, Karol Gregor, Demis Hassabis, et al · 2018
Later among the works it cites.
David Ha and Jürgen Schmidhuber · 2018
Later among the works it cites.
John D Co-Reyes, YuXuan Liu, Abhishek Gupta, Benjamin Eysenbach, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Brandon Amos, Laurent Dinh, Serkan Cabi, Thomas Rothörl, Sergio Gómez Colmenarejo, Alistair Muldal, Tom Erez, Yuval Tassa, Nando de Freitas, and Misha Denil · 2018
Later among the works it cites.
Geometric consistency for self-supervised end-to-end visual odometry
Ganesh Iyer, J Krishna Murthy, Gunshi Gupta, Madhava Krishna, and Liam Paull · 2018
Later among the works it cites.
Neural map: Structured memory for deep reinforcement learning
Emilio Parisotto and Ruslan Salakhutdinov · 2018
Later among the works it cites.
Navigation and planning in latent maps
Baris Kayalibay, Atanas Mirchev, Maximilian Soelch, Patrick van der Smagt, and Justin Bayer · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Learning attractor dynamics for generative memory
Yan Wu, Gregory Wayne, Karol Gregor, and Timothy Lillicrap · 2018
Later among the works it cites.
Conditional adversarial generative flow for controllable image synthesis
Rui Liu, Yu Liu, Xinyu Gong, Xiaogang Wang, and Hongsheng Li · 2019
Closest in time.
Preventing posterior collapse with delta-vaes
Ali Razavi, Aäron van den Oord, Ben Poole, and Oriol Vinyals · 2019
Closest in time.