Fetching the paper…
Reading the bibliography…
Deep reinforcement learning is successful in decision making for sophisticated games, such as Atari, Go, etc.
Principles of statistics
Michael George Bulmer · 1979
Earlier work this paper cites.
The complexity of markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bram Bakker · 2002
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang · 2006
Earlier work this paper cites.
Solving deep memory pomdps with recurrent policy gradients
Daan Wierstra, Alexander Foerster, Jan Peters, and Juergen Schmidhuber · 2007
Earlier work this paper cites.
Sarsop: Efficient point-based pomdp planning by approximating optimally reachable belief spaces
Hanna Kurniawati, David Hsu, and Wee Sun Lee · 2008
Earlier work this paper cites.
Monte-carlo planning in large pomdps
David Silver and Joel Veness · 2010
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
DESPOT: Online POMDP planning with regularization
Adhiraj Somani, Nan Ye, David Hsu, and Wee Sun Lee · 2013
Earlier work this paper cites.
Intention-aware online pomdp planning for autonomous driving in a crowd
Haoyu Bai, Shaojun Cai, Nan Ye, David Hsu, and Wee Sun Lee · 2015
Earlier work this paper cites.
A recurrent latent variable model for sequential data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio · 2015
Earlier work this paper cites.
Draw: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra · 2015
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Learning to communicate to solve riddles with deep distributed recurrent q-networks
Jakob N Foerster, Yannis M Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Backprop KF: Learning discriminative deterministic state estimators
Tuomas Haarnoja, Anurag Ajay, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
End-to-end learnable histogram filters
Rico Jonschkowski and Oliver Brock · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Cited alongside, same era.
Optimizing agent behavior over long time scales by transporting value
Chia-Chun Hung, Timothy Lillicrap, Josh Abramson, Yan Wu, Mehdi Mirza, Federico Carnevale, Arun Ahuja, and Greg Wayne · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Differentiable Particle Filters: End-to-End Learning with Algorithmic Priors
Rico Jonschkowski, Divyam Rastogi, and Oliver Brock · 2018
Later among the works it cites.
Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation
Gregory Kahn, Adam Villaflor, Bosen Ding, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Particle filter networks with application to visual localization
Peter Karkus, David Hsu, and Wee Sun Lee · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
A brief survey of deep reinforcement learning
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath · 2017
Cited alongside, same era.
QMDP-net: Deep learning for planning under partial observability
Peter Karkus, David Hsu, and Wee Sun Lee · 2017
Cited alongside, same era.
Playing fps games with deep reinforcement learning
Guillaume Lample and Devendra Singh Chaplot · 2017
Cited alongside, same era.
Filtering variational objectives
Chris J Maddison, John Lawson, George Tucker, Nicolas Heess, Mohammad Norouzi, Andriy Mnih, Arnaud Doucet, and Yee Teh · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Auto-encoding sequential monte carlo
Tuan Anh Le, Maximilian Igl, Tom Rainforth, Tom Jin, and Frank Wood · 2018
Later among the works it cites.
Variational sequential monte carlo
Christian Naesseth, Scott Linderman, Rajesh Ranganath, and David Blei · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Gibson env: real-world perception for embodied agents
Fei Xia, Amir R. Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, and Silvio Savarese · 2018
Later among the works it cites.
Natural environment benchmarks for reinforcement learning
Amy Zhang, Yuxin Wu, and Joelle Pineau · 2018
Later among the works it cites.
On improving deep reinforcement learning for pomdps
Pengfei Zhu, Xin Li, Pascal Poupart, and Guanghui Miao · 2018
Later among the works it cites.
Modular visual navigation using active neural mapping
Devendra Singh Chaplot, Saurabh Gupta, Abhinav Gupta, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Deepmdp: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Later among the works it cites.
Shaping belief states with generative environment models for rl
Karol Gregor, Danilo Jimenez Rezende, Frederic Besse, Yan Wu, Hamza Merzic, and Aaron van den Oord · 2019
Later among the works it cites.
Particle filter recurrent neural networks
Xiao Ma, Peter Karkus, David Hsu, and Wee Sun Lee · 2019
Later among the works it cites.
Habitat: A platform for embodied AI research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al · 2019
Later among the works it cites.