Fetching the paper…
Reading the bibliography…
Reinforcement learning aims to learn optimal policies from interaction with environments whose dynamics are unknown.
The optimal control of partially observable markov processes over a finite horizon
Richard D Smallwood and Edward J Sondik · 1973
Earlier work this paper cites.
Asymptotic evaluation of certain Markov process expectations for large time, I
Monroe D Donsker and SR Srinivasa Varadhan · 1975
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bram Bakker · 2001
Earlier work this paper cites.
Probabilistic robotics
Sebastian Thrun · 2002
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Value iteration for continuous-state POMDPs
Josep M. Porta, Matthijs T. J. Spaan, and Nikos Vlassis · 2004
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I , volume 1
Dimitri Bertsekas · 2012
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Deep recurrent Q-learning for partially observable MDPs
Matthew Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap, and David Silver · 2015
Cited alongside, same era.
QMDP-net: Deep learning for planning under partial observability
Peter Karkus, David Hsu, and Wee Sun Lee · 2017
Later among the works it cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Later among the works it cites.
On improving deep reinforcement learning for POMDPs
Pengfei Zhu, Xin Li, Pascal Poupart, and Guanghui Miao · 2017
Later among the works it cites.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Minimal gated unit for recurrent neural networks
Guo-Bing Zhou, Jianxin Wu, Chen-Lin Zhang, and Zhi-Hua Zhou · 2016
Cited alongside, same era.
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Meta-trained agents implement Bayes-optimal agents
Vladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein, Miljan Martic, Shane Legg, and Pedro Ortega · 2020
Later among the works it cites.
A bio-inspired bistable recurrent cell allows for long-lasting memory
Nicolas Vecoven, Damien Ernst, and Guillaume Drion · 2021
Later among the works it cites.