Fetching the paper…
Reading the bibliography…
Many real-world sequential decision making problems are partially observable by nature, and the environment model is typically unknown.
Optimal control of markov decision processes with incomplete state estimation
Astrom, Karl J · 1965
Earlier work this paper cites.
Overcoming incomplete perception with utile distinction memory
McCallum, R Andrew · 1993
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Barto, Andrew G, Bradtke, Steven J, and Singh, Satinder P · 1995
Earlier work this paper cites.
The context-tree weighting method: basic properties
Willems, Frans MJ, Shtarkov, Yuri M, and Tjalkens, Tjalling J · 1995
Earlier work this paper cites.
Reinforcement learning with selective perception and hidden state
McCallum, Andrew Kachites and Ballard, Dana · 1996
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, Leslie Pack, Littman, Michael L, and Cassandra, Anthony R · 1998
Earlier work this paper cites.
Approximate planning for factored pomdps using belief state simplification
McAllester, David A and Singh, Satinder · 1999
Earlier work this paper cites.
Monte carlo pomdps
Thrun, Sebastian · 2000
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bakker, Bram · 2002
Earlier work this paper cites.
Point-based value iteration: An anytime algorithm for pomdps
Pineau, Joelle, Gordon, Geoff, Thrun, Sebastian, et al · 2003
Earlier work this paper cites.
Model-based online learning of pomdps
Shani, Guy, Brafman, Ronen I, and Shimony, Solomon E · 2005
Earlier work this paper cites.
Solving deep memory pomdps with recurrent policy gradients
Wierstra, Daan, Foerster, Alexander, Peters, Jan, and Schmidhuber, Juergen · 2007
Earlier work this paper cites.
Exploiting locality of interaction in factored dec-pomdps
Oliehoek, Frans A, Spaan, Matthijs TJ, Whiteson, Shimon, and Vlassis, Nikos · 2008
Earlier work this paper cites.
Online planning algorithms for pomdps
Ross, Stéphane, Pineau, Joelle, Paquet, Sébastien, and Chaib-Draa, Brahim · 2008
Earlier work this paper cites.
Particle filter-based policy gradient in pomdps
Coquelin, Pierre-Arnaud, Deguest, Romain, and Munos, Rémi · 2009
Earlier work this paper cites.
A tutorial on particle filtering and smoothing: Fifteen years later
Doucet, Arnaud and Johansen, Adam M · 2009
Earlier work this paper cites.
Monte-carlo planning in large pomdps
Silver, David and Veness, Joel · 2010
Earlier work this paper cites.
A bayesian approach for learning and planning in partially observable markov decision processes
Ross, Stéphane, Pineau, Joelle, Chaib-draa, Brahim, and Kreitmann, Pierre · 2011
Cited alongside, same era.
Solving nonlinear continuous state-action-observation pomdps for mechanical systems with gaussian noise
Deisenroth, Marc Peter and Peters, Jan · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Cited alongside, same era.
Skip context tree switching
Bellemare, Marc, Veness, Joel, and Talvitie, Erik · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
Kingma, Diederik P and Welling, Max · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, Danilo Jimenez, Mohamed, Shakir, and Wierstra, Daan · 2014
Importance weighted autoencoders
Burda, Yuri, Grosse, Roger, and Salakhutdinov, Ruslan · 2016
Later among the works it cites.
Learning to communicate to solve riddles with deep distributed recurrent q-networks
Foerster, Jakob N, Assael, Yannis M, de Freitas, Nando, and Whiteson, Shimon · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Value iteration networks
Tamar, Aviv, Wu, Yi, Thomas, Garrett, Levine, Sergey, and Abbeel, Pieter · 2016
Later among the works it cites.
Openai baselines, 2017
Dhariwal, Prafulla, Hesse, Christopher, Klimov, Oleg, Nichol, Alex, Plappert, Matthias, Radford, Alec, Schulman, John, Sidor, Szymon, and Wu, Yuhuai · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Count-based frequency estimation with bounded memory
Bellemare, Marc G · 2015
Cited alongside, same era.
A recurrent latent variable model for sequential data
Chung, Junyoung, Kastner, Kyle, Dinh, Laurent, Goel, Kratarth, Courville, Aaron C, and Bengio, Yoshua · 2015
Cited alongside, same era.
Bayesian nonparametric methods for partially-observable reinforcement learning
Doshi-Velez, Finale, Pfau, David, Wood, Frank, and Roy, Nicholas · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable MDPs
Hausknecht, Matthew and Stone, Peter · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Heess, Nicolas, Hunt, Jonathan J, Lillicrap, Timothy P, and Silver, David · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Later among the works it cites.
Qmdp-net: Deep learning for planning under partial observability
Karkus, Peter, Hsu, David, and Lee, Wee Sun · 2017
Later among the works it cites.
Learning in pomdps with monte carlo tree search
Katt, Sammie, Oliehoek, Frans A, and Amato, Christopher · 2017
Later among the works it cites.
Playing fps games with deep reinforcement learning
Lample, Guillaume and Chaplot, Devendra Singh · 2017
Later among the works it cites.
Filtering variational objectives
Maddison, Chris J, Lawson, John, Tucker, George, Heess, Nicolas, Norouzi, Mohammad, Mnih, Andriy, Doucet, Arnaud, and Teh, Yee · 2017
Later among the works it cites.
Dynamic-depth context tree weighting
Messias, João V and Whiteson, Shimon · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
Wu, Yuhuai, Mansimov, Elman, Grosse, Roger B, Liao, Shun, and Ba, Jimmy · 2017
Later among the works it cites.
On improving deep reinforcement learning for POMDPs
Zhu, Pengfei, Li, Xin, and Poupart, Pascal · 2017
Later among the works it cites.
Belief state representation in the dopamine system
Babayan, Benedicte M, Uchida, Naoshige, and Gershman, Samuel J · 2018
Closest in time.
Auto-encoding sequential Monte Carlo
Le, Tuan Anh, Igl, Maximilian, Jin, Tom, Rainforth, Tom, and Wood, Frank · 2018
Closest in time.
Variational sequential monte carlo
Naesseth, Christian A, Linderman, Scott W, Ranganath, Rajesh, and Blei, David M · 2018
Closest in time.