Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) agents typically learn memoryless policies---policies that only consider the last observation when selecting actions.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
An optimization-based categorization of reinforcement learning environments
Michael L. Littman · 1993
Earlier work this paper cites.
Monte Carlo matrix inversion and reinforcement learning
Andrew Barto and Michael Duff · 1994
Earlier work this paper cites.
Memoryless policies: Theoretical limitations and practical results
Michael L. Littman · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
Satinder P. Singh, Tommi Jaakkola, and Michael I. Jordan · 1994
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable Markov decision problems
Tommi Jaakkola, Satinder P. Singh, and Michael I. Jordan · 1995
Earlier work this paper cites.
Learning policies with external memory
Leonid Peshkin, Nicolas Meuleau, and Leslie Pack Kaelbling · 1999
Earlier work this paper cites.
Predictive representations of state
Michael L. Littman, Richard S. Sutton, and Satinder Singh · 2002
Earlier work this paper cites.
Model-based Bayesian reinforcement learning in partially observable domains
Pascal Poupart and Nikos Vlassis · 2008
Earlier work this paper cites.
Finding optimal memoryless policies of POMDPs under the expected average reward criterion
Yanjie Li, Baoqun Yin, and Hongsheng Xi · 2011
Earlier work this paper cites.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
Ivo Grondman, Lucian Busoniu, Gabriel A. D. Lopes, and Robert Babuska · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Bayesian nonparametric methods for partially-observable reinforcement learning
Finale Doshi-Velez, David Pfau, Frank Wood, and Nicholas Roy · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
True online TD (lambda)
Harm Seijen and Richard S. Sutton · 2014
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, Aviv Tamar, et al · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable MDPs
Matthew Hausknecht and Peter Stone · 2015
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
Learning deep neural network policies with continuous memory states
Marvin Zhang, Zoe McCarthy, Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
OpenAI baselines
Christopher Hesse, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2017
Later among the works it cites.
Memory augmented control networks
Arbaaz Khan, Clark Zhang, Nikolay Atanasov, Konstantinos Karydis, Vijay Kumar, and Daniel D. Lee · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Control of memory, active perception, and action in minecraft
Junhyuk Oh, Valliappa Chockalingam, Satinder Singh, and Honglak Lee · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Minimalistic gridworld environment for OpenAI gym
Maxime Chevalier-Boisvert and Lucas Willems · 2018
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Chia-Chun Hung, Timothy Lillicrap, Josh Abramson, Yan Wu, Mehdi Mirza, Federico Carnevale, Arun Ahuja, and Greg Wayne · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Learning reward machines for partially observable reinforcement learning
Rodrigo Toro Icarte, Ethan Waldie, Toryn Q. Klassen, Rick Valenzano, Margarita P. Castro, and Sheila A. McIlraith · 2019
Later among the works it cites.
RL starter files
Lucas Willems · 2019
Later among the works it cites.
Learning causal state representations of partially observable environments
Amy Zhang, Zachary C. Lipton, Luis Pineda, Kamyar Azizzadenesheli, Anima Anandkumar, Laurent Itti, Joelle Pineau, and Tommaso Furlanello · 2019
Later among the works it cites.