Fetching the paper…
Reading the bibliography…
Real world applications of Reinforcement Learning (RL) are often partially observable, thus requiring memory.
A Short Survey On Memory Based Reinforcement Learning
Dhruv Ramani · 1904
Earlier work this paper cites.
On the power of the compass (or, why mazes are easier to search than graphs)
Manuel Blum and Dexter Kozen · 1978
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Temporal Difference Learning in Continuous Time and Space
Kenji Doya · 1995
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A survey of POMDP applications
Anthony R Cassandra · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, November 2020
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2005
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, and others · 2015
Earlier work this paper cites.
Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Hassabis, Shane Legg, and Stig Petersen · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, and others · 2016
Earlier work this paper cites.
ViZDoom: A Doom-based AI Research Platform for Visual Reinforcement Learning
Micha\l Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Earlier work this paper cites.
OpenAI Baselines, 2017
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Learning Overcomplete HMMs
Vatsal Sharan, Sham M Kakade, Percy S Liang, and Gregory Valiant · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Cited alongside, same era.
MiniWorld: Minimalistic 3D Environment for RL & Robotics Research, 2018
Maxime Chevalier-Boisvert · 2018
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Embodied Visual Navigation With Automatic Curriculum Learning in Real Environments
Steven D. Morad, Roberto Mecca, Rudra P. K. Poudel, Stephan Liwicki, and Roberto Cipolla · 2020
Later among the works it cites.
Behaviour Suite for Reinforcement Learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvári, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado van Hasselt · 2020
Later among the works it cites.
Stabilizing Transformers for Reinforcement Learning
Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphaël Lopez Kaufman, Aidan Clark, Seb Noury, Matthew Botvinick, Nicolas Heess, and Raia Hadsell · 2020
Later among the works it cites.
Can RL From Pixels be as Efficient as RL From State?, July 2020
Daniel Seita · 2020
Later among the works it cites.
pomdp_py: A Framework to Build and Solve POMDP Problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao · 2018
Cited alongside, same era.
RLlib: Abstractions for distributed reinforcement learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
RECURRENT EXPERIENCE REPLAY IN DISTRIBUTED REINFORCEMENT LEARNING
Steven Kapturowski, Georg Ostrovski, John Quan, Remi Munos, and Will Dabney · 2019
Cited alongside, same era.
The starcraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder De Witt, Gregory Farquhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson · 2019
Cited alongside, same era.
Kaiyu Zheng and Stefanie Tellex · 2020
Later among the works it cites.
gym-gridverse: Gridworld domains for fully and partially observable reinforcement learning, 2021
Andrea Baisero and Sammie Katt · 2021
Later among the works it cites.
Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability
Dibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang, Ryan P Adams, and Sergey Levine · 2021
Later among the works it cites.
CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms
Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, and Jeff Braga · 2021
Later among the works it cites.
RL for Latent MDPs: Regret Guarantees and a Lower Bound
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, and Shie Mannor · 2021
Later among the works it cites.
Stable-Baselines3: Reliable Reinforcement Learning Implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Linear Transformers Are Secretly Fast Weight Programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Later among the works it cites.
Recurrent Off-policy Baselines for Memory-based Continuous Control
Zhihan Yang and Hai Nguyen · 2021
Later among the works it cites.
VMAS: A Vectorized Multi-Agent Simulator for Collective Robot Learning
Matteo Bettini, Ryan Kortvelesy, Jan Blumenkamp, and Amanda Prorok · 2022
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2022
Later among the works it cites.
Efficiently Modeling Long Sequences with Structured State Spaces
Albert Gu, Karan Goel, and Christopher Re · 2022
Later among the works it cites.
Modeling Partially Observable Systems using Graph-Based Memory and Topological Priors
Steven Morad, Stephan Liwicki, Ryan Kortvelesy, Roberto Mecca, and Amanda Prorok · 2022
Later among the works it cites.
Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs
Tianwei Ni, Benjamin Eysenbach, and Ruslan Salakhutdinov · 2022
Later among the works it cites.