Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is a branch of machine learning which is employed to solve various sequential decision making problems without proper supervision.
Random sampling with a reservoir
Jeffrey S Vitter · 1985
Earlier work this paper cites.
The neurobiology of learning and memory
Richard F Thompson · 1986
Earlier work this paper cites.
Configural association theory: The role of the hippocampal formation in learning, memory, and amnesia
Robert J Sutherland and Jerry W Rudy · 1989
Earlier work this paper cites.
Simple memory: a theory for archicortex
David Marr, David Willshaw, and Bruce McNaughton · 1991
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Hippocampal mediation of stimulus representation: A computational theory
Mark A Gluck and Catherine E Myers · 1993
Earlier work this paper cites.
Incremental multi-step q-learning
Jing Peng and Ronald J Williams · 1994
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
James L McClelland, Bruce L McNaughton, and Randall C O’reilly · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximationtechnical
JN Tsitsiklis and B Van Roy · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control
Nathaniel D Daw, Yael Niv, and Peter Dayan · 2005
Earlier work this paper cites.
The hippocampus book
Per Andersen, Richard Morris, David Amaral, John O’Keefe, Tim Bliss, et al · 2007
Earlier work this paper cites.
Hippocampal contributions to control: the third way
Máté Lengyel and Peter Dayan · 2008
Earlier work this paper cites.
Decision making and reward in frontal cortex: complementary evidence from neurophysiological and neuropsychological studies
Steven W Kennerley and Mark E Walton · 2011
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Memory augmented control networks
Arbaaz Khan, Clark Zhang, Nikolay Atanasov, Konstantinos Karydis, Vijay Kumar, and Daniel D Lee · 2017
Later among the works it cites.
Neural map: Structured memory for deep reinforcement learning
Emilio Parisotto and Ruslan Salakhutdinov · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Later among the works it cites.
Neural episodic control
Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adria Puigdomenech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Cited alongside, same era.
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Sample-efficient deep reinforcement learning via episodic backward update
Su Young Lee, Sungik Choi, and Sae-Young Chung · 2018
Later among the works it cites.
Episodic memory deep q-networks
Zichuan Lin, Tianqi Zhao, Guangwen Yang, and Lintao Zhang · 2018
Later among the works it cites.
Now i remember! episodic memory for reinforcement learning, 2018
Ricky Loynd, Matthew Hausknecht, Lihong Li, and Li Deng · 2018
Later among the works it cites.
Episodic curiosity through reachability
Nikolay Savinov, Anton Raichuk, Raphaël Marinier, Damien Vincent, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly · 2018
Later among the works it cites.
University of waterloo - cs885 ”reinforcement learning” paper presentation, 2018
Andreas Stöckel · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Unsupervised predictive memory in a goal-directed agent
Greg Wayne, Chia-Chun Hung, David Amos, Mehdi Mirza, Arun Ahuja, Agnieszka Grabska-Barwinska, Jack Rae, Piotr Mirowski, Joel Z Leibo, Adam Santoro, et al · 2018
Later among the works it cites.
Vizdoom competitions: Playing doom from pixels
Marek Wydmuch, Michał Kempka, and Wojciech Jaśkowski · 2018
Later among the works it cites.
Integrating episodic memory into a reinforcement learning agent using reservoir sampling, 2018
Kenny J. Young, Shuo Yang, and Richard S. Sutton · 2018
Later among the works it cites.