Fetching the paper…
Reading the bibliography…
We propose a recurrent RL agent with an episodic exploration mechanism that helps discovering good policies in text-based game environments.
Zork I, 1980
Infocom · 1980
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, Leslie Pack, Littman, Michael L, and Cassandra, Anthony R · 1998
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, Alexander L and Littman, Michael L · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, J Zico and Ng, Andrew Y · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, Matthew J. and Stone, Peter · 2015
Earlier work this paper cites.
Deep reinforcement learning with a natural language action space
He, Ji, Chen, Jianshu, He, Xiaodong, Gao, Jianfeng, Li, Lihong, Deng, Li, and Ostendorf, Mari · 2015
Cited alongside, same era.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, Karthik, Kulkarni, Tejas, and Barzilay, Regina · 2015
Cited alongside, same era.
Playing FPS games with deep reinforcement learning
Lample, Guillaume and Chaplot, Devendra Singh · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Osband, Ian, Blundell, Charles, Pritzel, Alexander, and Van Roy, Benjamin · 2016
Cited alongside, same era.
What can you do with a rock? affordance extraction via word embeddings
Count-based exploration in feature space for reinforcement learning
Martin, Jarryd, Sasikumar, Suraj Narayanan, Everitt, Tom, and Hutter, Marcus · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, Georg, Bellemare, Marc G, Oord, Aaron van den, and Munos, Rémi · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, Adam, Gross, Sam, Chintala, Soumith, Chanan, Gregory, Yang, Edward, DeVito, Zachary, Lin, Zeming, Desmaison, Alban, Antiga, Luca, and Lerer, Adam · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, Matthias, Houthooft, Rein, Dhariwal, Prafulla, Sidor, Szymon, Chen, Richard Y, Chen, Xi, Asfour, Tamim, Abbeel, Pieter, and Andrychowicz, Marcin · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fulda, Nancy, Ricks, Daniel, Murdoch, Ben, and Wingate, David · 2017
Cited alongside, same era.
Reinforcement learning and episodic memory in humans and animals: an integrative framework
Gershman, Samuel J and Daw, Nathaniel D · 2017
Cited alongside, same era.
Tang, Haoran, Houthooft, Rein, Foote, Davis, Stooke, Adam, Chen, Xi, Duan, Yan, Schulman, John, DeTurck, Filip, and Abbeel, Pieter · 2017
Later among the works it cites.
Textworld: A learning environment for text-based games
Côté, Marc-Alexandre, Kádár, Ákos, Yuan, Xingdi, Kybartas, Ben, Barnes, Tavian, Fine, Emery, Moore, James, Hausknecht, Matthew, Asri, Layla El, Adada, Mahmoud, Tay, Wendy, and Trischler, Adam · 2018
Closest in time.