Fetching the paper…
Reading the bibliography…
We introduce Recurrent Predictive State Policy (RPSP) networks, a recurrent architecture that brings insights from predictive state representations to reinforcement learning in partially observable environments.
Backpropagation through time: what does it do and how to do it
Werbos, P · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, Evan, Bartlett, Peter L., and Baxter, Jonathan · 2001
Earlier work this paper cites.
Predictive representations of state
Littman, Michael L., Sutton, Richard S., and Singh, Satinder · 2001
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R., Mcallester, D., Singh, S., and Mansour, Y · 2001
Earlier work this paper cites.
Learning low dimensional predictive representations
Rosencrantz, Matthew and Gordon, Geoff · 2004
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Singh, Satinder, James, Michael R., and Rudary, Matthew R · 2004
Earlier work this paper cites.
Learning tetris using the noisy cross-entropy method
Szita, I. and Lörincz, A · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, Ali and Recht, Benjamin · 2008
Earlier work this paper cites.
Predictive state temporal difference learning
Boots, Byron and Gordon, Geoffrey J · 2010
Earlier work this paper cites.
Recurrent policy gradients
Wierstra, D., Förster, A., Peters, J., and Schmidhuber, J · 2010
Cited alongside, same era.
Closing the learning planning loop with predictive state representations
Boots, Byron, Siddiqi, Sajid, and Gordon, Geoffrey · 2011
Cited alongside, same era.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Halko, N., Martinsson, P. G., and Tropp, J. A · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, Stéphane, Gordon, Geoffrey J., and Bagnell, Drew · 2011
Cited alongside, same era.
Hilbert Space Embeddings of Predictive State Representations
Boots, Byron, Gretton, Arthur, and Gordon, Geoffrey J · 2013
Cited alongside, same era.
Kernel bayes’ rule: Bayesian inference with positive definite kernels
Reinforcement learning of pomdp’s using spectral methods
Azizzadenesheli, Kamyar, Lazaric, Alessandro, and Anandkumar, Animashree · 2016
Later among the works it cites.
End to end learning for self-driving cars
Bojarski, Mariusz, Testa, Davide Del, Dworakowski, Daniel, Firner, Bernhard, Flepp, Beat, Goyal, Prasoon, Jackel, Lawrence D., Monfort, Mathew, Muller, Urs, Zhang, Jiakai, Zhang, Xin, Zhao, Jake, and Zieba, Karol · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Later among the works it cites.
Backprop kf: Learning discriminative deterministic state estimators
Haarnoja, Tuomas, Ajay, Anurag, Levine, Sergey, and Abbeel, Pieter · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J., Guez, Arthur, Sifre, Laurent, van den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, Dieleman, Sander, Grewe, Dominik, Nham, John, Kalchbrenner, Nal, Sutskever, Ilya, Lillicrap, Timothy, Leach, Madeleine, Kavukcuoglu, Koray, Graepel, Thore, and Hassabis, Demis · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fukumizu, Kenji, Song, Le, and Gretton, Arthur · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning, 2013
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, Kyunghyun, van Merrienboer, Bart, Gülçehre, Çaglar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
Efficient learning and planning with compressed predictive states
Hamilton, William, Fard, Mahdi Milani, and Pineau, Joelle · 2014
Cited alongside, same era.
Supervised learning for dynamical system learning
Hefny, Ahmed, Downey, Carlton, and Gordon, Geoffrey J · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael, and Moritz, Philipp · 2015
Cited alongside, same era.
Later among the works it cites.
Learning to filter with predictive state inference machines
Sun, Wen, Venkatraman, Arun, Boots, Byron, and Bagnell, J. Andrew · 2016
Later among the works it cites.
Value iteration networks
Tamar, Aviv, Levine, Sergey, Abbeel, Pieter, Wu, Yi, and Thomas, Garrett · 2016
Later among the works it cites.
Predictive state recurrent neural networks
Downey, Carlton, Hefny, Ahmed, Boots, Byron, Gordon, Geoffrey J, and Li, Boyue · 2017
Later among the works it cites.
Predictive state decoders: Encoding the future into recurrent networks
Venkatraman, Arun, Rhinehart, Nicholas, Sun, Wen, Pinto, Lerrel, Boots, Byron, Kitani, Kris, and Bagnell., James Andrew · 2017
Later among the works it cites.
An efficient, expressive and local minima-free method for learning controlled dynamical systems
Hefny, Ahmed, Downey, Carlton, and Gordon, Geoffrey J · 2018
Closest in time.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 2018
Closest in time.