Fetching the paper…
Reading the bibliography…
State construction is important for learning in partially observable environments.
Statistical Inference for Probabilistic Functions of Finite State Markov Chains
Baum, L. E., and Petrie, T. (1966) · 1966
Earlier work this paper cites.
Intelligence: Its Organization and Development
Cunningham, M. (1972) · 1972
Earlier work this paper cites.
A model for the encoding of experiential information
Becker, J. D. (1973) · 1973
Earlier work this paper cites.
Neural Network and Physical Systems with Emergent Collective Computational Abilities
Hopfield, J. J. (1982) · 1982
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M., and Cohen, N. J. (1989) · 1989
Earlier work this paper cites.
A Learning Algorithm for Continually Running Fully Recurrent Neural Networks.
Williams, R. J., and Zipser, D. (1989) · 1989
Earlier work this paper cites.
Made-up minds: a constructivist approach to artificial intelligence
Drescher, G. L. (1991) · 1991
Earlier work this paper cites.
Reinforcement learning with hidden states
Lin, L.-J., and Mitchell, T. M. (1993) · 1993
Earlier work this paper cites.
Fast Exact Multiplication by the Hessian
Pearlmutter, B. A. (1994) · 1994
Earlier work this paper cites.
Learning to use selective attention and short-term memory in sequential tasks
McCallum, R. A. (1996) · 1996
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S., and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Predictive representations of state
Littman, M. L., Sutton, R. S., and Singh, S. (2001) · 2001
Earlier work this paper cites.
Harnessing Nonlinearity: Predicting Chaotic Systems and Saving Energy in Wireless Communication
Jaeger, H., and Haas, H. (2004) · 2004
Earlier work this paper cites.
Temporal-Difference Networks
Sutton, R. S., and Tanner, B. (2004) · 2004
Earlier work this paper cites.
Online discovery and learning of predictive state representations
McCracken, P., and Bowling, M. H. (2005) · 2005
Earlier work this paper cites.
Using predictive representations to improve generalization in reinforcement learning
Rafols, E. J., Ring, M. B., Sutton, R. S., and Tanner, B. (2005) · 2005
Earlier work this paper cites.
Predictive linear-Gaussian models of stochastic dynamical systems
Rudary, M., Singh, S., and Wingate, D. (2005) · 2005
Earlier work this paper cites.
Temporal Abstraction in Temporal-difference Networks.
Sutton, R. S., Rafols, E. J., and Koop, A. (2005) · 2005
Earlier work this paper cites.
Temporal-Difference Networks with History.
Tanner, B., and Sutton, R. S. (2005) · 2005
Earlier work this paper cites.
Simultaneous localization and mapping
Durrant-Whyte, H., and Bailey, T. (2006) · 2006
Earlier work this paper cites.
Learning predictive representations in autonomous driving to improve deep reinforcement learning
Graves, D., Nguyen, N. M., Hassanzadeh, K., and Jin, J. (2020) · 2006
Earlier work this paper cites.
Javed, K., White, M., and Bengio, Y. (2020) · 2006
Earlier work this paper cites.
Mixtures of predictive linear gaussian models for nonlinear, stochastic dynamical systems
Wingate, D., and Singh, S. (2006) · 2006
Earlier work this paper cites.
Predictive state representations with options.
Wolfe, B., and Singh, S. P. (2006) · 2006
Earlier work this paper cites.
On-line discovery of temporal-difference networks
Makino, T., and Takagi, T. (2008) · 2008
Cited alongside, same era.
Coordinating with the future: the anticipatory nature of representation
Pezzulo, G. (2008) · 2008
Cited alongside, same era.
Temporal-Difference Networks for Dynamical Systems with Continuous Observations and Actions.
Vigorito, C. M. (2009) · 2009
Cited alongside, same era.
Perspectives on system identification
Ljung, L. (2010) · 2010
Cited alongside, same era.
Toward Off-Policy Learning Control with Function Approximation
Maei, H., Szepesvári, C., Bhatnagar, S., and Sutton, R. (2010) · 2010
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Later among the works it cites.
Lstm: A search space odyssey
Greff, K., Srivastava, R. K., Koutník, J., Steunebrink, B. R., and Schmidhuber, J. (2017) · 2017
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
English conversational telephone speech recognition by humans and machines
Saon, G., Kurata, G., Sercu, T., Audhkhasi, K., Thomas, S., Dimitriadis, D., Cui, X., Ramabhadran, B., Picheny, M., Lim, L.-L., et al. (2017) · 2017
Later among the works it cites.
Predictive-State Decoders: Encoding the Future into Recurrent Networks
Venkatraman, A., Rhinehart, N., Sun, W., Pinto, L., Hebert, M., Boots, B., Kitani, K., and Bagnell, J. (2017) · 2017
Later among the works it cites.
Unifying task specification in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Subramanian, J., Sinha, A., Seraj, R., and Mahajan, A. (2020) · 2010
Cited alongside, same era.
Closing the learning-planning loop with predictive state representations
Boots, B., Siddiqi, S., and Gordon, G. (2011) · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P., White, A., and Precup, D. (2011) · 2011
Cited alongside, same era.
A spectral algorithm for learning Hidden Markov Models
Hsu, D., Kakade, S., and Zhang, T. (2012) · 2012
Cited alongside, same era.
Gradient Temporal Difference Networks
Silver, D. (2012) · 2012
Cited alongside, same era.
Representation Search through Generate and Test.
Mahmood, A. R., and Sutton, R. S. (2013) · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks.
Pascanu, R., Mikolov, T., and Bengio, Y. (2013) · 2013
Cited alongside, same era.
White, M. (2017) · 2017
Later among the works it cites.
Vector-based navigation using grid-like representations in artificial agents
Banino, A., Barry, C., Uria, B., Blundell, C., Lillicrap, T., Mirowski, P., Pritzel, A., Chadwick, M. J., Degris, T., Modayil, J., et al. (2018) · 2018
Later among the works it cites.
Initialization matters: Orthogonal Predictive State Recurrent Neural Networks
Choromanski, K., Downey, C., and Boots, B. (2018) · 2018
Later among the works it cites.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. (2018) · 2018
Later among the works it cites.
Ghiassian, S., Patterson, A., White, M., Sutton, R. S., and White, A. (2018) · 2018
Later among the works it cites.
Predictions, surprise, and predictions of surprise in general value function architectures
Günther, J., Kearney, A., Dawson, M. R., Sherstan, C., and Pilarski, P. M. (2018) · 2018
Later among the works it cites.
Flux: Elegant Machine Learning with Julia
Innes, M. (2018) · 2018
Later among the works it cites.
Measuring catastrophic forgetting in neural networks
Kemker, R., McClure, M., Abitino, A., Hayes, T., and Kanan, C. (2018) · 2018
Later among the works it cites.
Predicting the future with multi-scale successor representations
Momennejad, I., and Howard, M. W. (2018) · 2018
Later among the works it cites.
Approximating real-time recurrent learning with random kronecker factors
Mujika, A., Meier, F., and Steger, A. (2018) · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018) · 2018
Later among the works it cites.
Directly estimating the variance of the { \{ \ \backslash lambda } \} -return using temporal-difference methods
Sherstan, C., Bennett, B., Young, K., Ashley, D. R., White, A., White, M., and Sutton, R. S. (2018) · 2018
Later among the works it cites.
Unbiased Online Recurrent Optimization
Tallec, C., and Ollivier, Y. (2018) · 2018
Later among the works it cites.
Learning longer-term dependencies in rnns with auxiliary losses
Trinh, T. H., Dai, A. M., Luong, M.-T., and Le, Q. V. (2018) · 2018
Later among the works it cites.
Optimal kronecker-sum approximation of real time recurrent learning
Benzing, F., Gauy, M. M., Mujika, A., Martinsson, A., and Steger, A. (2019) · 2019
Later among the works it cites.
Continuous learning in single-incremental-task scenarios
Maltoni, D., and Lomonaco, V. (2019) · 2019
Later among the works it cites.
Discovery of useful questions as auxiliary tasks
Veeriah, V., Hessel, M., Xu, Z., Rajendran, J., Lewis, R. L., Oh, J., van Hasselt, H. P., Silver, D., and Singh, S. (2019) · 2019
Later among the works it cites.
Representation and General Value Functions
Sherstan, C. (2020) · 2020
Later among the works it cites.
Investigating Objectives for Off-policy Value Estimation in Reinforcement Learning
Patterson, A., Ghiassian, S., Gupta, D., White, A., and White, M. (2021) · 2021
Later among the works it cites.