Fetching the paper…
Reading the bibliography…
In this paper we study online Reinforcement Learning (RL) in partially observable dynamical systems.
Maximum matching and a polyhedron with 0, 1-vertices
Edmonds, J. (1965) · 1965
Earlier work this paper cites.
Nonnegative ranks, decompositions, and factorizations of nonnegative matrices
Cohen, J. E. and Rothblum, U. G. (1993) · 1993
Earlier work this paper cites.
Discrete-time, discrete-valued observable operator models: a tutorial
Jaeger, H. (1998) · 1998
Earlier work this paper cites.
Approximate planning in large pomdps via reusable trajectories
Kearns, M., Mansour, Y., and Ng, A. (1999) · 1999
Earlier work this paper cites.
Empirical Processes in M-estimation
Geer, S. A., van de Geer, S., and Williams, D. (2000) · 2000
Earlier work this paper cites.
Predictive representations of state
Littman, M. and Sutton, R. S. (2001) · 2001
Earlier work this paper cites.
Regret minimization in partially observable linear quadratic control
Lale, S., Azizzadenesheli, K., Hassibi, B., and Anandkumar, A. (2020) · 2002
Earlier work this paper cites.
Learning low dimensional predictive representations
Rosencrantz, M., Gordon, G., and Thrun, S. (2004) · 2004
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Singh, S., James, M. R., and Rudary, M. R. (2004) · 2004
Earlier work this paper cites.
Reinforcement learning in pomdps without resets
Even-Dar, E., Kakade, S. M., and Mansour, Y. (2005) · 2005
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A., Kakade, S., Krishnamurthy, A., and Sun, W. (2020) · 2006
Earlier work this paper cites.
Hilbert Space Embeddings of Hidden Markov Models
Boots, B., Siddiqi, S. M., Gordon, G., and Smola, A. (2010) · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Closing the learning-planning loop with predictive state representations
Boots, B., Siddiqi, S. M., and Gordon, G. J. (2011) · 2011
Earlier work this paper cites.
A spectral algorithm for learning hidden markov models
Hsu, D., Kakade, S. M., and Zhang, T. (2012) · 2012
Earlier work this paper cites.
Hilbert space embeddings of predictive state representations
Boots, B., Gordon, G., and Gretton, A. (2013) · 2013
Earlier work this paper cites.
Combinatorial bounds on nonnegative rank and extended formulations
Fiorini, S., Kaibel, V., Pashkovich, K., and Theis, D. O. (2013) · 2013
Cited alongside, same era.
Efficient learning and planning with compressed predictive states
Hamilton, W., Fard, M. M., and Pineau, J. (2014) · 2014
Cited alongside, same era.
Low-rank spectral learning
Kulesza, A., Rao, N. R., and Singh, S. (2014) · 2014
Cited alongside, same era.
Supervised learning for dynamical system learning
Hefny, A., Downey, C., and Gordon, G. J. (2015) · 2015
Cited alongside, same era.
Links between multiplicity automata, observable operator models and predictive state representations: a unified learning framework
Thon, M. R. and Jaeger, H. (2015) · 2015
Cited alongside, same era.
Reinforcement learning of pomdps using spectral methods
Azizzadenesheli, K., Lazaric, A., and Anandkumar, A. (2016) · 2016
Completing state representations using spectral learning
Jiang, N., Kulesza, A., and Singh, S. (2018) · 2018
Later among the works it cites.
Provably efficient Q-learning with function approximation via distribution shift error checking oracle
Du, S. S., Luo, Y., Wang, R., and Zhang, H. (2019) · 2019
Later among the works it cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J. (2019) · 2019
Later among the works it cites.
Sample-efficient reinforcement learning of undercomplete pomdps
Jin, C., Kakade, S., Krishnamurthy, A., and Liu, Q. (2020) · 2020
Later among the works it cites.
Improper learning for non-stochastic control
Simchowitz, M., Singh, K., and Hazan, E. (2020) · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A pac rl algorithm for episodic pomdps
Guo, Z. D., Doroudi, S., and Brunskill, E. (2016) · 2016
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2016) · 2016
Cited alongside, same era.
PAC reinforcement learning with rich observations
Krishnamurthy, A., Agarwal, A., and Langford, J. (2016) · 2016
Cited alongside, same era.
Learning to filter with predictive state inference machines
Sun, W., Venkatraman, A., Boots, B., and Bagnell, J. A. (2016) · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Predictive state recurrent neural networks
Downey, C., Hefny, A., Boots, B., Gordon, G. J., and Li, B. (2017) · 2017
Cited alongside, same era.
Yang, L. and Wang, M. (2020) · 2020
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Du, S. S., Kakade, S. M., Lee, J. D., Lovett, S., Mahajan, G., Sun, W., and Wang, R. (2021) · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Foster, D. J., Kakade, S. M., Qian, J., and Rakhlin, A. (2021) · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Jin, C., Liu, Q., and Miryoosefi, S. (2021) · 2021
Later among the works it cites.
Rl for latent mdps: Regret guarantees and a lower bound
Kwon, J., Efroni, Y., Caramanis, C., and Mannor, S. (2021) · 2021
Later among the works it cites.
Reinforcement learning under a multi-agent predictive state representation model: Method and theory
Zhang, Z., Yang, Z., Liu, H., Tokekar, P., and Huang, F. (2021) · 2021
Later among the works it cites.
Sample-efficient reinforcement learning for pomdps with linear function approximations
Cai, Q., Yang, Z., and Wang, Z. (2022) · 2022
Closest in time.
Provable reinforcement learning with a short-term memory
Efroni, Y., Jin, C., Krishnamurthy, A., and Miryoosefi, S. (2022) · 2022
Closest in time.
When is partially observable reinforcement learning not scary?
Liu, Q., Chung, A., Szepesvári, C., and Jin, C. (2022) · 2022
Closest in time.
Provably efficient reinforcement learning in partially observable dynamical systems
Uehara, M., Sekhari, A., Lee, J. D., Kallus, N., and Sun, W. (2022) · 2022
Closest in time.
Embed to control partially observed systems: Representation learning with provable sample efficiency
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2022) · 2022
Closest in time.