Fetching the paper…
Reading the bibliography…
Motivated by the prevailing paradigm of using unsupervised learning for efficient exploration in reinforcement learning (RL) problems [tang2017exploration,bellemare2016unifying], we investigate when this paradigm is provably efficient.
A theoretical analysis of deep Q-learning
Yang, Z., Xie, Y., and Wang, Z. (2019) · 1901
Earlier work this paper cites.
Efficient model-free reinforcement learning in metric spaces
Song, Z. and Sun, W. (2019) · 1905
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2019) · 1907
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2019) · 1908
Earlier work this paper cites.
Mixture models: Inference and applications to clustering
McLachlan, G. J. and Basford, K. E. (1988) · 1988
Earlier work this paper cites.
Error bounds for approximate value iteration
Munos, R. (2005) · 1999
Earlier work this paper cites.
A two-round variant of EM for Gaussian mixtures
Dasgupta, S. and Schulman, L. J. (2000) · 2000
Earlier work this paper cites.
Learning mixtures of arbitrary Gaussians
Arora, S. and Kannan, R. (2001) · 2001
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Charikar, M. S. (2002) · 2002
Earlier work this paper cites.
Du, S. S., Lee, J. D., Mahajan, G., and Wang, R. (2020b) · 2002
Earlier work this paper cites.
On the use of Bernoulli mixture models for text classification
Juan, A. and Vidal, E. (2002) · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Policy search by dynamic programming
Bagnell, J. A., Kakade, S. M., Schneider, J. G., and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Bernoulli mixture models for binary images
Juan, A. and Vidal, E. (2004) · 2004
Earlier work this paper cites.
Finite mixture models
McLachlan, G. J. and Peel, D. (2004) · 2004
Earlier work this paper cites.
A spectral algorithm for learning mixture models
Vempala, S. and Wang, G. (2004) · 2004
Earlier work this paper cites.
On spectral learning of mixtures of distributions
Achlioptas, D. and McSherry, F. (2005) · 2005
Earlier work this paper cites.
Model-based clustering for expression data via a Dirichlet process mixture model
Dahl, D. B. (2006) · 2006
Earlier work this paper cites.
Two-way Poisson mixture models for simultaneous document classification and word clustering
Li, J. and Zha, H. (2006) · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
Strehl, A. L., Li, L., Wiewiora, E., Langford, J., and Littman, M. L. (2006) · 2006
Cited alongside, same era.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R. (2008) · 2008
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
Knows what it knows: A framework for self-aware learning
Li, L., Littman, M. L., Walsh, T. J., and Strehl, A. L. (2011) · 2011
Cited alongside, same era.
Subspace clustering
Vidal, R. (2011) · 2011
Cited alongside, same era.
Clustering and variable selection for categorical multivariate data
Bontemps, D. and Toussile, W. (2013) · 2013
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2017) · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. (2017) · 2017
Later among the works it cites.
On learning mixtures of well-separated Gaussians
Regev, O. and Vijayaraghavan, A. (2017) · 2017
Later among the works it cites.
# Exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Efficient exploration through Bayesian deep Q-networks
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A. (2018) · 2018
Later among the works it cites.
On oracle-efficient PAC RL with rich observations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sparse subspace clustering: Algorithm, theory, and applications
Elhamifar, E. and Vidal, R. (2013) · 2013
Cited alongside, same era.
PAC optimal exploration in continuous space Markov decision processes
Pazis, J. and Parr, R. (2013) · 2013
Cited alongside, same era.
Provable subspace clustering: When LRR meets SSC
Wang, Y.-X., Xu, H., and Leng, C. (2013) · 2013
Cited alongside, same era.
Efficient exploration and value function generalization in deterministic systems
Wen, Z. and Van Roy, B. (2013) · 2013
Cited alongside, same era.
Local policy search in a convex space and conservative policy iteration as boosted policy search
Scherrer, B. and Geist, M. (2014) · 2014
Cited alongside, same era.
Robust subspace clustering
Soltanolkotabi, M., Elhamifar, E., Candes, E. J., et al. (2014) · 2014
Cited alongside, same era.
Dann, C., Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2018) · 2018
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Hessel, M., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S. (2018) · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Variance reduction methods for sublinear reinforcement learning
Kakade, S., Wang, M., and Yang, L. F. (2018) · 2018
Later among the works it cites.
BBQ-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Lipton, Z. C., Li, X., Gao, J., Li, L., Ahmed, F., and Deng, L. (2018) · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N. (2019) · 2019
Later among the works it cites.
A theory of regularized Markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Later among the works it cites.
Learning to control in metric space with optimal regret
Ni, C., Yang, L. F., and Wang, M. (2019) · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
Simchowitz, M. and Jamieson, K. G. (2019) · 2019
Later among the works it cites.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J. (2019) · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L. F. and Wang, M. (2019) · 2019
Later among the works it cites.
Zanette, A. and Brunskill, E. (2019) · 2019
Later among the works it cites.
Mixture models and applications
Bouguila, N. and Fan, W. (2020) · 2020
Closest in time.
Reliable clustering of Bernoulli mixture models
Najafi, A., Motahari, S. A., and Rabiee, H. R. (2020) · 2020
Closest in time.