Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) algorithms are known to suffer from the curse of dimensionality, which refers to the fact that large-scale problems often lead to exponentially high sample complexity.
Stochastic approximation with cone-contractive operators: Sharp l-infty bounds for q-learning
Wainwright, M. J. (2019a) · 1905
Earlier work this paper cites.
Variance-reduced q-learning is minimax optimal
Wainwright, M. J. (2019b) · 1906
Earlier work this paper cites.
Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales
Banach, S. (1922) · 1922
Earlier work this paper cites.
Computer solutions of the traveling salesman problem
Lin, S. (1965) · 1965
Earlier work this paper cites.
On chromatic number of graphs and set-systems
Erdős, P. and Hajnal, A. (1966) · 1966
Earlier work this paper cites.
The complexity of satisfiability problems
Schaefer, T. J. (1978) · 1978
Earlier work this paper cites.
Computers and intractability: A guide to the theory of np-completeness
Gary, M. R. and Johnson, D. S. (1979) · 1979
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T., Jordan, M., and Singh, S. (1993) · 1993
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
Tsitsiklis, J. N. (1994) · 1994
Earlier work this paper cites.
Exploiting structure in policy construction
Boutilier, C., Dearden, R., Goldszmidt, M., et al. (1995) · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. and Van Roy, B. (1996) · 1996
Earlier work this paper cites.
The chromatic number of oriented graphs
Sopena, E. (1997) · 1997
Earlier work this paper cites.
The asymptotic convergence-rate of q-learning
Szepesvári, C. (1997) · 1997
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Kearns, M. and Singh, S. (1998) · 1998
Earlier work this paper cites.
Decision-theoretic planning: Structural assumptions and computational leverage
Boutilier, C., Dean, T., and Hanks, S. (1999) · 1999
Earlier work this paper cites.
Improved inapproximability results for maxclique, chromatic number and approximate graph coloring
Khot, S. (2001) · 2001
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
Kearns, M., Mansour, Y., and Ng, A. Y. (2002) · 2002
Earlier work this paper cites.
Learning rates for q-learning
Even-Dar, E., Mansour, Y., and Bartlett, P. (2003) · 2003
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
Guestrin, C., Koller, D., Parr, R., and Venkataraman, S. (2003) · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. (2003) · 2003
Cited alongside, same era.
Multidimensional knapsack problems
Kellerer, H., Pferschy, U., Pisinger, D., Kellerer, H., Pferschy, U., and Pisinger, D. (2004) · 2004
Cited alongside, same era.
Reducibility among combinatorial problems
Karp, R. M. (2010) · 2010
Cited alongside, same era.
Graph coloring problems
Jensen, T. R. and Toft, B. (2011) · 2011
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2012) · 2012
Cited alongside, same era.
The s-packing chromatic number of a graph
Goddard, W. and Xu, H. (2012) · 2012
Cited alongside, same era.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., and Yang, L. F. (2020) · 2020
Later among the works it cites.
Efficient reinforcement learning in factored MDPs with application to constrained RL
Chen, X., Hu, J., Li, L., and Wang, L. (2020) · 2020
Later among the works it cites.
A theoretical analysis of deep Q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z. (2020) · 2020
Later among the works it cites.
Deep reinforcement learning for intelligent transportation systems: A survey
Haydari, A. and Yılmaz, Y. (2020) · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2020) · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Gheshlaghi Azar, M., Munos, R., and Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Cited alongside, same era.
Near-optimal reinforcement learning in factored MDPs
Osband, I. and Van Roy, B. (2014) · 2014
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2020) · 2020
Later among the works it cites.
Scalable reinforcement learning of localized policies for multi-agent networked systems
Qu, G., Wierman, A., and Li, N. (2020) · 2020
Later among the works it cites.
Robot modeling and control
Spong, M. W., Hutchinson, S., and Vidyasagar, M. (2020) · 2020
Later among the works it cites.
Towards minimax optimal reinforcement learning in factored markov decision processes
Tian, Y., Qian, J., and Sra, S. (2020) · 2020
Later among the works it cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded Eluder dimension
Wang, R., Salakhutdinov, R. R., and Yang, L. (2020) · 2020
Later among the works it cites.
A finite-time analysis of Q-learning with neural network function approximation
Xu, P. and Gu, Q. (2020) · 2020
Later among the works it cites.
Reinforcement learning in factored MDPs: Oracle-efficient algorithms and tighter regret bounds for the non-episodic setting
Xu, Z. and Tewari, A. (2020) · 2020
Later among the works it cites.
Reinforcement learning in economics and finance
Charpentier, A., Elie, R., and Remlinger, C. (2021) · 2021
Later among the works it cites.
Electricity Price Data
California ISO (2021) · 2022
Later among the works it cites.
Reinforcement learning for selective key applications in power systems: Recent advances and future challenges
Chen, X., Qu, G., Tang, Y., Low, S., and Li, N. (2022) · 2022
Later among the works it cites.
Target network and truncation overcome the deadly triad in Q Q -learning
Chen, Z., Clarke, J.-P., and Maguluri, S. T. (2023) · 2023
Later among the works it cites.
Variance reduced value iteration and faster algorithms for solving markov decision processes
Sidford, A., Wang, M., Wu, X., and Ye, Y. (2023) · 2023
Later among the works it cites.
Sustaingym: Reinforcement learning environments for sustainable energy systems
Yeh, C., Li, V., Datta, R., Arroyo, J., Christianson, N., Zhang, C., Chen, Y., Hosseini, M. M., Golmohammadi, A., Shi, Y., et al. (2023) · 2023
Later among the works it cites.
Is q-learning minimax optimal? a tight sample complexity analysis
Li, G., Cai, C., Chen, Y., Wei, Y., and Chi, Y. (2024) · 2024
Closest in time.
The curious price of distributional robustness in reinforcement learning with a generative model
Shi, L., Li, G., Wei, Y., Chen, Y., Geist, M., and Chi, Y. (2024) · 2024
Closest in time.
Sample efficient reinforcement learning in mixed systems through augmented samples and its applications to queueing networks
Wei, H., Liu, X., Wang, W., and Ying, L. (2024) · 2024
Closest in time.