Fetching the paper…
Reading the bibliography…
This paper investigates the potential of quantum acceleration in addressing infinite horizon Markov Decision Processes (MDPs) to enhance average reward outcomes.
Quantum measurements and the abelian stabilizer problem
Kitaev, A. Y · 1995
Earlier work this paper cites.
A fast quantum mechanical algorithm for database search
Grover, L. K · 1996
Earlier work this paper cites.
Quantum amplitude amplification and estimation
Brassard, G., Hoyer, P., Mosca, M., and Tapp, A · 2002
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J · 2003
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P., Jaksch, T., and Ortner, R · 2008
Earlier work this paper cites.
Quantum algorithm for linear systems of equations
Harrow, A. W., Hassidim, A., and Lloyd, S · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Quantum computation and quantum information
Nielsen, M. A. and Chuang, I. L · 2010
Earlier work this paper cites.
Quantum algorithms for supervised and unsupervised machine learning
Lloyd, S., Mohseni, M., and Rebentrost, P · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
A quantum approximate optimization algorithm
Farhi, E., Goldstone, J., and Gutmann, S · 2014
Earlier work this paper cites.
Quantum speedup for active learning agents
Paparo, G. D., Dunjko, V., Makmal, A., Martin-Delgado, M. A., and Briegel, H. J · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Quantum speedup of monte carlo methods
Montanaro, A · 2015
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Quantum machine learning
Biamonte, J., Wittek, P., Pancotti, N., Rebentrost, P., Wiebe, N., and Lloyd, S · 2017
Cited alongside, same era.
Advances in quantum reinforcement learning
Dunjko, V., Taylor, J. M., and Briegel, H. J · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, R., Pirotta, M., Lazaric, A., and Ortner, R · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Quantum exploration algorithms for multi-armed bandits
Wang, D., You, X., Li, T., and Childs, A. M · 2021
Later among the works it cites.
Decision making in monopoly using a hybrid deep reinforcement learning approach
Bonjour, T., Haliem, M., Alsalem, A., Thomas, S., Li, H., Aggarwal, V., Kejriwal, M., and Bhargava, B · 2022
Later among the works it cites.
Near-optimal quantum algorithms for multivariate mean estimation
Cornelissen, A., Hamoudi, Y., and Jerbi, S · 2022
Later among the works it cites.
Quantum policy gradient algorithms
Jerbi, S., Cornelissen, A., Ozols, M., and Dunjko, V · 2022
Later among the works it cites.
Quantum policy iteration via amplitude estimation and grover search–towards quantum advantage for reinforcement learning
Wiedemann, S., Hein, D., Udluft, S., and Mendl, C. B · 2022
Later among the works it cites.
Reinforcement learning for joint optimization of multiple rewards
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: Theory and algorithms
Agarwal, A., Jiang, N., Kakade, S. M., and Sun, W · 2019
Cited alongside, same era.
Deeppool: Distributed model-free algorithm for ride-sharing using deep reinforcement learning
Al-Abbasi, A. O., Ghosh, A., and Aggarwal, V · 2019
Cited alongside, same era.
Quantum singular value transformation and beyond: exponential improvements for quantum matrix arithmetics
Gilyén, A., Su, Y., Low, G. H., and Wiebe, N · 2019
Cited alongside, same era.
Online convex optimization in adversarial markov decision processes
Rosenberg, A. and Mansour, Y · 2019
Cited alongside, same era.
Quantum bandits
Casalé, B., Di Molfetta, G., Kadri, H., and Ralaivola, L · 2020
Cited alongside, same era.
Model-free reinforcement learning in infinite-horizon average-reward markov decision processes
Wei, C.-Y., Jahromi, M. J., Luo, H., Sharma, H., and Jain, R · 2020
Cited alongside, same era.
Agarwal, M. and Aggarwal, V · 2023
Closest in time.
Quantum speedups for zero-sum games via improved dynamic gibbs sampling
Bouland, A., Getachew, Y. M., Jin, Y., Sidford, A., and Tian, K · 2023
Closest in time.
Quantum computing provides exponential regret improvement in episodic reinforcement learning
Ganguly, B., Wu, Y., Wang, D., and Aggarwal, V · 2023
Closest in time.
Quantum speedups for stochastic optimization
Sidford, A. and Zhang, C · 2023
Closest in time.
Quantum multi-armed bandits and stochastic linear bandits enjoy logarithmic regrets
Wan, Z., Zhang, Z., Li, T., Zhang, J., and Sun, X · 2023
Closest in time.
Wu, Y., Guan, C., Aggarwal, V., and Wang, D · 2023
Closest in time.
Regret analysis of policy gradient algorithm for infinite horizon average reward markov decision processes
Bai, Q., Mondal, W. U., and Aggarwal, V · 2024
Closest in time.
Provably efficient exploration in quantum reinforcement learning with logarithmic worst-case regret
Zhong, H., Hu, J., Xue, Y., Li, T., and Wang, L · 2024
Closest in time.
Reinforcement learning for infinite-horizon average-reward linear mdps via approximation by discounted-reward mdps
Hong, K., Chae, W., Zhang, Y., Lee, D., and Tewari, A · 2025
Closest in time.
Accelerating quantum reinforcement learning with a quantum natural policy gradient based approach
Xu, Y. and Aggarwal, V · 2025
Closest in time.