Fetching the paper…
Reading the bibliography…
In complex environments with large discrete action spaces, effective decision-making is critical in reinforcement learning (RL).
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T., Jordan, M., and Singh, S · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Rubinstein, R · 1999
Earlier work this paper cites.
Reinforcement learning as classification: Leveraging modern classifiers
Lagoudakis, M. G. and Parr, R · 2003
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Sallans, B. and Hinton, G. E · 2004
Earlier work this paper cites.
Using continuous action spaces to solve discrete problems
Van Hasselt, H. and Wiering, M. A · 2009
Earlier work this paper cites.
Double q-learning
Hasselt, H · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Generalized value functions for large action sets
Pazis, J. and Parr, R · 2011
Earlier work this paper cites.
Fast reinforcement learning with large action sets using error-correcting output codes for mdp factorization
Dulac-Arnold, G., Denoyer, L., Preux, P., and Gallinari, P · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B · 2015
Earlier work this paper cites.
Deep reinforcement learning with an unbounded action space
He, J., Chen, J., He, X., Gao, J., Li, L., Deng, L., and Ostendorf, M · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Online symbolic gradient-based optimization for factored action mdps
Cui, H. and Khardon, R · 2016
Earlier work this paper cites.
Deep reinforcement learning with a combinatorial action space for predicting popular Reddit threads
He, J., Ostendorf, M., He, X., Chen, J., Gao, J., Li, L., and Deng, L · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Discrete sequential prediction of continuous actions for deep rl
Metz, L., Ibarz, J., Jaitly, N., and Davidson, J · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Weighted double q-learning
Zhang, Z., Pan, Z., and Kochenderfer, M. J · 2017
Cited alongside, same era.
Stable baselines
Hill, A., Raffin, A., Ernestus, M., Gleave, A., Kanervisto, A., Traore, R., Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2018
Cited alongside, same era.
Bic-ddpg: Bidirectionally-coordinated nets for deep multi-agent reinforcement learning
Wang, G., Shi, D., Xue, C., Jiang, H., and Wang, Y · 2020
Later among the works it cites.
Generating adjacency-constrained subgoals in hierarchical reinforcement learning
Zhang, T., Guo, S., Tan, T., Hu, X., and Chen, F · 2020
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Dulac-Arnold, G., Levine, N., Mankowitz, D. J., Li, J., Paduraru, C., Gowal, S., and Hester, T · 2021
Later among the works it cites.
Artificial intelligence for satellite communication: A review
Fourati, F. and Alouini, M.-S · 2021
Later among the works it cites.
A distributed model-free ride-sharing approach for joint matching, pricing, and dispatching using deep reinforcement learning
Haliem, M., Mani, G., Aggarwal, V., and Bhargava, B · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods
Quillen, D., Jang, E., Nachum, O., Finn, C., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Action branching architectures for deep reinforcement learning
Tavakoli, A., Pardo, F., and Kormushev, P · 2018
Cited alongside, same era.
Learn what not to learn: Action elimination with deep reinforcement learning
Zahavy, T., Haroush, M., Merlis, N., Mankowitz, D. J., and Mannor, S · 2018
Cited alongside, same era.
Deeppool: Distributed model-free algorithm for ride-sharing using deep reinforcement learning
Al-Abbasi, A. O., Ghosh, A., and Aggarwal, V · 2019
Cited alongside, same era.
Applications of deep reinforcement learning in communications and networking: A survey
Luong, N. C., Hoang, D. T., Gong, S., Niyato, D., Wang, P., Liang, Y.-C., and Kim, D. I · 2019
Cited alongside, same era.
Learning collaborative policies to solve np-hard routing problems
Kim, M., Park, J., et al · 2021
Later among the works it cites.
Reinforcement learning in factored action spaces using tensor decompositions
Mahajan, A., Samvelyan, M., Mao, L., Makoviychuk, V., Garg, A., Kossaifi, J., Whiteson, S., Zhu, Y., and Anandkumar, A · 2021
Later among the works it cites.
Reinforcement learning for combinatorial optimization: A survey
Mazyavkina, N., Sviridov, S., Ivanov, S., and Burnaev, E · 2021
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Peng, B., Rashid, T., Schroeder de Witt, C., Kamienny, P.-A., Torr, P., Böhmer, W., and Whiteson, S · 2021
Later among the works it cites.
Is bang-bang control all you need? solving continuous control with bernoulli policies
Seyde, T., Gilitschenski, I., Schwarting, W., Stellato, B., Riedmiller, M., Wulfmeier, M., and Rus, D · 2021
Later among the works it cites.
Adaptive ensemble q-learning: Minimizing estimation bias via error feedback
Wang, H., Lin, S., and Zhang, J · 2021
Later among the works it cites.
Combining decision making and trajectory planning for lane changing using deep reinforcement learning
Li, S., Wei, C., and Wang, Y · 2022
Later among the works it cites.
Solving continuous control via q-learning
Seyde, T., Werner, P., Schwarting, W., Gilitschenski, I., Riedmiller, M., Rus, D., and Wulfmeier, M · 2022
Later among the works it cites.
Deep reinforcement learning: a survey
Wang, X., Wang, S., Liang, X., Zhao, D., Huang, J., Xu, X., Dai, B., and Miao, Q · 2022
Later among the works it cites.
Handling large discrete action spaces via dynamic neighborhood construction
Akkerman, F., Luy, J., van Heeswijk, W., and Schiffer, M · 2023
Later among the works it cites.
Hybrid multi-agent deep reinforcement learning for autonomous mobility on demand systems
Enders, T., Harrison, J., Pavone, M., and Schiffer, M · 2023
Later among the works it cites.
Randomized greedy learning for non-monotone stochastic submodular maximization under full-bandit feedback
Fourati, F., Aggarwal, V., Quinn, C., and Alouini, M.-S · 2023
Later among the works it cites.
Asap: A semi-autonomous precise system for telesurgery during communication delays
Gonzalez, G., Balakuntala, M., Agarwal, M., Low, T., Knoth, B., Kirkpatrick, A. W., McKee, J., Hager, G., Aggarwal, V., Xue, Y., et al · 2023
Later among the works it cites.
Revalued: Regularised ensemble value-decomposition for factorisable markov decision processes
Ireland, D. and Montana, G · 2024
Closest in time.