Fetching the paper…
Reading the bibliography…
Large discrete action spaces (LDAS) remain a central challenge in reinforcement learning.
Simulated annealing
Dimitris Bertsimas and John Tsitsiklis · 1993
Earlier work this paper cites.
Local branching
Matteo Fischetti and Andrea Lodi · 2003
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Brian Sallans and Geoffrey E. Hinton · 2004
Earlier work this paper cites.
Using continuous action spaces to solve discrete problems
Hado van Hasselt and Marco A Wiering · 2009
Earlier work this paper cites.
Value function approximation in reinforcement learning using the fourier basis
George Konidaris, Sarah Osentoski, and Philip Thomas · 2011
Earlier work this paper cites.
Generalized value functions for large action sets
Jason Pazis and Ron Parr · 2011
Earlier work this paper cites.
Fast reinforcement learning with large action sets using error-correcting output codes for MDP factorization
Gabriel Dulac-Arnold, Ludovic Denoyer, Philippe Preux, and Patrick Gallinari · 2012
Earlier work this paper cites.
Motor primitive discovery
Philip S. Thomas and Andrew G. Barto · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning, 2013
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Scalable nearest neighbor algorithms for high dimensional data
Marius Muja and David G. Lowe · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Gabriel Dulac-Arnold, Richard Evans, Hado van Hasselt, Peter Sunehag, Timothy Lillicrap, Jonathan Hunt, Timothy Mann, Theophane Weber, Thomas Degris, and Ben Coppin · 2015
Earlier work this paper cites.
Online symbolic gradient-based optimization for factored action mdps
Hao Cui and Roni Khardon · 2016
Earlier work this paper cites.
Deep reinforcement learning with a natural language action space
Ji He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, and Mari Ostendorf · 2016
Earlier work this paper cites.
A unified view of piecewise linear neural network verification
Rudy Bunel, Ilker Turkaslan, Philip H. S. Torr, Pushmeet Kohli, and Pawan Kumar Mudigonda · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Sahil Sharma, Aravind Suresh, Rahul Ramesh, and Balaraman Ravindran · 2017
Cited alongside, same era.
Lifted stochastic planning, belief propagation and marginal map
Hao Cui and Roni Khardon · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Action branching architectures for deep reinforcement learning
Arash Tavakoli, Fabio Pardo, and Petar Kormushev · 2018
Cited alongside, same era.
Learn what not to learn: Action elimination with deep reinforcement learning
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Gabriel Dulac-Arnold, Nir Levine, Daniel J. Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2021
Later among the works it cites.
Learning collaborative policies to solve np-hard routing problems
Minsu Kim, Jinkyoo Park, and joungho kim · 2021
Later among the works it cites.
A deep reinforcement learning approach for traffic signal control optimization, 2021
Zhenning Li, Chengzhong Xu, and Guohui Zhang · 2021
Later among the works it cites.
Reinforcement learning in factored action spaces using tensor decompositions
Anuj Mahajan, Mikayel Samvelyan, Lei Mao, Viktor Makoviychuk, Animesh Garg, Jean Kossaifi, Shimon Whiteson, Yuke Zhu, and Anima Anandkumar · 2021
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Boehmer, and Shimon Whiteson · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Zahavy, Matan Haroush, Nadav Merlis, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Algorithms for Optimization
Mykel J. Kochenderfer and Tim A. Wheeler · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
The natural language of actions
Guy Tennenholtz and Shie Mannor · 2019
Cited alongside, same era.
Learning to navigate the synthetically accessible chemical space using reinforcement learning
Sai Krishna Gottipati, Boris Sattarov, Sufeng Niu, Yashaswi Pathak, Haoran Wei, Shengchao Liu, Shengchao Liu, Simon Blackburn, Karam Thomas, Connor Coley, Jian Tang, Sarath Chandar, and Yoshua Bengio · 2020
Cited alongside, same era.
Generalization to new actions in reinforcement learning
Ayush Jain, Andrew Szot, and Joseph J. Lim · 2020
Cited alongside, same era.
Continuous-discrete reinforcement learning for hybrid control in robotics
Michael Neunert, Abbas Abdolmaleki, Markus Wulfmeier, Thomas Lampe, Tobias Springenberg, Roland Hafner, Francesco Romano, Jonas Buchli, Nicolas Heess, and Martin Riedmiller · 2020
Cited alongside, same era.
Jointly-learned state-action embedding for efficient reinforcement learning
Paul J. Pritz, Liang Ma, and Kin K. Leung · 2021
Later among the works it cites.
Bic-ddpg: Bidirectionally-coordinated nets for deep multi-agent reinforcement learning
Gongju Wang, Dianxi Shi, Chao Xue, Hao Jiang, and Yajie Wang · 2021
Later among the works it cites.
Deep reinforcement learning for inventory control: A roadmap
Robert N. Boute, Joren Gijsbrechts, Willem van Jaarsveld, and Nathalie Vanvuchelen · 2022
Later among the works it cites.
Learning pseudometric-based action representations for offline reinforcement learning
Pengjie Gu, Mengchen Zhao, Chen Chen, Dong Li, Jianye Hao, and Bo An · 2022
Later among the works it cites.
Know your action set: Learning action relations for reinforcement learning
Ayush Jain, Norio Kosaka, Kyung-Min Kim, and Joseph J Lim · 2022
Later among the works it cites.
Leveraging factored action spaces for efficient offline reinforcement learning in healthcare
Shengpu Tang, Maggie Makar, Michael Sjoding, Finale Doshi-Velez, and Jenna Wiens · 2022
Later among the works it cites.
The use of continuous action representations to scale deep reinforcement learning for inventory control
Nathalie Vanvuchelen, Bram de Moor, and Robert N. Boute · 2022
Later among the works it cites.
Hybrid multi-agent deep reinforcement learning for autonomous mobility on demand systems
Tobias Enders, James Harrison, Marco Pavone, and Maximilian Schiffer · 2023
Closest in time.
A spatial pyramid pooling-based deep reinforcement learning model for dynamic job-shop scheduling problem
Xinquan Wu and Xuefeng Yan · 2023
Closest in time.
Two-sided deep reinforcement learning for dynamic mobility-on-demand management with mixed autonomy
Jiaohong Xie, Yang Liu, and Nan Chen · 2023
Closest in time.