Fetching the paper…
Reading the bibliography…
This paper presents a framework to tackle constrained combinatorial optimization problems using deep Reinforcement Learning (RL).
The complexity of flowshop and jobshop scheduling
M. R. Garey, D. S. Johnson, and R. Sethi · 1976
Earlier work this paper cites.
Computers and intractability: A guide to the theory of np-completeness, 1979
M. R. Gary and D. S. Johnson · 1979
Earlier work this paper cites.
Neural computation of decisions in optimization problems
J. J. Hopfield and D. W. Tank · 1985
Earlier work this paper cites.
Or-library: distributing test problems by electronic mail
J. E. Beasley · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist Reinforcement Learning
R. J. Williams · 1992
Earlier work this paper cites.
A tutorial survey of job-shop scheduling problems using genetic algorithms—i. representation
R. Cheng, M. Gen, and Y. Tsujimura · 1996
Earlier work this paper cites.
Handbook of evolutionary computation
T. Bäck, D. B. Fogel, and Z. Michalewicz · 1997
Earlier work this paper cites.
Nonlinear programming
D. P. Bertsekas · 1997
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
E. Altman · 1999
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
F. A. Gers, J. Schmidhuber, and F. Cummins · 1999
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
V. S. Borkar · 2005
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Operations management
B. Mahadevan · 2010
Cited alongside, same era.
Energy-aware resource allocation heuristics for efficient management of data centers for cloud computing
A. Beloglazov, J. Abawajy, and R. Buyya · 2012
Cited alongside, same era.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska · 2012
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Openai baselines, 2017
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Later among the works it cites.
Device placement optimization with reinforcement learning
A. Mirhoseini, H. Pham, Q. V. Le, B. Steiner, R. Larsen, Y. Zhou, N. Kumar, M. Norouzi, S. Bengio, and J. Dean · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Learning heuristics for the tsp by policy gradient
M. Deudon, P. Cournut, A. Lacoste, Y. Adulyasak, and L.-M. Rousseau · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
M.-T. Luong, H. Pham, and C. D. Manning · 2015
Cited alongside, same era.
Pointer networks
O. Vinyals, M. Fortunato, and N. Jaitly · 2015
Cited alongside, same era.
Neural combinatorial optimization with reinforcement learning
I. Bello, H. Pham, Q. V. Le, M. Norouzi, and S. Bengio · 2016
Cited alongside, same era.
A research survey: review of flexible job shop scheduling techniques
I. A. Chaudhry and A. A. Khan · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Resource management with deep reinforcement learning
H. Mao, M. Alizadeh, I. Menache, and S. Kandula · 2016
Cited alongside, same era.
W. Kool, H. van Hoof, and M. Welling · 2018
Later among the works it cites.
Reinforcement learning for solving the vehicle routing problem
M. Nazari, A. Oroojlooy, L. Snyder, and M. Takác · 2018
Later among the works it cites.
Reward constrained policy optimization
C. Tessler, D. J. Mankowitz, and S. Mannor · 2018
Later among the works it cites.
Learning to perform local rewriting for combinatorial optimization
X. Chen and Y. Tian · 2019
Later among the works it cites.
Learning scheduling algorithms for data processing clusters
H. Mao, M. Schwarzkopf, S. B. Venkatakrishnan, Z. Meng, and M. Alizadeh · 2019
Later among the works it cites.
Constrained reinforcement learning has zero duality gap
S. Paternain, L. Chamon, M. Calvo-Fullana, and A. Ribeiro · 2019
Later among the works it cites.