Fetching the paper…
Reading the bibliography…
This paper gives a detailed review of reinforcement learning (RL) in combinatorial optimization, introduces the history of combinatorial optimization starting in the 1950s, and compares it with the RL algorithms of recent years.
R. Bellman, “Dynamic programming,” Science , vol. 153, pp. 34 – 37, 1957
1957
Earlier work this paper cites.
R. L. Karg and G. L. Thompson, “A heuristic approach to solving travelling salesman problems,” Management Science , vol. 10, pp. 225–248, 1964
1964
Earlier work this paper cites.
D. E. Kirk, “Optimal control theory: an introduction,” 1970
1970
Earlier work this paper cites.
G. W. Graves and A. Whinston, “An algorithm for the quadratic assignment problem,” Management Science , vol. 16, pp. 453–471, 1970
1970
Earlier work this paper cites.
C. J. C. H. Watkins, “Learning with delayed rewards,” 1989
1989
Earlier work this paper cites.
T. H. Cormen, C. E. Leiserson, and R. L. Rivest, “Introduction to algorithms,” 1990
1990
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, pp. 229–256, 1992
1992
Earlier work this paper cites.
L. M. Gambardella and M. Dorigo, “Ant-q: A reinforcement learning approach to the traveling salesman problem,” in Machine Learning Proceedings 1995 , A. Prieditis and S. Russell, Eds. San Francisco (CA): Morgan Kaufmann, 1995, pp. 252–260
1995
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” J. Artif. Intell. Res. , vol. 4, pp. 237–285, 1996
1996
Earlier work this paper cites.
M. Dorigo, V. Maniezzo, and A. Colorni, “Ant system: optimization by a colony of cooperating agents,” IEEE transactions on systems, man, and cybernetics. Part B, Cybernetics : a publication of the IEEE Systems, Man, and Cybernetics Society , vol. 26 1, pp. 29–41, 1996
1996
Cited alongside, same era.
V. G. Deineko and E. Çela, “The quadratic assignment problem: Theory and algorithms,” 1998
1998
Cited alongside, same era.
R. Carr, “Simulated annealing,” 2002
2002
Cited alongside, same era.
K. M. Anstreicher, “Recent advances in the solution of quadratic assignment problems,” Mathematical Programming , vol. 97, pp. 27–42, 2003
2003
Cited alongside, same era.
R. S. Sutton and A. G. Barto, “Reinforcement learning: An introduction,” IEEE Transactions on Neural Networks , vol. 16, pp. 285–286, 2005
2005
Cited alongside, same era.
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Processing Magazine , vol. 34, pp. 26–38, 2017
2017
Later among the works it cites.
2018
Later among the works it cites.
V. François-Lavet, P. Henderson, R. Islam, M. G. Bellemare, and J. Pineau, “An introduction to deep reinforcement learning,” Found. Trends Mach. Learn. , vol. 11, pp. 219–354, 2018
2018
Later among the works it cites.
W. Kool, H. van Hoof, and M. Welling, “Attention, learn to solve routing problems!” in International Conference on Learning Representations , 2018
2018
Later among the works it cites.
S. Tanwar, “Bellman equation and dynamic programming,” 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. van Otterlo and M. A. Wiering, “Reinforcement learning and markov decision processes,” in Reinforcement Learning , 2012
2012
Cited alongside, same era.
S. H. Benton, “The hamilton-jacobi equation: A global approach,” 2012
2012
Cited alongside, same era.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Cited alongside, same era.
J. Graves, “Understanding rl: The bellman equations,” 2017
2017
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Mazyavkina, S. V. Sviridov, S. Ivanov, and E. Burnaev, “Reinforcement learning for combinatorial optimization: A survey,” Comput. Oper. Res. , vol. 134, p. 105400, 2020
2020
Closest in time.
C. L. Bajaj, Y. Wang, and Y. Yang, “Reinforcement learning of self enhancing camera image and signal processing,” 2021
2021
Closest in time.
Y. Yang, Z. Xue, and A. Whinston, “Self-enhancing multi-filter sequence-to-sequence model,” Procedia Computer Science , vol. 215, pp. 537–545, 2022
2022
Closest in time.