Fetching the paper…
Reading the bibliography…
Action and observation delays exist prevalently in the real-world cyber-physical systems which may pose challenges in reinforcement learning design.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International conference on machine learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
Z. Artstein, “Linear systems with delayed controls: A reduction,” IEEE Transactions on Automatic control , vol. 27, no. 4, pp. 869–879, 1982
1982
Earlier work this paper cites.
M. Tan, “Multi-agent reinforcement learning: Independent vs. cooperative agents,” in Proceedings of the tenth international conference on machine learning , 1993, pp. 330–337
1993
Earlier work this paper cites.
K. J. Astrom, C. C. Hang, and B. Lim, “A new smith predictor for controlling a process with an integrator and long dead-time,” IEEE transactions on Automatic Control , vol. 39, no. 2, pp. 343–345, 1994
1994
Earlier work this paper cites.
S. P. Singh, T. Jaakkola, and M. I. Jordan, “Learning without state-estimation in partially observable markovian decision processes,” in Machine Learning Proceedings 1994 . Elsevier, 1994, pp. 284–292
1994
Earlier work this paper cites.
M. R. Matausek and A. Micic, “On the modified smith predictor for controlling a process with an integrator and long dead-time,” IEEE Transactions on Automatic Control , vol. 44, no. 8, pp. 1603–1606, 1999
1999
Earlier work this paper cites.
L. Mirkin, “On the extraction of dead-time controllers from delay-free parametrizations,” IFAC Proceedings Volumes , vol. 33, no. 23, pp. 169–174, 2000
2000
Earlier work this paper cites.
S.-I. Niculescu, Delay effects on stability: a robust control approach . Springer Science & Business Media, 2001, vol. 269
2001
Earlier work this paper cites.
K. Gu and S.-I. Niculescu, “Survey on recent results in the stability and control of time-delay systems,” Journal of dynamic systems, measurement, and control , vol. 125, no. 2, pp. 158–165, 2003
2003
Earlier work this paper cites.
K. V. Katsikopoulos and S. E. Engelbrecht, “Markov decision processes with delays and asynchronous cost collection,” IEEE transactions on automatic control , vol. 48, no. 4, pp. 568–574, 2003
2003
Earlier work this paper cites.
A. K. Agogino and K. Tumer, “Unifying temporal and structural credit assignment problems,” in Proceedings of the Third International Joint Conference on Autonomous Agents and Multiagent Systems-Volume 2 . IEEE Computer Society, 2004, pp. 980–987
2004
Earlier work this paper cites.
M. Lazarević, “Finite time stability analysis of pd α \alpha fractional control of robotic time-delay systems,” Mechanics research communications , vol. 33, no. 2, pp. 269–279, 2006
2006
Earlier work this paper cites.
S. Biswas, R. Tatchikou, and F. Dion, “Vehicle-to-vehicle wireless communication protocols for enhancing highway traffic safety,” IEEE communications magazine , vol. 44, no. 1, pp. 74–82, 2006
2006
Earlier work this paper cites.
S. Ammoun, F. Nashashibi, and C. Laurgeau, “Real-time crash avoidance system on crossroads based on 802.11 devices and gps receivers,” in 2006 IEEE Intelligent Transportation Systems Conference . IEEE, 2006, pp. 1023–1028
2006
Earlier work this paper cites.
L. Bu, R. Babu, B. De Schutter et al. , “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) , vol. 38, no. 2, pp. 156–172, 2008
2008
Earlier work this paper cites.
E. Moulay, M. Dambrine, N. Yeganefar, and W. Perruquetti, “Finite-time stability and stabilization of time-delay systems,” Systems & Control Letters , vol. 57, no. 7, pp. 561–566, 2008
2008
Cited alongside, same era.
F. P. Bayan, A. D. Cornetto, A. Dunn, and E. Sauer, “Brake timing measurements for a tractor-semitrailer under emergency braking,” SAE International Journal of Commercial Vehicles , vol. 2, no. 2009-01-2918, pp. 245–255, 2009
2009
Cited alongside, same era.
T. J. Walsh, A. Nouri, L. Li, and M. L. Littman, “Learning and planning in environments with delayed feedback,” Autonomous Agents and Multi-Agent Systems , vol. 18, no. 1, p. 83, 2009
2009
Cited alongside, same era.
E. Schuitema, L. Buşoniu, R. Babuška, and P. Jonker, “Control delay in reinforcement learning for real-time dynamic systems: a memoryless approach,” in 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2010, pp. 3226–3231
2010
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
R. Lowe, Y. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in Advances in neural information processing systems , 2017, pp. 6379–6390
2017
Later among the works it cites.
I. Mordatch and P. Abbeel, “Emergence of grounded compositional language in multi-agent populations,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Rajamani, Vehicle dynamics and control . Springer Science & Business Media, 2011
2011
Cited alongside, same era.
L. Matignon, L. Jeanpierre, and A.-I. Mouaddib, “Coordinated multi-robot exploration under communication constraints using decentralized markov decision processes,” in Twenty-sixth AAAI conference on artificial intelligence , 2012
2012
Cited alongside, same era.
L. Matignon, G. J. Laurent, and N. Le Fort-Piat, “Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems,” The Knowledge Engineering Review , vol. 27, no. 1, pp. 1–31, 2012
2012
Cited alongside, same era.
2013
Cited alongside, same era.
2015
Cited alongside, same era.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, p. 484, 2016
2016
Cited alongside, same era.
J. Foerster, I. A. Assael, N. De Freitas, and S. Whiteson, “Learning to communicate with deep multi-agent reinforcement learning,” in Advances in neural information processing systems , 2016, pp. 2137–2145
2016
Cited alongside, same era.
S. Sukhbaatar, R. Fergus et al. , “Learning multiagent communication with backpropagation,” in Advances in neural information processing systems , 2016, pp. 2244–2252
2016
Cited alongside, same era.
R. Hannah and W. Yin, “On unbounded delays in asynchronous parallel fixed-point algorithms,” Journal of Scientific Computing , vol. 76, no. 1, pp. 299–326, 2018
2018
Later among the works it cites.
M. Wang, S. P. Hoogendoorn, W. Daamen, B. van Arem, B. Shyrokau, and R. Happee, “Delay-compensating strategy to enhance string stability of adaptive cruise controlled vehicles,” Transportmetrica B: Transport Dynamics , vol. 6, no. 3, pp. 211–229, 2018
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
E. Leurent, “An environment for autonomous driving decision-making,” https://github.com/eleurent/highway-env , 2018
2018
Later among the works it cites.
P. Hernandez-Leal, B. Kartal, and M. E. Taylor, “A survey and critique of multiagent deep reinforcement learning,” Autonomous Agents and Multi-Agent Systems , vol. 33, no. 6, pp. 750–797, 2019
2019
Later among the works it cites.
S. Ramstedt and C. Pal, “Real-time reinforcement learning,” in Advances in Neural Information Processing Systems , 2019, pp. 3067–3076
2019
Later among the works it cites.
2019
Later among the works it cites.
B. Chen, M. Xu, L. Li, and D. Zhao, “Delay-aware model-based reinforcement learning for continuous control,” 2020
2020
Closest in time.