Fetching the paper…
Reading the bibliography…
The standard Markov Decision Process (MDP) formulation hinges on the assumption that an action is executed immediately after it was chosen.
Dynamic programming and Markov processes
Ronald A Howard · 1960
Earlier work this paper cites.
Explicit solution of inventory problems with delivery lags
Avner Bar-Ilan and Agnès Sulem · 1995
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas, Dimitri P Bertsekas, Dimitri P Bertsekas, and Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Stability and control of time-delay systems , volume 228
Luc Dugard and Erik I Verriest · 1998
Earlier work this paper cites.
Markov decision processes with delays and asynchronous cost collection
Konstantinos V Katsikopoulos and Sascha E Engelbrecht · 2003
Earlier work this paper cites.
Time-delay systems: an overview of some recent advances and open problems
Jean-Pierre Richard · 2003
Earlier work this paper cites.
Delay-aware model-based reinforcement learning for continuous control
Baiming Chen, Mengdi Xu, Liang Li, and Ding Zhao · 2005
Earlier work this paper cites.
Delay-aware multi-agent reinforcement learning
Baiming Chen, Mengdi Xu, Zuxin Liu, Liang Li, and Ding Zhao · 2005
Earlier work this paper cites.
Impulse control problem on finite horizon with execution delay
Benjamin Bruder and Huyên Pham · 2009
Earlier work this paper cites.
Learning and planning in environments with delayed feedback
Thomas J Walsh, Ali Nouri, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Exponential lower bounds for policy iteration
John Fearnley · 2010
Earlier work this paper cites.
Lower bounds for howard’s algorithm for finding minimum mean-cost cycles
Thomas Dueholm Hansen and Uri Zwick · 2010
Cited alongside, same era.
The complexity of policy iteration is exponential for discounted markov decision processes
Romain Hollanders, Jean-Charles Delvenne, and Raphaël M Jungers · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Texplore: real-time sample-efficient reinforcement learning for robots
Todd Hester and Peter Stone · 2013
Cited alongside, same era.
Online learning under delayed feedback
Pooria Joulani, Andras Gyorgy, and Csaba Szepesvári · 2013
Cited alongside, same era.
Introduction to time-delay systems: Analysis and control
Emilia Fridman · 2014
Bandits with delayed anonymous feedback
Ciara Pike-Burke, Shipra Agrawal, Csaba Szepesvari, and Steffen Grünewälder · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Later among the works it cites.
At human speed: Deep reinforcement learning with action delay
Vlad Firoiu, Tina Ju, and Josh Tenenbaum · 2018
Later among the works it cites.
Challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Daniel Mankowitz, and Todd Hester · 2019
Later among the works it cites.
26ms inference time for resnet-50: Towards real-time execution of all dnns on smartphone
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Multiple model q-learning for stochastic asynchronous rewards
Jeffrey S Campbell, Sidney N Givigi, and Howard M Schwartz · 2016
Cited alongside, same era.
Improved and generalized upper bounds on the complexity of policy iteration
Bruno Scherrer et al · 2016
Cited alongside, same era.
Wei Niu, Xiaolong Ma, Yanzhi Wang, and Bin Ren · 2019
Later among the works it cites.
Real-time reinforcement learning
Simon Ramstedt and Chris Pal · 2019
Later among the works it cites.
Characterizing perception module performance and robustness in production-scale autonomous driving system
Alessandro Toschi, Mustafa Sanic, Jingwen Leng, Quan Chen, Chunlin Wang, and Minyi Guo · 2019
Later among the works it cites.
Towards safety-aware computing system design in autonomous vehicles
Hengyu Zhao, Yubo Zhang, Pingfan Meng, Hui Shi, Li Erran Li, Tiancheng Lou, and Jishen Zhao · 2019
Later among the works it cites.
Learning to simulate dynamic environments with gamegan
Seung Wook Kim, Yuhao Zhou, Jonah Philion, Antonio Torralba, and Sanja Fidler · 2020
Later among the works it cites.
Thinking while moving: Deep reinforcement learning with concurrent control
Ted Xiao, Eric Jang, Dmitry Kalashnikov, Sergey Levine, Julian Ibarz, Karol Hausman, and Alexander Herzog · 2020
Later among the works it cites.