Fetching the paper…
Reading the bibliography…
In environments with delayed observation, state augmentation by including actions within the delay window is adopted to retrieve Markovian property to enable reinforcement learning (RL).
Closed-loop control with delayed information
E. Altman and P. Nain · 1992
Earlier work this paper cites.
Markov decision processes with delays and asynchronous cost collection
K. V. Katsikopoulos and S. E. Engelbrecht · 2003
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
M. Toussaint · 2009
Earlier work this paper cites.
Learning and planning in environments with delayed feedback
T. J. Walsh, A. Nouri, L. Li, and M. L. Littman · 2009
Earlier work this paper cites.
Control delay in reinforcement learning for real-time dynamic systems: A memoryless approach
E. Schuitema, L. Buşoniu, R. Babuška, and P. Jonker · 2010
Earlier work this paper cites.
Variational inference for policy search in changing situations
G. Neumann · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
M. Gheshlaghi Azar, R. Munos, and H. J. Kappen · 2013
Earlier work this paper cites.
Low-latency trading
J. Hasbrouck and G. Saar · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
N. Tishby and N. Zaslavsky · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2017
Earlier work this paper cites.
Control of a quadrotor with reinforcement learning
J. Hwangbo, I. Sa, R. Siegwart, and M. Hutter · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Relative entropy regularized policy iteration
A. Abdolmaleki, J. T. Springenberg, J. Degrave, S. Bohez, Y. Tassa, D. Belov, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
At human speed: Deep reinforcement learning with action delay
V. Firoiu, T. Ju, and J. Tenenbaum · 2018
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Later among the works it cites.
Learning a belief representation for delayed reinforcement learning
P. Liotet, E. Venneri, and M. Restelli · 2021
Later among the works it cites.
Revisiting state augmentation methods for reinforcement learning with stochastic delays
S. Nath, M. Baranwal, and H. Khadilkar · 2021
Later among the works it cites.
Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms
S. Huang, R. F. J. Dossa, C. Ye, J. Braga, D. Chakraborty, K. Mehta, and J. G. AraÚjo · 2022
Later among the works it cites.
Delayed reinforcement learning by imitation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Cited alongside, same era.
Setting up a reinforcement learning task with a real-world robot
A. R. Mahmood, D. Korenkevych, B. J. Komer, and J. Bergstra · 2018
Cited alongside, same era.
Reinforcement learning: Theory and algorithms
A. Agarwal, N. Jiang, and S. M. Kakade · 2019
Cited alongside, same era.
Virel: A variational inference framework for reinforcement learning
M. Fellows, A. Mahajan, T. G. Rudner, and S. Whiteson · 2019
Cited alongside, same era.
Reinforcement learning with random delays
Y. Bouteiller, S. Ramstedt, G. Beltrame, C. Pal, and J. Binas · 2020
Cited alongside, same era.
Using reinforcement learning to minimize the probability of delay occurrence in transportation
Z. Cao, H. Guo, W. Song, K. Gao, Z. Chen, L. Zhang, and X. Zhang · 2020
Cited alongside, same era.
P. Liotet, D. Maran, L. Bisi, and M. Restelli · 2022
Later among the works it cites.
Constrained variational policy optimization for safe reinforcement learning
Z. Liu, Z. Cen, V. Isenbaev, W. Liu, S. Wu, B. Li, and D. Zhao · 2022
Later among the works it cites.
Belief projection-based reinforcement learning for environments with delayed feedback
J. Kim, H. Kim, J. Kang, J. Baek, and S. Han · 2023
Later among the works it cites.
Delays in reinforcement learning
P. Liotet · 2023
Later among the works it cites.
Addressing signal delay in deep reinforcement learning
W. Wang, D. Han, X. Luo, and D. Li · 2023
Later among the works it cites.
Joint differentiable optimization and verification for certified reinforcement learning
Y. Wang, S. Zhan, Z. Wang, C. Huang, Z. Wang, Z. Yang, and Q. Zhu · 2023
Later among the works it cites.
Enforcing hard constraints with soft barriers: Safe reinforcement learning in unknown stochastic environments
Y. Wang, S. S. Zhan, R. Jiao, Z. Wang, W. Jin, Z. Yang, Z. Wang, C. Huang, and Q. Zhu · 2023
Later among the works it cites.
State-wise safe reinforcement learning with pixel observations
S. S. Zhan, Y. Wang, Q. Wu, R. Jiao, C. Huang, and Q. Zhu · 2023
Later among the works it cites.
Boosting long-delayed reinforcement learning with auxiliary short-delayed task
Q. Wu, S. S. Zhan, Y. Wang, C.-W. Lin, C. Lv, Q. Zhu, and C. Huang · 2024
Closest in time.