Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is challenging in the common case of delays between events and their sensory perceptions.
Lineare differentialgleichungen mit nacheilendem argument
Myshkis, A. D · 1955
Earlier work this paper cites.
Differential—difference equations
Cooke, K. L · 1963
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Sutton, R. S · 1984
Earlier work this paper cites.
Closed-loop control with delayed information
Altman, E. and Nain, P · 1992
Earlier work this paper cites.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G · 1994
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S · 1995
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Learning rates for q-learning
Even-Dar, E., Mansour, Y., and Bartlett, P · 2003
Earlier work this paper cites.
On reachable sets for linear systems with delay and bounded peak inputs
Fridman, E. and Shaked, U · 2003
Earlier work this paper cites.
Markov decision processes with delays and asynchronous cost collection
Katsikopoulos, K. V. and Engelbrecht, S. E · 2003
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Villani, C. et al · 2009
Earlier work this paper cites.
Learning and planning in environments with delayed feedback
Walsh, T. J., Nouri, A., Li, L., and Littman, M. L · 2009
Earlier work this paper cites.
On the locality of action domination in sequential decision making
Rachelson, E. and Lagoudakis, M. G · 2010
Earlier work this paper cites.
Control delay in reinforcement learning for real-time dynamic systems: A memoryless approach
Schuitema, E., Buşoniu, L., Babuška, R., and Jonker, P · 2010
Earlier work this paper cites.
Speedy q-learning
Azar, M. G., Munos, R., Ghavamzadeh, M., and Kappen, H · 2011
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I , volume 4
Bertsekas, D · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Low-latency trading
Hasbrouck, J. and Saar, G · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Control of a quadrotor with reinforcement learning
Hwangbo, J., Sa, I., Siegwart, R., and Hutter, M · 2017
Automatic p2p energy trading model based on reinforcement learning using long short-term delayed reward
Kim, J.-G. and Lee, B · 2020
Later among the works it cites.
Over-and underapproximating reach sets for perturbed delay differential equations
Xue, B., Wang, Q., Feng, S., and Zhan, N · 2020
Later among the works it cites.
Delay-aware model-based reinforcement learning for continuous control
Chen, B., Xu, M., Li, L., and Zhao, D · 2021
Later among the works it cites.
Acting in delayed environments with non-stationary markov policies
Derman, E., Dalal, G., and Mannor, S · 2021
Later among the works it cites.
Learning a belief representation for delayed reinforcement learning
Liotet, P., Venneri, E., and Restelli, M · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
At human speed: Deep reinforcement learning with action delay
Firoiu, V., Ju, T., and Tenenbaum, J · 2018
Cited alongside, same era.
Setting up a reinforcement learning task with a real-world robot
Mahmood, A. R., Korenkevych, D., Komer, B. J., and Bergstra, J · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Nath, S., Baranwal, M., and Khadilkar, H · 2021
Later among the works it cites.
Learning-based framework for sensor fault-tolerant building hvac control with model-assisted learning
Xu, S., Fu, Y., Wang, Y., O’Neill, Z., and Zhu, Q · 2021
Later among the works it cites.
Reach-avoid analysis for delay differential equations
Xue, B., Bai, Y., Zhan, N., Liu, W., and Jiao, L · 2021
Later among the works it cites.
Big data analytics in weather forecasting: A systematic review
Fathi, M., Haghi Kashani, M., Jameii, S. M., and Mahdipour, E · 2022
Later among the works it cites.
Off-policy reinforcement learning with delayed rewards
Han, B., Ren, Z., Wu, Z., Zhou, Y., and Peng, J · 2022
Later among the works it cites.
Delayed reinforcement learning by imitation
Liotet, P., Maran, D., Bisi, L., and Restelli, M · 2022
Later among the works it cites.
Accelerate online reinforcement learning for building hvac control with heterogeneous expert guidances
Xu, S., Fu, Y., Wang, Y., Yang, Z., O’Neill, Z., Wang, Z., and Zhu, Q · 2022
Later among the works it cites.
Belief projection-based reinforcement learning for environments with delayed feedback
Kim, J., Kim, H., Kang, J., Baek, J., and Han, S · 2023
Later among the works it cites.
Joint differentiable optimization and verification for certified reinforcement learning
Wang, Y., Zhan, S., Wang, Z., Huang, C., Wang, Z., Yang, Z., and Zhu, Q · 2023
Later among the works it cites.
Highway reinforcement learning
Wang, Y., Strupl, M., Faccio, F., Wu, Q., Liu, H., Grudzień, M., Tan, X., and Schmidhuber, J · 2024
Closest in time.
State-wise safe reinforcement learning with pixel observations
Zhan, S. S., Wang, Y., Wu, Q., Jiao, R., Huang, C., and Zhu, Q · 2024
Closest in time.