S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation
1997
Earlier work this paper cites.
H. Rahmandad, N. Repenning, and J. Sterman, “Effects of feedback delay on learning,” System Dynamics Review
2009
Earlier work this paper cites.
PhD thesis, University of Alberta, 2011
H. R. Maei, Gradient temporal-difference learning algorithms · 2011
Earlier work this paper cites.
M. Valko, N. Korda, R. Munos, I. Flaounas, and N. Cristianini, “Finite-time analysis of kernelised contextual bandits,” arXiv preprint arXiv:1309.6869
Original
2013
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
J. A. Tropp, “An introduction to matrix concentration inequalities,” arXiv preprint arXiv:1501.01571
Original
2015
Earlier work this paper cites.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research
2016
Earlier work this paper cites.
S. Shalev-Shwartz, S. Shammah, and A. Shashua, “Safe, multi-agent, reinforcement learning for autonomous driving,” arXiv preprint arXiv:1610.03295
Original
2016
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al
2017
Earlier work this paper cites.
M. Olivecrona, T. Blaschke, O. Engkvist, and H. Chen, “Molecular de-novo design through deep reinforcement learning,” Journal of Cheminformatics
2017
Earlier work this paper cites.
D. Hein, S. Depeweg, M. Tokic, S. Udluft, A. Hentschel, T. A. Runkler, and V. Sterzing, “A benchmark environment motivated by industrial control problems,” in Proc. Symposium Series on Computational Intelligence (SSCI)
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. Advances in Neural Information Processing Systems (NeurIPS)
2017
Earlier work this paper cites.
S. R. Chowdhury and A. Gopalan, “On kernelized multi-armed bandits,” in Proc. International Conference on Machine Learning (ICML)
2017
Earlier work this paper cites.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 2018
Earlier work this paper cites.
K. Lin, R. Zhao, Z. Xu, and J. Zhou, “Efficient large-scale fleet management via multi-agent deep reinforcement learning,” in Proc. International Conference on Knowledge Discovery & Data Mining
2018
Earlier work this paper cites.
Z. Zheng, J. Oh, and S. Singh, “On learning intrinsic rewards for policy gradient methods,” Proc. Advances in Neural Information Processing Systems (NeurIPS)
2018
Earlier work this paper cites.
J. Oh, Y. Guo, S. Singh, and H. Lee, “Self-imitation learning,” in Proc. International Conference on Machine Learning (ICML)
2018
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: convergence and generalization in neural networks,” Proc. Advances in Neural Information Processing Systems (NeurIPS)
2018
Earlier work this paper cites.
Cambridge university press, 2018
R. Vershynin, High-dimensional probability: An introduction with applications in data science · 2018
Earlier work this paper cites.
MIT press, 2018
M. Mohri, A. Rostamizadeh, and A. Talwalkar, Foundations of machine learning · 2018
Earlier work this paper cites.
Y. Gong, M. Abdel-Aty, Q. Cai, and M. S. Rahman, “Decentralized network level adaptive signal control by multi-agent deep reinforcement learning,” Transportation Research Interdisciplinary Perspectives
2019
Earlier work this paper cites.
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter, “Rudder: return decomposition for delayed rewards,” Proc. Advances in Neural Information Processing Systems (NeurIPS)
2019
Earlier work this paper cites.
Y. Liu, Y. Luo, Y. Zhong, X. Chen, Q. Liu, and J. Peng, “Sequence modeling of temporal credit assignment for episodic reinforcement learning,” arXiv preprint arXiv:1905.13420
Original
2019
Earlier work this paper cites.
O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi, “Guidelines for reinforcement learning in healthcare,” Nature medicine
2019
Earlier work this paper cites.