Fetching the paper…
Reading the bibliography…
With extensive pre-trained knowledge and high-level general capabilities, large language models (LLMs) emerge as a promising avenue to augment reinforcement learning (RL) in aspects such as multi-task learning, sample efficiency, and high-level task planning.
1901
Earlier work this paper cites.
1910
Earlier work this paper cites.
2005
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems , vol. 25, 2012
2012
Earlier work this paper cites.
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez et al. , “A survey of monte carlo tree search methods,” IEEE Transactions on Computational Intelligence and AI in games , vol. 4, no. 1, pp. 1–43, 2012
2012
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , vol. 16, no. 1, pp. 1437–1480, 2015
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou et al. , “Mastering the game of go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. Dragan, “Inverse reward design,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Neural Information Processing Systems , 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:13756489
2017
Earlier work this paper cites.
J. Andreas, D. Klein, and S. Levine, “Modular multitask reinforcement learning with policy sketches,” in International conference on machine learning . PMLR, 2017, pp. 166–175
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning . PMLR, 2018, pp. 1861–1870
2018
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould et al. , “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford and K. Narasimhan, “Improving language understanding by generative pre-training,” 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:49313245
2018
Earlier work this paper cites.
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly et al. , “Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” Nature , vol. 575, no. 7782, pp. 350–354, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Jiang, S. S. Gu, K. P. Murphy, and C. Finn, “Language as an abstraction for hierarchical deep reinforcement learning,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Luketina, N. Nardelli, G. Farquhar, J. Foerster, J. Andreas, E. Grefenstette, S. Whiteson, and T. Rocktäschel, “A Survey of Reinforcement Learning Informed by Natural Language,” Jun. 2019
2019
Earlier work this paper cites.
R. Yang, X. Sun, and K. Narasimhan, “A generalized algorithm for multi-objective reinforcement learning and policy adaptation,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in International conference on machine learning . PMLR, 2019, pp. 2555–2565
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
N. Brown, A. Lerer, S. Gross, and T. Sandholm, “Superhuman ai for multiplayer poker,” Science , vol. 365, no. 6456, pp. 885–890, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
K. Wang, B. Kang, J. Shao, and J. Feng, “Improving generalization in reinforcement learning with mixture regularization,” Advances in Neural Information Processing Systems , vol. 33, pp. 7968–7978, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Yang, W. Huang, W. Tu, Q. Qu, Y. Shen, and K. Lei, “Multitask learning and reinforcement learning for personalized dialog generation: An empirical study,” IEEE transactions on neural networks and learning systems , vol. 32, no. 1, pp. 49–62, 2020
2020
Earlier work this paper cites.
A. Srinivas, M. Laskin, and P. Abbeel, “CURL: Contrastive Unsupervised Representations for Reinforcement Learning,” Sep. 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
R. Patel, E. Pavlick, and S. Tellex, “Grounding language to non-markovian tasks with no supervision of task specifications.” in Robotics: Science and Systems , vol. 2020, 2020
2020
Earlier work this paper cites.
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra, “Improving vision-and-language navigation with image-text pairs from the web,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16 . Springer, 2020, pp. 259–274
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Hausknecht, P. Ammanabrolu, M.-A. Côté, and X. Yuan, “Interactive fiction games: A colossal adventure,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 7903–7910
2020
Earlier work this paper cites.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to Control: Learning Behaviors by Latent Imagination,” Mar. 2020
2020
Earlier work this paper cites.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” 2020
2020
Earlier work this paper cites.
N. B. Schmid, A. Botev, A. Hennig, A. Lerer, Q. Wu, D. Yarats, J. Foerster, T. Rocktäschel et al. , “Rebel: A general game playing ai,” Science , vol. 373, no. 6556, pp. 664–670, 2021
2021
Earlier work this paper cites.
A. Stooke, K. Lee, P. Abbeel, and M. Laskin, “Decoupling representation learning from reinforcement learning,” in International Conference on Machine Learning . PMLR, 2021, pp. 9870–9879
2021
Earlier work this paper cites.
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese et al. , “What matters in learning from offline human demonstrations for robot manipulation,” 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Rao, Y. Wu, Z. Yang, W. Zhang, S. Lu, W. Lu, and Z. Zha, “Visual navigation with multiple goals based on deep reinforcement learning,” IEEE Transactions on Neural Networks and Learning Systems , vol. 32, no. 12, pp. 5445–5455, 2021
2021
Earlier work this paper cites.
C. Huang, R. Zhang, M. Ouyang, P. Wei, J. Lin, J. Su, and L. Lin, “Deductive reinforcement learning for visual autonomous urban driving navigation,” IEEE Transactions on Neural Networks and Learning Systems , vol. 32, no. 12, pp. 5379–5391, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
T. R. Sumers, M. K. Ho, R. D. Hawkins, K. Narasimhan, and T. L. Griffiths, “Learning rewards from linguistic feedback,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 7, 2021, pp. 6002–6010
2021
Earlier work this paper cites.
J. Eschmann, “Reward function design in reinforcement learning,” Reinforcement Learning Algorithms: Analysis and Applications , pp. 25–33, 2021
2021
Earlier work this paper cites.
S. Mirchandani, S. Karamcheti, and D. Sadigh, “Ella: Exploration through learned language abstraction,” Advances in Neural Information Processing Systems , vol. 34, pp. 29 529–29 540, 2021
2021
Earlier work this paper cites.
M. Janner, Q. Li, and S. Levine, “Offline reinforcement learning as one big sequence modeling problem,” Advances in neural information processing systems , vol. 34, pp. 1273–1286, 2021
2021
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” 2021
2021
Earlier work this paper cites.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas et al. , “Decision transformer: Reinforcement learning via sequence modeling,” Advances in neural information processing systems , vol. 34, pp. 15 084–15 097, 2021
2021
Earlier work this paper cites.
C. Yu, J. Liu, S. Nemati, and G. Yin, “Reinforcement learning in healthcare: A survey,” ACM Computing Surveys (CSUR) , vol. 55, no. 1, pp. 1–36, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Zhao, Z. Liu, X. Mai, J. Zhao, J. Qiu, G. Liu, Z. Y. Dong, and A. M. Ghias, “Mobile battery energy storage system control with knowledge-assisted deep reinforcement learning,” Energy Conversion and Economics , vol. 3, no. 6, pp. 381–391, 2022
2022
Earlier work this paper cites.
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence, “Interactive language: Talking to robots in real time,” 2022
2022
Earlier work this paper cites.
Z. Yang, K. Ren, X. Luo, M. Liu, W. Liu, J. Bian, W. Zhang, and D. Li, “Towards applicable reinforcement learning: Improving the generalization and sample efficiency with policy ensemble,” in International Joint Conference on Artificial Intelligence , 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:248887230
2022
Earlier work this paper cites.
F. Dworschak, S. Dietze, M. Wittmann, B. Schleich, and S. Wartzack, “Reinforcement learning for engineering design automation,” Advanced Engineering Informatics , vol. 52, p. 101612, 2022
2022
Earlier work this paper cites.
L. L. Di Langosco, J. Koch, L. D. Sharkey, J. Pfau, and D. Krueger, “Goal misgeneralization in deep reinforcement learning,” in International Conference on Machine Learning . PMLR, 2022, pp. 12 004–12 019
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei et al. , “Learning to summarize from human feedback,” 2022
2022
Cited alongside, same era.
R. Devidze, P. Kamalaruban, and A. Singla, “Exploration-guided reward shaping for reinforcement learning under sparse rewards,” Advances in Neural Information Processing Systems , vol. 35, pp. 5829–5842, 2022
2022
Cited alongside, same era.
L. Sun, H. Zhang, W. Xu, and M. Tomizuka, “Paco: Parameter-compositional multi-task reinforcement learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 21 495–21 507, 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
L. Gao, J. Schulman, and J. Hilton, “Scaling laws for reward model overoptimization,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 23–29 Jul 2023, pp. 10 835–10 866
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma et al. , “Emergent abilities of large language models,” Trans. Mach. Learn. Res. , vol. 2022, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:249674500
2022
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal et al. , “Training language models to follow instructions with human feedback,” 2022
2022
Cited alongside, same era.
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu et al. , “Do As I Can, Not As I Say: Grounding Language in Robotic Affordances,” Aug. 2022
2022
Cited alongside, same era.
M. Kwon, S. M. Xie, K. Bullard, and D. Sadigh, “Reward Design with Language Models,” in The Eleventh International Conference on Learning Representations , Sep. 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang et al. , “Ego4d: Around the world in 3,000 hours of egocentric video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 995–19 012
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart et al. , “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” in Conference on Robot Learning . PMLR, 2023, pp. 2165–2183
2023
Later among the works it cites.
H. Hu and D. Sadigh, “Language Instructed Reinforcement Learning for Human-AI Coordination,” Jun. 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Liu, Z. Guo, Y. Yao, Z. Cen, W. Yu, T. Zhang, and D. Zhao, “Constrained decision transformer for offline safe reinforcement learning,” in International Conference on Machine Learning . PMLR, 2023, pp. 21 611–21 630
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language Agents with Verbal Reinforcement Learning,” Oct. 2023
2023
Later among the works it cites.
J. Robine, M. Höftmann, T. Uelwer, and S. Harmeling, “Transformer-based World Models Are Happy With 100k Interactions,” Mar. 2023
2023
Later among the works it cites.
R. P. K. Poudel, H. Pandya, C. Zhang, and R. Cipolla, “LanGWM: Language Grounded World Model,” Nov. 2023
2023
Later among the works it cites.
D. Das, S. Chernova, and B. Kim, “State2explanation: Concept-based explanations to benefit agent learning and user understanding,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023. [Online]. Available: https://openreview.net/forum?id=xGz0wAIJrS
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Lu, X. Zhao, S. Magg, M. Gromniak, M. Li, and S. Wermterl, “A closer look at reward decomposition for high-level robotic explanations,” in 2023 IEEE International Conference on Development and Learning (ICDL) . IEEE, 2023, pp. 429–436
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
O. Slumbers, D. H. Mguni, K. Shao, and J. Wang, “Leveraging large language models for optimised coordination in textual multi-agent reinforcement learning,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
R. S. Y. C. Tan, Q. Lin, G. H. Low, R. Lin, T. C. Goh, C. C. E. Chang, F. F. Lee, W. Y. Chan et al. , “Inferring cancer disease response from radiology reports using large language models with data augmentation and prompting,” Journal of the American Medical Informatics Association , vol. 30, no. 10, pp. 1657–1664, 2023
2023
Later among the works it cites.
Q. Gao, C. Zhao, Y. Sun, T. Xi, G. Zhang, B. Ghanem, and J. Zhang, “A unified continual learning framework with general parameter-efficient tuning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 483–11 493
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Jiang, Z. Ma, L. Zhang, and J. Chen, “Eplus-llm: A large language model-based computing platform for automated building energy modeling,” Applied Energy , vol. 367, p. 123431, 2024
2024
Closest in time.
H. Tan, Z. Guo, Z. Lin, Y. Chen, D. Huang, W. Yuan, H. Zhang, and J. Yan, “General generative ai-based image augmentation method for robust rooftop pv segmentation,” Applied Energy , vol. 368, p. 123554, 2024
2024
Closest in time.
2024
Closest in time.
M. Cho, J. Park, S. Lee, and Y. Sung, “Hard tasks first: Multi-task reinforcement learning through task scheduling,” in Forty-first International Conference on Machine Learning , 2024. [Online]. Available: https://openreview.net/forum?id=haUOhXo70o
2024
Closest in time.
2024
Closest in time.
G. Zhang, A. Jain, I. Hwang, S.-H. Sun, and J. J. Lim, “Efficient multi-task reinforcement learning via selective behavior sharing,” 2024. [Online]. Available: https://openreview.net/forum?id=LYGHdwyXUb
2024
Closest in time.
2024
Closest in time.
F. Paischer, T. Adler, M. Hofmarcher, and S. Hochreiter, “Semantic helm: A human-readable memory for reinforcement learning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
W. K. Kim, S. Kim, H. Woo et al. , “Efficient policy adaptation with contrastive prompt ensemble for embodied agents,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
R. P. Poudel, H. Pandya, S. Liwicki, and R. Cipolla, “Recore: Regularized contrastive representation learning of world model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 22 904–22 913
2024
Closest in time.
2024
Closest in time.
B. A. Spiegel, Z. Yang, W. Jurayj, B. Bachmann, S. Tellex, and G. Konidaris, “Informing reinforcement learning agents by grounding language to markov decision processes,” in Workshop on Training Agents with Foundation Models at RLC 2024 , 2024. [Online]. Available: https://openreview.net/forum?id=uFm9e4Ly26
2024
Closest in time.
Y. Wang, J. He, D. Wang, Q. Wang, B. Wan, and X. Luo, “Multimodal transformer with adaptive modality weighting for multimodal sentiment analysis,” Neurocomputing , vol. 572, p. 127181, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Xie, S. Zhao, C. H. Wu, Y. Liu, Q. Luo, V. Zhong, Y. Yang, and T. Yu, “Text2reward: Reward shaping with language models for reinforcement learning,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=tUM39YTRxH
2024
Closest in time.
I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Dernoncourt, T. Yu, R. Zhang et al. , “Bias and fairness in large language models: A survey,” Computational Linguistics , pp. 1–79, 2024
2024
Closest in time.
2024
Closest in time.
Z. Zeng, C. Zhang, S. Wang, and C. Sun, “Goal-conditioned predictive coding for offline reinforcement learning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
M. Dalal, T. Chiruvolu, D. S. Chaplot, and R. Salakhutdinov, “Plan-seq-learn: Language model guided RL for solving long horizon robotics tasks,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=hQVCCxQrYN
2024
Closest in time.
W. Huang, H. Liu, Z. Huang, and C. Lv, “Safety-aware human-in-the-loop reinforcement learning with shared control for autonomous driving,” IEEE Transactions on Intelligent Transportation Systems , pp. 1–12, 2024
2024
Closest in time.
2024
Closest in time.
J. Lee, A. Xie, A. Pacchiano, Y. Chandak, C. Finn, O. Nachum, and E. Brunskill, “Supervised pretraining can learn in-context reinforcement learning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
L. Wang, X. Zhang, H. Su, and J. Zhu, “A comprehensive survey of continual learning: theory, method and application,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Closest in time.
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, p. 186345, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
B. Gordon, Y. Bitton, Y. Shafir, R. Garg, X. Chen, D. Lischinski, D. Cohen-Or, and I. Szpektor, “Mismatch quest: Visual and textual feedback for image-text misalignment,” in European Conference on Computer Vision . Springer, 2025, pp. 310–328
2025
Closest in time.
N. Vithayathil Varghese and Q. H. Mahmoud, “A survey of multi-task deep reinforcement learning,” Electronics , vol. 9, no. 9, 2020. [Online]. Available: https://www.mdpi.com/2079-9292/9/9/1363
2079
Closest in time.