Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) often encounters delayed and sparse feedback in real-world applications, even with only episodic rewards.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
The StarCraft Multi-Agent Challenge
Samvelyan, M.; Rashid, T.; de Witt, C. S.; Farquhar, G.; Nardelli, N.; Rudner, T. G. J.; Hung, C.-M.; Torr, P. H. S.; Foerster, J.; and Whiteson, S. 2019 · 1902
Earlier work this paper cites.
Sequence modeling of temporal credit assignment for episodic reinforcement learning
Liu, Y.; Luo, Y.; Zhong, Y.; Chen, X.; Liu, Q.; and Peng, J. 2019 · 1905
Earlier work this paper cites.
Dynamic programming
Bellman, R. 1966 · 1966
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y.; Harada, D.; and Russell, S. 1999 · 1999
Earlier work this paper cites.
The information bottleneck method
Tishby, N.; Pereira, F. C.; and Bialek, W. 2000 · 2000
Earlier work this paper cites.
Pearson correlation coefficient
Cohen, I.; Huang, Y.; Chen, J.; Benesty, J.; Benesty, J.; Chen, J.; Huang, Y.; and Cohen, I. 2009 · 2009
Earlier work this paper cites.
Align-rudder: Learning from few demonstrations by reward redistribution
Patil, V. P.; Hofmarcher, M.; Dinu, M.-C.; Dorfer, M.; Blies, P. M.; Brandstetter, J.; Arjona-Medina, J. A.; and Hochreiter, S. 2020 · 2009
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y.; Pál, D.; and Szepesvári, C. 2011 · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S.; Modayil, J.; Delp, M.; Degris, T.; Pilarski, P. M.; White, A.; and Precup, D. 2011 · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
Deep Variational Information Bottleneck
Alemi, A. A.; Fischer, I.; Dillon, J. V.; and Murphy, K. 2017 · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R.; Wu, Y. I.; Tamar, A.; Harb, J.; Abbeel, O. P.; and Mordatch, I. 2017 · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D.; Agrawal, P.; Efros, A. A.; and Darrell, T. 2017 · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S.; Hoof, H.; and Meger, D. 2018 · 2018
Earlier work this paper cites.
Sparse attentive backtracking: Temporal credit assignment through reminding
Ke, N. R.; ALIAS PARTH GOYAL, A. G.; Bilaniuk, O.; Binas, J.; Mozer, M. C.; Pal, C.; and Bengio, Y. 2018 · 2018
Earlier work this paper cites.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T.; Samvelyan, M.; De Witt, C. S.; Farquhar, G.; Foerster, J.; and Whiteson, S. 2018 · 2018
Earlier work this paper cites.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A.; Gillhofer, M.; Widrich, M.; Unterthiner, T.; Brandstetter, J.; and Hochreiter, S. 2019 · 2019
Earlier work this paper cites.
Learning guidance rewards with trajectory-space smoothing
Gangwani, T.; Zhou, Y.; and Peng, J. 2020 · 2020
Earlier work this paper cites.
Learning to utilize shaping rewards: A new approach of reward shaping
Hu, Y.; Wang, W.; Jia, H.; Wang, Y.; Chen, Y.; Hao, J.; Wu, F.; and Fan, C. 2020 · 2020
Earlier work this paper cites.
Learning retrospective knowledge with reverse reinforcement learning
Zhang, S.; Veeriah, V.; and Whiteson, S. 2020 · 2020
Cited alongside, same era.
Reinforcement learning with trajectory feedback
Efroni, Y.; Merlis, N.; and Mannor, S. 2021 · 2021
Cited alongside, same era.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B. R.; Sobh, I.; Talpaert, V.; Mannion, P.; Al Sallab, A. A.; Yogamani, S.; and Pérez, P. 2021 · 2021
Cited alongside, same era.
Learning long-term reward redistribution via randomized return decomposition
Ren, Z.; Guo, R.; Zhou, Y.; and Peng, J. 2021 · 2021
Cited alongside, same era.
Pettingzoo: Gym for multi-agent reinforcement learning
Terry, J.; Black, B.; Grammel, N.; Jayakumar, M.; Hari, A.; Sullivan, R.; Santos, L. S.; Dieffendahl, C.; Horsch, C.; Perez-Vicente, R.; et al. 2021 · 2021
Cited alongside, same era.
Reward Design with Language Models
Kwon, M.; Xie, S. M.; Bullard, K.; and Sadigh, D. 2023 · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
Liang, J.; Huang, W.; Xia, F.; Xu, P.; Hausman, K.; Ichter, B.; Florence, P.; and Zeng, A. 2023 · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J.; Liang, W.; Wang, G.; Huang, D.-A.; Bastani, O.; Jayaraman, D.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2023 · 2023
Later among the works it cites.
Supported trust region optimization for offline reinforcement learning
Mao, Y.; Zhang, H.; Chen, C.; Xu, Y.; and Ji, X. 2023 · 2023
Later among the works it cites.
Self-driven Grounding: Large Language Model Agents with Automatical Language-aligned Skill Learning
Peng, S.; Hu, X.; Yi, Q.; Zhang, R.; Guo, J.; Huang, D.; Tian, Z.; Chen, R.; Du, Z.; Guo, Q.; et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Widrich, M.; Hofmarcher, M.; Patil, V. P.; Bitto-Nemling, A.; and Hochreiter, S. 2021 · 2021
Cited alongside, same era.
Episodic multi-agent reinforcement learning with curiosity-driven exploration
Zheng, L.; Chen, J.; Wang, J.; He, J.; Hu, Y.; Chen, Y.; Fan, C.; Gao, Y.; and Zhang, C. 2021 · 2021
Cited alongside, same era.
Off-policy reinforcement learning with delayed rewards
Han, B.; Ren, Z.; Wu, Z.; Zhou, Y.; and Peng, J. 2022 · 2022
Cited alongside, same era.
History compression via language models in reinforcement learning
Paischer, F.; Adler, T.; Patil, V.; Bitto-Nemling, A.; Holzleitner, M.; Lehner, S.; Eghbal-Zadeh, H.; and Hochreiter, S. 2022 · 2022
Cited alongside, same era.
Agent-time attention for sparse rewards multi-agent reinforcement learning
She, J.; Gupta, J. K.; and Kochenderfer, M. J. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
Agent-temporal attention for reward redistribution in episodic multi-agent reinforcement learning
Xiao, B.; Ramasubramanian, B.; and Poovendran, R. 2022 · 2022
Cited alongside, same era.
Hokoff: Real Game Dataset from Honor of Kings and its Offline Reinforcement Learning Benchmarks
Qu, Y.; Wang, B.; Shao, J.; Jiang, Y.; Chen, C.; Ye, Z.; Liu, L.; Feng, Y. J.; Lai, L.; Qin, H.; et al. 2023 · 2023
Later among the works it cites.
Complementary attention for multi-agent reinforcement learning
Shao, J.; Zhang, H.; Qu, Y.; Liu, C.; He, S.; Jiang, Y.; and Ji, X. 2023 · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K. R.; and Yao, S. 2023 · 2023
Later among the works it cites.
LgTS: Dynamic Task Sampling using LLM-generated sub-goals for Reinforcement Learning Agents
Shukla, Y.; Gao, W.; Sarathy, V.; Velasquez, A.; Wright, R.; and Sinapov, J. 2023 · 2023
Later among the works it cites.
Song, J.; Zhou, Z.; Liu, J.; Fang, C.; Shu, Z.; and Ma, L. 2023 · 2023
Later among the works it cites.
Subgoal Proposition Using a Vision-Language Model
Su, J.; and Zhang, Q. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Gymnasium
Towers, M.; Terry, J. K.; Kwiatkowski, A.; Balis, J. U.; Cola, G. d.; Deleu, T.; Goulão, M.; Kallinteris, A.; KG, A.; Krimmel, M.; Perez-Vicente, R.; Pierré, A.; Schulhoff, S.; Tai, J. J.; Shen, A. T. J.; and Younis, O. G. 2023 · 2023
Later among the works it cites.
Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods
Cao, Y.; Zhao, H.; Cheng, Y.; Shu, T.; Liu, G.; Liang, G.; Zhao, J.; and Li, Y. 2024 · 2024
Closest in time.
Episodic Return Decomposition by Difference of Implicitly Assigned Sub-trajectory Reward
Lin, H.; Wu, H.; Zhang, J.; Sun, Y.; Ye, J.; and Yu, Y. 2024 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
Choices are more important than efforts: Llm enables efficient multi-agent exploration
Qu, Y.; Wang, B.; Jiang, Y.; Shao, J.; Mao, Y.; Wang, C.; Liu, C.; and Ji, X. 2024 · 2024
Closest in time.
Counterfactual conservative Q learning for offline multi-agent reinforcement learning
Shao, J.; Qu, Y.; Chen, C.; Zhang, H.; and Ji, X. 2024 · 2024
Closest in time.
Unleashing the Power of Pre-trained Language Models for Offline Reinforcement Learning
Shi, R.; Liu, Y.; Ze, Y.; Du, S. S.; and Xu, H. 2024 · 2024
Closest in time.
LLM-Empowered State Representation for Reinforcement Learning
Wang, B.; Qu, Y.; Jiang, Y.; Shao, J.; Liu, C.; Yang, W.; and Ji, X. 2024 · 2024
Closest in time.