Fetching the paper…
Reading the bibliography…
Designing effective reward functions in multi-agent reinforcement learning (MARL) is a significant challenge, often leading to suboptimal or misaligned behaviors in complex, coordinated environments.
Sulla determinazione empirica di una legge di distribuzione
Kolmogorov, A. N · 1933
Earlier work this paper cites.
Markov decision processes
Puterman, M. L · 1990
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al · 2000
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L., Babuska, R., and De Schutter, B · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
Knox, W. B. and Stone, P · 2009
Earlier work this paper cites.
Where do rewards come from
Singh, S., Lewis, R. L., and Barto, A. G · 2009
Earlier work this paper cites.
Probability: Theory and Examples
Durrett, R · 2010
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning
Mannion, P., Devlin, S., Duggan, J., and Howley, E · 2018
Earlier work this paper cites.
Credit assignment for collective multiagent rl with global rewards
Nguyen, D. T., Kumar, A., and Lau, H. C · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Earlier work this paper cites.
Liir: Learning individual intrinsic reward in multi-agent reinforcement learning
Du, Y., Han, L., Fang, M., Liu, J., Dai, T., and Tao, D · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Is independent learning all you need in the starcraft multi-agent challenge?
De Witt, C. S., Gupta, T., Makoviichuk, D., Makoviychuk, V., Torr, P. H., Sun, M., and Whiteson, S · 2020
Earlier work this paper cites.
Google research football: A novel reinforcement learning environment
Kurach, K., Raichuk, A., Stańczyk, P., Zajac, M., Bachem, O., Espeholt, L., Riquelme, C., Vincent, D., Michalski, M., Bousquet, O., et al · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Multi-agent collaboration via reward attribution decomposition
Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J. E., and Tian, Y · 2020
Cited alongside, same era.
Learning implicit credit assignment for cooperative multi-agent reinforcement learning
Zhou, M., Liu, Z., Sui, P., Li, Y., and Chung, Y. Y · 2020
Cited alongside, same era.
Multi-agent reinforcement learning: A review of challenges and applications
Canese, L., Cardarilli, G. C., Di Nunzio, L., Fazzolari, R., Giardino, D., Re, M., and Spanò, S · 2021
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Later among the works it cites.
A review of cooperative multi-agent deep reinforcement learning
Oroojlooy, A. and Hajinezhad, D · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Arenas, M. G., Chiang, H.-T. L., Erez, T., Hasenclever, L., Humplik, J., et al · 2023
Later among the works it cites.
Understanding your agent: Leveraging large language models for behavior explanation
Zhang, X., Guo, Y., Stepputtis, S., Sycara, K., and Campbell, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ask your humans: Using human instructions to improve generalization in reinforcement learning
Chen, V., Gupta, A., and Marino, K · 2021
Cited alongside, same era.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Lee, K., Smith, L., and Abbeel, P · 2021
Cited alongside, same era.
Too many cooks: Bayesian inference for coordinating multi-agent collaboration
Wu, S. A., Wang, R. E., Evans, J. A., Tenenbaum, J. B., Parkes, D. C., and Kleiman-Weiner, M · 2021
Cited alongside, same era.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Zhang, K., Yang, Z., and Başar, T · 2021
Cited alongside, same era.
Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning
Liu, R., Bai, F., Du, Y., and Yang, Y · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
How to talk so ai will learn: Instructions, descriptions, and autonomy
Sumers, T., Hawkins, R., Ho, M. K., Griffiths, T., and Hadfield-Menell, D · 2022
Cited alongside, same era.
Chen, X.-H., Wang, Z., Du, Y., Jiang, S., Fang, M., Yu, Y., and Wang, J · 2024
Later among the works it cites.
Learning to learn faster from human feedback with language model predictive control
Liang, J., Xia, F., Yu, W., Zeng, A., Arenas, M. G., Attarian, M., Bauza, M., Bennice, M., Bewley, A., Dostmohamed, A., et al · 2024
Later among the works it cites.
Model balancing helps low-data training and fine-tuning
Liu, Z., Hu, Y., Pang, T., Zhou, Y., Ren, P., and Yang, Y · 2024
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2024
Later among the works it cites.
GPT-4o (gpt-4o-2024-11-20)
OpenAI · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Later among the works it cites.
Multi-turn reinforcement learning from preference human feedback
Shani, L., Rosenberg, A., Cassel, A., Lang, O., Calandriello, D., Zipori, A., Noga, H., Keller, O., Piot, B., Szpektor, I., Hassidim, A., Matias, Y., and Munos, R · 2024
Later among the works it cites.
An empirical study on google research football multi-agent scenarios
Song, Y., Jiang, H., Tian, Z., Zhang, H., Zhang, Y., Zhu, J., Dai, Z., Zhang, W., and Wang, J · 2024
Later among the works it cites.
Safe multi-agent reinforcement learning with natural language constraints
Wang, Z., Fang, M., Tomilin, T., Fang, F., and Du, Y · 2024
Later among the works it cites.
Multi-agent reinforcement learning from human feedback: Data coverage and algorithmic techniques
Zhang, N., Wang, X., Cui, Q., Zhou, R., Kakade, S. M., and Du, S. S · 2024
Later among the works it cites.
Pragmatic instruction following and goal assistance via cooperative language-guided inverse planning
Zhi-Xuan, T., Ying, L., Mansinghka, V., and Tenenbaum, J. B · 2024
Later among the works it cites.
Eigenspectrum analysis of neural networks without aspect ratio bias
Hu, Y., Goel, K., Killiakov, V., and Yang, Y · 2025
Closest in time.
M+: Extending memoryllm with scalable long-term memory
Wang, Y., Krotov, D., Hu, Y., Gao, Y., Zhou, W., McAuley, J., Gutfreund, D., Feris, R., and He, Z · 2025
Closest in time.
A survey on large language model based human-agent systems
Zou, H. P., Huang, W.-C., Wu, Y., Chen, Y., Miao, C., Nguyen, H., Zhou, Y., Zhang, W., Fang, L., He, L., et al · 2025
Closest in time.