Fetching the paper…
Reading the bibliography…
Credit assignment, the process of attributing credit or blame to individual agents for their contributions to a team's success or failure, remains a fundamental challenge in multi-agent reinforcement learning (MARL), particularly in environments with sparse rewards.
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Multi-agent reinforcement learning for traffic light control
Wiering, M. A. et al · 2000
Earlier work this paper cites.
Monotonic value function factorisation for deep multi-agent reinforcement learning, 2020
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S · 2003
Earlier work this paper cites.
Value-decomposition multi-agent actor-critics, 2020
Su, J., Adams, S., and Beling, P. A · 2007
Earlier work this paper cites.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Matignon, L., Laurent, G. J., and Le Fort-Piat, N · 2012
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y. I., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Value-decomposition networks for cooperative multi-agent learning
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., et al · 2017
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Cited alongside, same era.
Multi-agent deep reinforcement learning for large-scale traffic signal control
Chu, T., Wang, J., Codecà, L., and Li, Z · 2019
Cited alongside, same era.
Liir: Learning individual intrinsic reward in multi-agent reinforcement learning
Du, Y., Han, L., Fang, M., Liu, J., Dai, T., and Tao, D · 2019
Cited alongside, same era.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P., Strouse, D., Leibo, J. Z., and De Freitas, N · 2019
Cited alongside, same era.
The perils of trial-and-error reward design: misdesign through overfitting and invalid task specifications
Booth, S., Knox, W. B., Shah, J., Niekum, S., Stone, P., and Allievi, A · 2023
Later among the works it cites.
An adaptive entropy-regularization framework for multi-agent reinforcement learning
Kim, W. and Sung, Y · 2023
Later among the works it cites.
A variational approach to mutual information-based coordination for multi-agent reinforcement learning
Kim, W., Jung, W., Cho, M., and Sung, Y · 2023
Later among the works it cites.
Reward (mis) design for autonomous driving
Knox, W. B., Allievi, A., Banzhaf, H., Schmitt, F., and Stone, P · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Son, K., Kim, D., Kang, W. J., Hostallero, D. E., and Yi, Y · 2019
Cited alongside, same era.
Qplex: Duplex dueling multi-agent q-learning
Wang, J., Ren, Z., Liu, T., Yu, Y., and Zhang, C · 2020
Cited alongside, same era.
Pettingzoo: Gym for multi-agent reinforcement learning
Terry, J., Black, B., Grammel, N., Jayakumar, M., Hari, A., Sullivan, R., Santos, L. S., Dieffendahl, C., Horsch, C., Perez-Vicente, R., Williams, N., Lokesh, Y., and Ravi, P · 2021
Cited alongside, same era.
Defining and characterizing reward gaming
Skalse, J., Howe, N., Krasheninnikov, D., and Krueger, D · 2022
Cited alongside, same era.
Multi-agent reinforcement learning is a sequence modeling problem, 2022
Wen, M., Kuba, J. G., Lin, R., Zhang, W., Wen, Y., Wang, J., and Yang, Y · 2022
Cited alongside, same era.
The surprising effectiveness of ppo in cooperative, multi-agent games, 2022
Yu, C., Velu, A., Vinitsky, E., Gao, J., Wang, Y., Bayen, A., and Wu, Y · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Lazy agents: a new perspective on solving sparse reward problem in multi-agent reinforcement learning
Liu, B., Pu, Z., Pan, Y., Yi, J., Liang, Y., and Zhang, D · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Assigning credit with partial reward decoupling in multi-agent proximal policy optimization, 2024
Kapoor, A., Freed, B., Choset, H., and Schneider, J · 2024
Later among the works it cites.
Navigating noisy feedback: Enhancing reinforcement learning with error-prone language models
Lin, M., Shi, S., Guo, Y., Chalaki, B., Tadiparthi, V., Pari, E. M., Stepputtis, S., Campbell, J., and Sycara, K · 2024
Later among the works it cites.
A minimaximalist approach to reinforcement learning from human feedback
Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A · 2024
Later among the works it cites.
Multi-agent reinforcement learning for autonomous driving: A survey
Zhang, R., Hou, J., Walter, F., Gu, S., Guan, J., Röhrbein, F., Du, Y., Cai, P., Chen, G., and Knoll, A · 2024
Later among the works it cites.