Fetching the paper…
Reading the bibliography…
Being able to harness the power of large datasets for developing cooperative multi-agent controllers promises to unlock enormous value for real-world applications.
The complexity of decentralized control of markov decision processes
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein · 2002
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Openai gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
L. Kraemer and B. Banerjee · 2016
Earlier work this paper cites.
Cooperative multi-agent control using deep reinforcement learning
J. K. Gupta, M. Egorov, and M. Kochenderfer · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mordatch · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Earlier work this paper cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
T. Rashid, M. Samvelyan, C. Schroeder, G. Farquhar, J. Foerster, and S. Whiteson · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Towards sample efficient reinforcement learning
Y. Yu · 2018
Earlier work this paper cites.
Reducing overestimation bias in multi-agent domains using double centralized critics
J. Ackermann, V. Gabler, T. Osa, and M. Sugiyama · 2019
Earlier work this paper cites.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Earlier work this paper cites.
The starcraft multi-agent challenge
M. Samvelyan, T. Rashid, C. Schroeder de Witt, G. Farquhar, N. Nardelli, T. G. Rudner, C.-M. Hung, P. H. Torr, J. Foerster, and S. Whiteson · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, C. Gulcehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Earlier work this paper cites.
CityFlow: A multi-agent reinforcement learning environment for large scale city traffic scenario
H. Zhang, S. Feng, C. Liu, Y. Ding, Y. Zhu, Z. Zhou, W. Zhang, Y. Yu, H. Jin, and Z. Li · 2019
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
C. Gulcehre, Z. Wang, A. Novikov, T. Paine, S. Gómez, K. Zolna, R. Agarwal, J. S. Merel, D. J. Mankowitz, C. Paduraru, et al · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
Flatland-rl : Multi-agent reinforcement learning on trains
S. Mohanty, E. Nygren, F. Laurent, M. Schneider, C. Scheller, N. Bhattacharya, J. Watson, A. Egli, C. Eichenberger, C. Baumberger, G. Vienken, I. Sturm, G. Sartoretti, and G. Spigler · 2020
Cited alongside, same era.
Multi-agent routing value iteration network
Q. Sykora, M. Ren, and R. Urtasun · 2020
Cited alongside, same era.
Citylearn: Standardizing research in multi-agent reinforcement learning for demand response and urban energy management
J. R. Vazquez-Canteli, S. Dey, G. Henze, and Z. Nagy · 2020
Cited alongside, same era.
Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning
Y. Yang, X. Ma, L. Chenghao, Z. Zheng, Q. Zhang, G. Huang, J. Yang, and Q. Zhao · 2021
Later among the works it cites.
A review of deep reinforcement learning for smart building energy management
L. Yu, S. Qin, M. Zhang, C. Shen, T. Jiang, and X. Guan · 2021
Later among the works it cites.
Finite-sample analysis for decentralized batch multiagent reinforcement learning with networked agents
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Başar · 2021
Later among the works it cites.
Madiff: Offline multi-agent learning with diffusion models
Z. Zhu, M. Liu, L. Mao, B. Kang, M. Xu, Y. Yu, S. Ermon, and W. Zhang · 2021
Later among the works it cites.
Provably efficient offline multi-agent reinforcement learning via strategy-wise bonus
Q. Cui and S. S. Du · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Zhou, Z. Liu, P. Sui, Y. Li, and Y. Y. Chung · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare · 2021
Cited alongside, same era.
Decision transformer: reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
Minimax sample complexity for turn-based stochastic game
Q. Cui and L. F. Yang · 2021
Cited alongside, same era.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
G. Dulac-Arnold, N. Levine, D. J. Mankowitz, J. Li, C. Paduraru, S. Gowal, and T. Hester · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
S. Fujimoto and S. S. Gu · 2021
Cited alongside, same era.
Datasheets for datasets
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. I. au2, and K. Crawford · 2021
Cited alongside, same era.
Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning
B. Ellis, S. Moalla, M. Samvelyan, M. Sun, A. Mahajan, J. N. Foerster, and S. Whiteson · 2022
Later among the works it cites.
Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters
K. Ghasemipour, S. S. Gu, and O. Nachum · 2022
Later among the works it cites.
Towards a standardised performance evaluation protocol for cooperative MARL
R. Gorsane, O. Mahjoub, R. J. de Kock, R. Dubb, S. Singh, and A. Pretorius · 2022
Later among the works it cites.
A review of safe reinforcement learning: Methods, theory and applications
S. Gu, L. Yang, Y. Du, G. Chen, F. Walter, J. Wang, Y. Yang, and A. Knoll · 2022
Later among the works it cites.
Winning the citylearn challenge: Adaptive optimization with evolutionary search under trajectory-based guidance
V. Khattar and M. Jin · 2022
Later among the works it cites.
Showing your offline reinforcement learning work: Online evaluation budget matters
V. Kurenkov and S. Kolesnikov · 2022
Later among the works it cites.
Challenges and opportunities in offline reinforcement learning from visual observations
C. Lu, P. J. Ball, T. G. Rudner, J. Parker-Holder, M. A. Osborne, and Y. W. Teh · 2022
Later among the works it cites.
A deeper understanding of state-based critics in multi-agent reinforcement learning
X. Lyu, A. Baisero, Y. Xiao, and C. Amato · 2022
Later among the works it cites.
Plan better amid conservatism: Offline multi-agent reinforcement learning with actor rectification
L. Pan, L. Huang, T. Ma, and H. Xu · 2022
Later among the works it cites.
Tackling climate change with machine learning
D. Rolnick, P. L. Donti, L. H. Kaack, K. Kochanski, A. Lacoste, K. Sankaran, A. S. Ross, N. Milojevic-Dupont, N. Jaques, A. Waldman-Brown, A. S. Luccioni, T. Maharaj, E. D. Sherwin, S. K. Mukkavilli, K. P. Kording, C. P. Gomes, A. Y. Ng, D. Hassabis, J. C. Platt, F. Creutzig, J. Chayes, and Y. Bengio · 2022
Later among the works it cites.
Constraints penalized q-learning for safe offline reinforcement learning
H. Xu, X. Zhan, and X. Zhu · 2022
Later among the works it cites.
Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets
H. Zhong, W. Xiong, J. Tan, L. Wang, T. Zhang, Z. Wang, and Z. Yang · 2022
Later among the works it cites.
A model-based solution to the offline multi-agent reinforcement learning coordination problem
P. Barde, J. Foerster, D. Nowrouzezahrai, and A. Zhang · 2023
Closest in time.
Reduce, reuse, recycle: Selective reincarnation in multi-agent reinforcement learning
C. Formanek, C. R. Tilbury, J. P. Shock, K. ab Tessera, and A. Pretorius · 2023
Closest in time.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
M. Nakamoto, Y. Zhai, A. Singh, Y. Ma, C. Finn, A. Kumar, and S. Levine · 2023
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
R. F. Prudencio, M. R. O. A. Maximo, and E. L. Colombini · 2023
Closest in time.