Fetching the paper…
Reading the bibliography…
This paper serves to introduce the reader to the field of multi-agent reinforcement learning (MARL) and its intersection with methods from the study of causality.
W. Deming, J. Neumann, and O. Morgenstern, “Theory of games and economic behavior.” Journal of the American Statistical Association , vol. 40, p. 263, 1944
1944
Earlier work this paper cites.
R. Bellman, “The theory of dynamic programming,” Bulletin of the American Mathematical Society , vol. 60, pp. 503–515, 1954
1954
Earlier work this paper cites.
K. Åström, “Optimal control of markov processes with incomplete state information,” Journal of Mathematical Analysis and Applications , vol. 10, pp. 174–205, 1964
1964
Earlier work this paper cites.
J. Harsanyi, “Games with incomplete information played by "bayesian" players, i-iii: Part i. the basic model&,” Manag. Sci. , vol. 14, pp. 159–182, 1967
1967
Earlier work this paper cites.
P. Rosenbaum and D. Rubin, “The central role of the propensity score in observational studies for causal effects,” Biometrika , vol. 70, pp. 41–55, 1983
1983
Earlier work this paper cites.
M. Littman, “Markov games as a framework for multi-agent reinforcement learning,” 1994
1994
Earlier work this paper cites.
——, “Graphs, causality, and structural equation models,” Sociological Methods & Research , vol. 27, pp. 226 – 284, 1998
1998
Earlier work this paper cites.
P. Spirtes, C. Glymour, and R. Scheines, “Causation, prediction, and search, second edition,” 2000
2000
Earlier work this paper cites.
D. H. Wolpert, K. Tumer, and K. Swanson, “Optimal wonderful life utility functions in multi-agent systems,” 2000
2000
Earlier work this paper cites.
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein, “The complexity of decentralized control of markov decision processes,” Mathematics of operations research , vol. 27, no. 4, pp. 819–840, 2002
2002
Earlier work this paper cites.
S. Murphy, “Optimal dynamic treatment regimes,” Journal of The Royal Statistical Society Series B-statistical Methodology , vol. 65, pp. 331–355, 2003
2003
Earlier work this paper cites.
Y. Shoham, R. Powers, and T. Grenager, “Multi-agent reinforcement learning: a critical survey,” 2003
2003
Earlier work this paper cites.
E. Hansen, D. Bernstein, and S. Zilberstein, “Dynamic programming for partially observable stochastic games,” 2004
2004
Earlier work this paper cites.
R. Sutton and A. Barto, “Reinforcement learning: An introduction,” IEEE Transactions on Neural Networks , vol. 16, pp. 285–286, 2005
2005
Earlier work this paper cites.
S. Meganck, S. Maes, B. Manderick, and P. Leray, “Distributed learning of multi-agent causal models,” pp. 285–288, 2005
2005
Earlier work this paper cites.
R. MacLehose, S. Kaufman, J. Kaufman, and C. Poole, “Bounding causal effects under uncontrolled confounding using counterfactuals,” Epidemiology , vol. 16, pp. 548–555, 2005
2005
Earlier work this paper cites.
D. Rubin, “Causal inference using potential outcomes,” Journal of the American Statistical Association , vol. 100, pp. 322 – 331, 2005
2005
Earlier work this paper cites.
S. Maes, S. Meganck, and B. Manderick, “Inference in multi-agent causal models,” Int. J. Approx. Reason. , vol. 46, pp. 274–299, 2007
2007
Earlier work this paper cites.
P. Hoyer, D. Janzing, J. Mooij, J. Peters, and B. Schölkopf, “Nonlinear causal discovery with additive noise models,” 2008
2008
Earlier work this paper cites.
F. A. Oliehoek, M. T. Spaan, N. Vlassis, and S. Whiteson, “Exploiting locality of interaction in factored dec-pomdps,” pp. 517–524, 2008
2008
Earlier work this paper cites.
J. Pearl, Causality . Cambridge university press, 2009
2009
Earlier work this paper cites.
M. Wooldridge, An introduction to multiagent systems . John wiley & sons, 2009
2009
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on Knowledge and Data Engineering , vol. 22, pp. 1345–1359, 2010
2010
Earlier work this paper cites.
D. Bertsekas, Dynamic programming and optimal control: Volume I . Athena scientific, 2012, vol. 1
2012
Earlier work this paper cites.
L. Wang, A. Rotnitzky, X. Lin, R. Millikan, and P. Thall, “Evaluation of viable dynamic treatment regimes in a sequentially randomized trial of advanced prostate cancer,” Journal of the American Statistical Association , vol. 107, pp. 493 – 508, 2012
2012
Earlier work this paper cites.
J. Pearl, “The causal foundations of structural equation modeling,” 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Pearl and E. Bareinboim, “External validity: From do-calculus to transportability across populations,” Statistical Science , vol. 29, no. 4, Nov 2014. [Online]. Available: http://dx.doi.org/10.1214/14-STS486
2014
Earlier work this paper cites.
S. Devlin, L. Yliniemi, D. Kudenko, and K. Tumer, “Potential-based difference rewards for multiagent reinforcement learning,” pp. 165–172, 2014
2014
Earlier work this paper cites.
M. Zimmer, P. Viappiani, and P. Weng, “Teacher-student framework: a reinforcement learning approach,” 2014
2014
Earlier work this paper cites.
P. K.J., H. K. A.N, and S. Bhatnagar, “Multi-agent reinforcement learning for traffic signal control,” pp. 2529–2534, 2014
2014
Earlier work this paper cites.
A. Marcellesi, “External validity: Is there still a problem?” Philosophy of Science , vol. 82, pp. 1308 – 1317, 2015
2015
Earlier work this paper cites.
E. Bareinboim, A. Forney, and J. Pearl, “Bandits with unobserved confounders: A causal approach,” 2015
2015
Earlier work this paper cites.
E. Bareinboim and J. Pearl, “Causal inference and the data-fusion problem,” Proceedings of the National Academy of Sciences , vol. 113, pp. 7345 – 7352, 2016
2016
Earlier work this paper cites.
S. O. Becker, “Using instrumental variables to establish causality,” The IZA World of Labor , pp. 250–250, 2016
2016
Cited alongside, same era.
J. Zhang and E. Bareinboim, “Markov decision processes with unobserved confounders: A causal approach,” 2016
2016
Cited alongside, same era.
J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson, “Learning to communicate with deep multi-agent reinforcement learning,” 2016
2016
Cited alongside, same era.
F. A. Oliehoek and C. Amato, A concise introduction to decentralized POMDPs . Springer, 2016
2016
Cited alongside, same era.
A. Holzinger, “Interactive machine learning for health informatics: when do we need the human-in-the-loop?” Brain Informatics , vol. 3, pp. 119 – 131, 2016
2016
Cited alongside, same era.
S. Omidshafiei, D.-K. Kim, M. Liu, G. Tesauro, M. Riemer, C. Amato, M. Campbell, and J. How, “Learning to teach in cooperative multiagent reinforcement learning,” 2019
2019
Later among the works it cites.
E. Ilhan, J. Gow, and D. P. Liebana, “Teaching on a budget in multi-agent deep reinforcement learning,” 2019 IEEE Conference on Games (CoG) , pp. 1–8, 2019
2019
Later among the works it cites.
B. Schölkopf, “Causality for machine learning,” arXiv preprint arXiv:1911.10500 , 2019
2019
Later among the works it cites.
E. Bareinboim, J. D. Correa, D. Ibeling, and T. Icard, “On pearl’s hierarchy and the foundations of causal inference,” ACM Special Volume in Honor of Judea Pearl (provisional title) , vol. 2, no. 3, p. 4, 2020
2020
Later among the works it cites.
E. Bareinboim, “Causal reinforcement learning,” 2020, iCML 2020. [Online]. Available: https://icml.cc/virtual/2020/tutorial/5752
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Forney, J. Pearl, and E. Bareinboim, “Counterfactual data-fusion for online reinforcement learners,” 2017
2017
Cited alongside, same era.
J. Peters, D. Janzing, and B. Schölkopf, “Elements of causal inference: Foundations and learning algorithms,” 2017
2017
Cited alongside, same era.
J. Zhang and E. Bareinboim, “Transfer learning in multi-armed bandits: A causal approach,” pp. 1340–1346, 2017. [Online]. Available: https://doi.org/10.24963/ijcai.2017/186
2017
Cited alongside, same era.
2018
Cited alongside, same era.
J. Pearl and D. Mackenzie, The Book of Why . New York: Basic Books, 2018
2018
Cited alongside, same era.
V. Landeiro and A. Culotta, “Robust text classification under confounding shift,” J. Artif. Intell. Res. , vol. 63, pp. 391–419, 2018
2018
Cited alongside, same era.
S. Lee and E. Bareinboim, “Structural causal bandits: where to intervene?” Advances in Neural Information Processing Systems 31 , vol. 31, 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin, “Addressing distribution shift in online reinforcement learning with offline datasets,” 2020
2020
Later among the works it cites.
——, “Characterizing optimal mixed policies: Where to intervene and what to observe,” Advances in neural information processing systems , vol. 33, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
——, “Designing optimal dynamic treatment regimes: A causal reinforcement learning approach,” 2020
2020
Later among the works it cites.
J. G. Richens, C. M. Lee, and S. Johri, “Improving the accuracy of medical diagnosis with causal machine learning,” Nature Communications , vol. 11, 2020
2020
Later among the works it cites.
——, “Can humans be out of the loop?” 2020
2020
Later among the works it cites.
J. Zhang, D. Kumor, and E. Bareinboim, “Causal imitation learning with unobserved confounders,” Advances in neural information processing systems , vol. 33, 2020
2020
Later among the works it cites.
A. Jaber, M. Kocaoglu, K. Shanmugam, and E. Bareinboim, “Causal discovery from soft interventions with unknown targets: Characterization and learning,” Advances in neural information processing systems , vol. 33, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Gonzalez-Soto, L. Sucar, and H. Escalante, “Causal games and causal nash equilibrium,” Res. Comput. Sci. , vol. 149, pp. 123–133, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
N. Rieke, J. Hancox, W. Li, F. Milletari, H. Roth, S. Albarqouni, S. Bakas, M. Galtier, B. Landman, K. H. Maier-Hein, S. Ourselin, M. J. Sheller, R. M. Summers, A. Trask, D. Xu, M. Baust, and M. Cardoso, “The future of digital health with federated learning,” NPJ Digital Medicine , vol. 3, 2020
2020
Later among the works it cites.
J. Lussange, I. Lazarevich, S. Bourgeois-Gironde, S. Palminteri, and B. Gutkin, “Modelling stock markets by multi-agent reinforcement learning,” Computational Economics , vol. 57, pp. 113–147, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
J. Li, Y. Luo, and X. Zhang, “Causal reinforcement learning: An instrumental variable approach,” Available at SSRN 3792824 , 2021
2021
Closest in time.
2021
Closest in time.
S. Gronauer and K. Diepold, “Multi-agent deep reinforcement learning: a survey,” Artificial Intelligence Review , pp. 1–49, 2021
2021
Closest in time.
2021
Closest in time.
A. Pretorius, K. ab Tessera, A. P. Smit, C. Formanek, S. J. Grimbly, K. Eloff, S. Danisa, L. Francis, J. Shock, H. Kamper, W. Brink, H. Engelbrecht, A. Laterre, and K. Beguir, “Mava: a research framework for distributed multi-agent reinforcement learning,” 2021
2021
Closest in time.
P. Sharma, R. Fernandez, E. G. Zaroukian, M. Dorothy, A. Basak, and D. E. Asher, “Survey of recent multi-agent reinforcement learning algorithms utilizing centralized training,” 2021
2021
Closest in time.
2021
Closest in time.
W. Kim, J. Park, and Y. Sung, “Communication in multi-agent reinforcement learning: Intention sharing,” 2021
2021
Closest in time.