Fetching the paper…
Reading the bibliography…
Multi-agent reinforcement learning (MARL) under partial observability has long been considered challenging, primarily due to the requirement for each agent to maintain a belief over all other agents' local histories -- a domain that generally grows exponentially over time.
K. Pearson, “On lines and planes of closest fit to systems of points in space,” The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science , vol. 2, no. 11, pp. 559–572, 1901
1901
Earlier work this paper cites.
E. J. Sondik, “The optimal control of partially observable Markov processes,” Ph.D. Thesis , 1971
1971
Earlier work this paper cites.
H. S. Witsenhausen, “A standard form for sequential stochastic control,” Mathematical Systems Theory , vol. 7, no. 1, pp. 5–11, 1973
1973
Earlier work this paper cites.
H. P. Sankappanavar and S. Burris, “A course in universal algebra,” Graduate Texts Math , vol. 78, 1981
1981
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine Learning , vol. 8, no. 3-4, pp. 279–292, 1992
1992
Earlier work this paper cites.
S. P. Singh, T. S. Jaakkola, and M. I. Jordan, “Learning without state-estimation in partially observable Markovian decision processes,” in Proc. ICML , 1994, pp. 284–292
1994
Earlier work this paper cites.
C. C. White III and W. T. Scherer, “Finite-memory suboptimal design for partially observed Markov decision processes,” Operations Research , vol. 42, no. 3, pp. 439–455, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
W. Li, H. H. Yue, S. Valle-Cervantes, and S. J. Qin, “Recursive PCA for adaptive process monitoring,” Journal of Process Control , vol. 10, no. 5, pp. 471–486, 2000
2000
Earlier work this paper cites.
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein, “The complexity of decentralized control of Markov decision processes,” Math. Operations Research , vol. 27, no. 4, pp. 819–840, 2002
2002
Earlier work this paper cites.
M. L. Littman and R. S. Sutton, “Predictive representations of state,” in Proc. NeurIPS , 2002, pp. 1555–1561
2002
Earlier work this paper cites.
R. Nair, M. Tambe, M. Yokoo, D. Pynadath, and S. Marsella, “Taming decentralized POMDPs: towards efficient policy computation for multiagent settings,” in Proc. IJCAI , 2003, pp. 705–711
2003
Earlier work this paper cites.
J. W. Lee, J. Park, O. Jangmin, J. Lee, and E. Hong, “A multiagent approach to Q-learning for daily stock trading,” IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans , vol. 37, no. 6, pp. 864–877, 2007
2007
Earlier work this paper cites.
S. Seuken and S. Zilberstein, “Improved memory-bounded dynamic programming for decentralized POMDPs,” in Proc. UAI , 2007, pp. 344–351
2007
Cited alongside, same era.
C. Amato, J. S. Dibangoye, and S. Zilberstein, “Incremental policy generation for finite-horizon Dec-POMDPs,” in ICAPS , 2009
2009
Cited alongside, same era.
D. A. Edwards, “On the Kantorovich–Rubinstein theorem,” Expositiones Mathematicae , vol. 29, no. 4, pp. 387–398, 2011
2011
Cited alongside, same era.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in Proc. AISTATS , 2011, pp. 315–323
2011
Cited alongside, same era.
Y. Duan, B. X. Cui, and X. H. Xu, “A multi-agent reinforcement learning approach to robot soccer,” Artificial Intelligence Review , vol. 38, no. 3, pp. 193–211, 2012
2012
Cited alongside, same era.
F. A. Oliehoek, C. Amato et al. , A concise introduction to decentralized POMDPs . Springer, 2016, vol. 1
2016
Later among the works it cites.
J. K. Gupta, M. Egorov, and M. Kochenderfer, “Cooperative multi-agent control using deep reinforcement learning,” in AAMAS , 2017
2017
Later among the works it cites.
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian, “Deep decentralized multi-task multi-agent reinforcement learning under partial observability,” in Proc. ICML , 2017, pp. 2681–2690
2017
Later among the works it cites.
J. S. Dibangoye and O. Buffet, “Learning to act in decentralized partially observable MDPs,” in Proc. ICML , 2018, pp. 1233–1242
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Banerjee, J. Lyle, L. Kraemer, and R. Yellamraju, “Sample bounded distributed reinforcement learning for decentralized POMDPs,” in Proc. AAAI , 2012, pp. 1256–1262
2012
Cited alongside, same era.
A. Nayyar, A. Mahajan, and D. Teneketzis, “Decentralized stochastic control with partial history sharing: A common information approach,” IEEE Trans. Automatic Control , vol. 58, no. 7, pp. 1644–1658, 2013
2013
Cited alongside, same era.
A. Nayyar, A. Mahajan, and D. Teneketzis, “The common-information approach to decentralized stochastic control,” in Information and Control in Networks . Springer, 2014, pp. 123–156
2014
Cited alongside, same era.
J. S. Dibangoye, O. Buffet, and F. Charpillet, “Error-bounded approximations for infinite-horizon discounted decentralized POMDPs,” in Proc. ECML/PKDD , 2014, pp. 338–353
2014
Cited alongside, same era.
2015
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Cited alongside, same era.
S. Wang, J. Wan, D. Zhang, D. Li, and C. Zhang, “Towards smart factory for industry 4.0: a self-organized multi-agent system with big data based feedback and coordination,” Computer Networks , vol. 101, pp. 158–168, 2016
2016
Cited alongside, same era.
2018
Later among the works it cites.
X. Ma, P. Karkus, D. Hsu, and W. S. Lee, “PF-LSTM: Belief state particle filter for LSTM,” in NeurIPS RLPO Workshop , 2018
2018
Later among the works it cites.
P. Moreno, J. Humplik, G. Papamakarios, B. A. Pires, L. Buesing, N. Heess, and T. Weber, “Neural belief states for partially observed domains,” in NeurIPS RLPO Workshop , 2018
2018
Later among the works it cites.
T. Lesort, N. Díaz-Rodríguez, J.-F. Goudou, and D. Filliat, “State representation learning for control: An overview,” Neural Networks , vol. 108, pp. 379–392, 2018
2018
Later among the works it cites.
2019
Later among the works it cites.
J. Subramanian and A. Mahajan, “Approximate information state for partially observed systems,” in Proc. CDC , 2019, pp. 1629–1636
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Zhang, E. Miehling, and T. Başar, “Online planning for decentralized stochastic control with partial history sharing,” in Prof. ACC , 2019, pp. 3544–3550
2019
Later among the works it cites.