Fetching the paper…
Reading the bibliography…
Reinforcement Learning in large action spaces is a challenging problem.
Introduction to lambda calculus
Barendregt, H. P · 1984
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
An introduction to computational learning theory
Kearns, M. J., Vazirani, U. V., and Vazirani, U · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Boutilier, C · 1996
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Actor-Critic Algorithms
Konda, V. and Tsitsiklis, J. N · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J · 2004
Earlier work this paper cites.
Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
Matignon, L., Laurent, G. J., and Le Fort-Piat, N · 2007
Earlier work this paper cites.
Qplex: Duplex dueling multi-agent q-learning
Wang, J., Ren, Z., Liu, T., Yu, Y., and Zhang, C · 2008
Earlier work this paper cites.
Tensor decompositions and applications
Kolda, T. G. and Bader, B. W · 2009
Earlier work this paper cites.
Multi-agent reinforcement learning: An overview
Buşoniu, L., Babuška, R., and De Schutter, B · 2010
Earlier work this paper cites.
Rode: Learning roles to decompose multi-agent tasks
Wang, T., Gupta, T., Mahajan, A., Peng, B., Whiteson, S., and Zhang, C · 2010
Earlier work this paper cites.
Efficient offline communication policies for factored multiagent pomdps
Messias, J. a. V., Spaan, M. T. J., and Lima, P. U · 2011
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2011
Earlier work this paper cites.
A method of moments for mixture models and hidden markov models
Anandkumar, A., Hsu, D., and Kakade, S. M · 2012
Cited alongside, same era.
A tensor factorization approach to generalization in multi-agent reinforcement learning
Bromuri, S · 2012
Cited alongside, same era.
Most tensor problems are np-hard
Hillar, C. J. and Lim, L.-H · 2013
Cited alongside, same era.
The sample-complexity of general reinforcement learning
Lattimore, T., Hutter, M., and Sunehag, P · 2013
Cited alongside, same era.
Tensor decompositions for learning latent variable models
Anandkumar, A., Ge, R., Hsu, D., Kakade, S. M., and Telgarsky, M · 2014
Cited alongside, same era.
Provable tensor factorization with missing data
Jain, P. and Oh, S · 2014
Cited alongside, same era.
Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., and Graepel, T · 2017
Later among the works it cites.
Learning to coordinate with coordination graphs in repeated single-stage multi-agent decision problems
Bargiacchi, E., Verstraeten, T., Roijers, D., Nowé, A., and Hasselt, H · 2018
Later among the works it cites.
Factorized q-learning for large-scale multi-agent systems
Chen, Y., Zhou, M., Wen, Y., Yang, Y., Su, Y., Zhang, W., Zhang, D., Wang, J., and Liu, H · 2018
Later among the works it cites.
Counterfactual multi-agent policy gradients
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Later among the works it cites.
QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning
Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Reinforcement learning in rich-observation mdps using spectral methods
Azizzadenesheli, K., Lazaric, A., and Anandkumar, A · 2016
Cited alongside, same era.
On the expressive power of deep learning: A tensor analysis
Cohen, N., Sharir, O., and Shashua, A · 2016
Cited alongside, same era.
Pac reinforcement learning with rich observations
Krishnamurthy, A., Agarwal, A., and Langford, J · 2016
Cited alongside, same era.
A Concise Introduction to Decentralized POMDPs
Oliehoek, F. A. and Amato, C · 2016
Cited alongside, same era.
Regularized policy gradients: direct variance reduction in policy gradient estimation
Zhao, T., Niu, G., Xie, N., Yang, J., and Sugiyama, M · 2016
Cited alongside, same era.
Later among the works it cites.
Multiagent soft q-learning
Wei, E., Wicke, D., Freelan, D., and Luke, S · 2018
Later among the works it cites.
Reinforcement Learning in Structured and Partially Observable Environments
Azizzadenesheli, K · 2019
Later among the works it cites.
T-net: Parametrizing fully convolutional nets with a single high-order tensor
Kossaifi, J., Bulat, A., Tzimiropoulos, G., and Pantic, M · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration
Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S · 2019
Later among the works it cites.
The StarCraft Multi-Agent Challenge
Samvelyan, M., Rashid, T., de Witt, C. S., Farquhar, G., Nardelli, N., Rudner, T. G., Hung, C.-M., Torr, P. H., Foerster, J., and Whiteson, S · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Son, K., Kim, D., Kang, W. J., Hostallero, D. E., and Yi, Y · 2019
Later among the works it cites.
Smix: Enhancing centralized value functions for cooperative multi-agent reinforcement learning
Yao, X., Wen, C., Wang, Y., and Tan, X · 2019
Later among the works it cites.
Incremental multi-domain learning with network latent tensor factorization
Bulat, A., Kossaifi, J., Tzimiropoulos, G., and Pantic, M · 2020
Later among the works it cites.
Uneven: Universal value exploration for multi-agent reinforcement learning
Gupta, T., Mahajan, A., Peng, B., Böhmer, W., and Whiteson, S · 2020
Later among the works it cites.
Factorized higher-order cnns with an application to spatio-temporal emotion estimation
Kossaifi, J., Toisoul, A., Bulat, A., Panagakis, Y., Hospedales, T., and Pantic, M · 2020
Later among the works it cites.
Qatten: A general framework for cooperative multiagent reinforcement learning
Yang, Y., Hao, J., Liao, B. L., Shao, K., Chen, G., Liu, W., and Tang, H · 2020
Later among the works it cites.