Fetching the paper…
Reading the bibliography…
Opponent modeling is essential to exploit sub-optimal opponents in strategic interactions.
Iterative solution of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
Stochastic games
Shapley, L. S · 1953
Earlier work this paper cites.
Attention, intentions, and the structure of discourse
Grosz, B. and Sidner, C. L · 1986
Earlier work this paper cites.
Active learner modelling
McCalla, G., Vassileva, J., Greer, J., and Bull, S · 2000
Earlier work this paper cites.
Learning to play bayesian games
Dekel, E., Fudenberg, D., and Levine, D. K · 2004
Earlier work this paper cites.
Identifying terrorist activity with ai plan recognition technology
Jarvis, P. A., Lunt, T. F., and Myers, K. L · 2005
Earlier work this paper cites.
Beliefs in repeated games
Nachbar, J. H · 2005
Earlier work this paper cites.
Learning against opponents with bounded memory
Powers, R. and Shoham, Y · 2005
Earlier work this paper cites.
An integrated trust and reputation model for open multi-agent systems
Huynh, T. D., Jennings, N. R., and Shadbolt, N. R · 2006
Earlier work this paper cites.
Reaching pareto-optimality in prisoner’s dilemma using conditional joint action learning
Banerjee, D. and Sen, S · 2007
Earlier work this paper cites.
A kernel method for the two-sample-problem
Gretton, A., Borgwardt, K., Rasch, M., Schölkopf, B., and Smola, A. J · 2007
Earlier work this paper cites.
Policy recognition for multi-player tactical scenarios
Sukthankar, G. and Sycara, K · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2007
Earlier work this paper cites.
Generating diverse opponents with multiobjective evolution
Agapitos, A., Togelius, J., Lucas, S. M., Schmidhuber, J., and Konstantinidis, A · 2008
Earlier work this paper cites.
Kernel choice and classifiability for rkhs embeddings of probability distributions
Fukumizu, K., Gretton, A., Lanckriet, G. R., Schölkopf, B., and Sriperumbudur, B. K · 2009
Earlier work this paper cites.
A data mining approach to strategy prediction
Weber, B. G. and Mateas, M · 2009
Earlier work this paper cites.
Non-parametric estimation of integral probability metrics
Sriperumbudur, B. K., Fukumizu, K., Gretton, A., Schölkopf, B., and Lanckriet, G. R · 2010
Cited alongside, same era.
A Bayesian model for opening prediction in rts games with application to StarCraft
Synnaeve, G. and Bessiere, P · 2011
Cited alongside, same era.
A kernel two-sample test
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A · 2012
Cited alongside, same era.
Opponent modeling by expectation–maximization and sequence prediction in simplified poker
Mealing, R. and Shapiro, J. L · 2015
Cited alongside, same era.
Planning over multi-agent epistemic states: A classical planning approach
Muise, C., Belle, V., Felli, P., McIlraith, S., Miller, T., Pearce, A. R., and Sonenberg, L · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Later among the works it cites.
Learning with opponent-learning awareness
Foerster, J., Chen, R. Y., Al-Shedivat, M., Whiteson, S., Abbeel, P., and Mordatch, I · 2018
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Hessel, M., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., et al · 2018
Later among the works it cites.
Unsupervised meta-learning for reinforcement learning
Gupta, A., Eysenbach, B., Finn, C., and Levine, S · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He, H., Boyd-Graber, J., Kwok, K., and Daumé III, H · 2016
Cited alongside, same era.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Cited alongside, same era.
Reasoning about hypothetical agent behaviours and their parameters
Albrecht, S. V. and Stone, P · 2017
Cited alongside, same era.
Negotiating with other minds: the role of recursive theory of mind in negotiation with incomplete information
de Weerd, H., Verbrugge, R., and Verheij, B · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al · 2017
Cited alongside, same era.
Robust deep reinforcement learning with adversarial attacks
Pattanaik, A., Tang, Z., Liu, S., Bommannan, G., and Chowdhary, G · 2018
Later among the works it cites.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2018
Later among the works it cites.
Machine theory of mind
Rabinowitz, N., Perbet, F., Song, F., Zhang, C., Eslami, S. A., and Botvinick, M · 2018
Later among the works it cites.
Probabilistic recursive reasoning for multi-agent reinforcement learning
Wen, Y., Yang, Y., Luo, R., Wang, J., and Pan, W · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D · 2018
Later among the works it cites.
Wasserstein robust reinforcement learning
Abdullah, M. A., Ren, H., Ammar, H. B., Milenkovic, V., Luo, R., Zhang, M., and Wang, J · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Later among the works it cites.
Single deep counterfactual regret minimization
Steinberger, E · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 2019
Later among the works it cites.
Meta-learning in neural networks: A survey
Hospedales, T., Antoniou, A., Micaelli, P., and Storkey, A · 2020
Later among the works it cites.