Fetching the paper…
Reading the bibliography…
Agents can achieve effective interaction with previously unknown other agents by maintaining beliefs over a set of hypothetical behaviours, or types, that these agents may have.
C. Watkins and P. Dayan · 1906
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. Thompson · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
H. Robbins · 1952
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
P. Hart, N. Nilsson, and B. Raphael · 1968
Earlier work this paper cites.
Rational learning leads to Nash equilibrium
E. Kalai and E. Lehrer · 1993
Earlier work this paper cites.
Weak and strong merging of opinions
E. Kalai and E. Lehrer · 1994
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. Schapire · 1995
Earlier work this paper cites.
Learning models of intelligent agents
D. Carmel and S. Markovitch · 1996
Earlier work this paper cites.
Tractable inference for complex stochastic processes
X. Boyen and D. Koller · 1998
Earlier work this paper cites.
Evolving aspirations and cooperation
R. Karandikar, D. Mookherjee, D. Ray, and F. Vega-Redondo · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
Exploration strategies for model-based learning in multi-agent systems
D. Carmel and S. Markovitch · 1999
Earlier work this paper cites.
Introduction to Global Optimization
R. Horst, P. Pardalos, and N. Thoai · 2000
Earlier work this paper cites.
The factored frontier algorithm for approximate inference in DBNs
K. Murphy and Y. Weiss · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Cited alongside, same era.
Coordination in multiagent reinforcement learning: a Bayesian approach
G. Chalkiadakis and C. Boutilier · 2003
Cited alongside, same era.
Exploration-exploitation tradeoffs for experts algorithms in reactive environments
D. de Farias and N. Megiddo · 2004
Cited alongside, same era.
Predicting opponent actions by observation
A. Ledezma, R. Aler, A. Sanchis, and D. Borrajo · 2004
Cited alongside, same era.
Coordination and adaptation in impromptu teams
M. Bowling and P. McCracken · 2005
Cited alongside, same era.
A framework for sequential planning in multiagent settings
P. Gmytrasiewicz and P. Doshi · 2005
Cited alongside, same era.
Multivariate polynomial integration and differentiation are polynomial time inapproximable unless P = NP
B. Fu · 2012
Later among the works it cites.
Practical Bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. Adams · 2012
Later among the works it cites.
A game-theoretic model and best-response learning method for ad hoc coordination in multiagent systems
S. Albrecht and S. Ramamoorthy · 2013
Later among the works it cites.
Teamwork with limited knowledge of teammates
S. Barrett, P. Stone, S. Kraus, and A. Rosenfeld · 2013
Later among the works it cites.
Bayesian approach to global optimization: theory and applications
J. Mockus · 2013
Later among the works it cites.
On convergence and optimality of best-response learning with policy types in multiagent systems
S. Albrecht and S. Ramamoorthy · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beliefs in repeated games
J. Nachbar · 2005
Cited alongside, same era.
Bayes’ bluff: opponent modelling in poker
F. Southey, M. Bowling, B. Larson, C. Piccione, N. Burch, D. Billings, and C. Rayner · 2005
Cited alongside, same era.
On the difficulty of achieving equilibrium in interactive POMDPs
P. Doshi and P. Gmytrasiewicz · 2006
Cited alongside, same era.
Bandit based Monte-Carlo planning
L. Kocsis and C. Szepesvári · 2006
Cited alongside, same era.
Gaussian Processes for Machine Learning
C. Rasmussen and C. Williams · 2006
Cited alongside, same era.
Ad hoc autonomous agent teams: collaboration without pre-coordination
P. Stone, G. Kaminka, S. Kraus, and J. Rosenschein · 2010
Cited alongside, same era.
Later among the works it cites.
Team behavior in interactive dynamic influence diagrams with applications to ad hoc teams
M. Chandrasekaran, P. Doshi, Y. Zeng, and Y. Chen · 2014
Later among the works it cites.
BayesOpt: A Bayesian optimization library for nonlinear optimization, experimental design and bandits
R. Martinez-Cantin · 2014
Later among the works it cites.
Cooperating with unknown teammates in complex domains: a robot soccer case study of ad hoc teamwork
S. Barrett and P. Stone · 2015
Later among the works it cites.
Belief and truth in hypothesised behaviours
S. Albrecht, J. Crandall, and S. Ramamoorthy · 2016
Later among the works it cites.
Special issue on multiagent interaction without prior coordination: Guest editorial
S. Albrecht, S. Liemhetcharat, and P. Stone · 2016
Later among the works it cites.
Exploiting causality for selective belief filtering in dynamic Bayesian networks
S. Albrecht and S. Ramamoorthy · 2016
Later among the works it cites.
Interactive POMDPs with finite-state models of other agents
A. Panella and P. Gmytrasiewicz · 2017
Later among the works it cites.