Fetching the paper…
Reading the bibliography…
How can one detect friendly and adversarial behavior from raw data? Detecting whether an environment is a friend, a foe, or anything in between, remains a poorly understood yet desirable ability for safe and robust agents.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Theory of games and economic behavior, 2nd rev
J. Von Neumann and O. Morgenstern · 1947
Earlier work this paper cites.
A course in game theory
M. J. Osborne and A. Rubinstein · 1994
Earlier work this paper cites.
Quantal Response Equilibria for Normal Form Games
R. McKelvey and T. Palfrey · 1995
Earlier work this paper cites.
Rationality and intelligence
S. Russell · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Data clustering by markovian relaxation and the information bottleneck method
N. Tishby and N. Slonim · 2001
Earlier work this paper cites.
Friend-or-Foe Q-Learning in General-Sum Games
M. L. Littman · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 2002
Earlier work this paper cites.
Correlated Q-Learning
A. Greenwald and K. Hall · 2003
Earlier work this paper cites.
Clustering with bregman divergences
A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh · 2005
Earlier work this paper cites.
Information bottleneck for Gaussian variables
G. Chechik, A. Globerson, N. Tishby, and Y. Weiss · 2005
Earlier work this paper cites.
New criteria and a new algorithm for learning in multi-agent systems
R. Powers and Y. Shoham · 2005
Cited alongside, same era.
Risk sensitive path integral control
B. van den Broek, W. Wiegerinck, and B. Kappen · 2010
Cited alongside, same era.
Information theory of decisions and actions
N. Tishby and D. Polani · 2011
Cited alongside, same era.
Information, utility and bounded rationality
P. A. Ortega and D. A. Braun · 2011
Cited alongside, same era.
Learning to compete, coordinate, and cooperate in repeated games using reinforcement learning
J. W. Crandall and M. A. Goodrich · 2011
Cited alongside, same era.
An empirical evaluation of Thompson Sampling
O. Chappelle and L. Li · 2011
Cited alongside, same era.
One practical algorithm for both stochastic and adversarial bandits
Y. Seldin and A. Silvkins · 2014
Later among the works it cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Later among the works it cites.
Reactive bandits with attitude
P. A. Ortega, K.-E. Kim, and D. D. Lee · 2015
Later among the works it cites.
Bounded Rationality, Abstraction, and Hierarchical Decision-Making: An Information-Theoretic Optimality Principle
T. Genewein, F. Leibfried, J. Grau-Moya, and D. A. Braun · 2015
Later among the works it cites.
Concrete Problems in AI Safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Hysteresis effects of changing the parameters of noncooperative games
D. H. Wolpert, M. Harré, E. Olbrich, N. Bertschinger, and J. Jost · 2012
Cited alongside, same era.
The best of both worlds: stochastic and adversarial bandits
S. Bubeck and A. Slivkins · 2012
Cited alongside, same era.
Thermodynamics as a theory of decision-making with information-processing costs
P. A. Ortega and D. A. Braun · 2013
Cited alongside, same era.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Cited alongside, same era.
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
P. Auer and C. Chao-Kai · 2016
Later among the works it cites.
J. Leike, M. Martic, V. Krakovna, P.A Ortega, T. Everitt, L. Orseau, and S. Legg · 2017
Later among the works it cites.
The numerics of gans
Lars Mescheder, Sebastian Nowozin, and Andreas Geiger · 2017
Later among the works it cites.
Learning against sequential opponents in repeated stochastic games
P. Hernandez-Leal and M. Kaisers · 2017
Later among the works it cites.
The Mechanics of n-Player Differentiable Games
D. Balduzzi, S. Racaniere, J. Martens, J. Foerster, and T. Tuyls, K. Graepel · 2018
Closest in time.
What game are we playing? end-to-end learning in normal and extensive form games
C. K. Ling, F. Fang, and J. Z. Kolter · 2018
Closest in time.