Fetching the paper…
Reading the bibliography…
In this paper, we study the class of games known as hidden-role games in which players are assigned privately to teams and are faced with the challenge of recognizing and cooperating with teammates.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Equilibrium points in n-person games
John Nash · 1950
Earlier work this paper cites.
Game with incomplete information played by Bayesian players
John Harsanyi · 1967
Earlier work this paper cites.
An approach to communication equilibria
Francoise Forges · 1986
Earlier work this paper cites.
Multistage games with communication
Roger B Myerson · 1986
Earlier work this paper cites.
Verifiable secret sharing and multiparty protocols with honest majority
Tal Rabin and Michael Ben-Or · 1989
Earlier work this paper cites.
Multiparty protocols tolerating half faulty processors
Donald Beaver · 1990
Earlier work this paper cites.
The complexity of two-person zero-sum games in extensive form
Daphne Koller and Nimrod Megiddo · 1992
Earlier work this paper cites.
Fast algorithms for finding randomized strategies in game trees
Daphne Koller, Nimrod Megiddo, and Bernhard von Stengel · 1994
Earlier work this paper cites.
Robust sharing of secrets when the dealer is honest or cheating
Tal Rabin · 1994
Earlier work this paper cites.
Efficient computation of behavior strategies
Bernhard von Stengel · 1996
Earlier work this paper cites.
Team-maxmin equilibria
Bernhard von Stengel and Daphne Koller · 1997
Earlier work this paper cites.
Some optimal inapproximability results
Johan Håstad · 2001
Earlier work this paper cites.
Computational complexity and communication: Coordination in two–player games
Amparo Urbano and Jose E Vila · 2002
Earlier work this paper cites.
Rational secure computation and ideal mechanism design
Sergei Izmalkov, Silvio Micali, and Matt Lepinski · 2005
Earlier work this paper cites.
Distributed computing meets game theory: robust mechanisms for rational secret sharing and multiparty computation
Ittai Abraham, Danny Dolev, Rica Gonen, and Joe Halpern · 2006
Cited alongside, same era.
Mafia: A theoretical study of players and coalitions in a partial information environment
Mark Braverman, Omid Etesami, and Elchanan Mossel · 2006
Cited alongside, same era.
Cooperative security in distributed sensor networks
Oscar Garcia Morchon, Heribert Baldus, Tobias Heer, and Klaus Wehrle · 2007
Cited alongside, same era.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Bowling, Michael Johanson, and Carmelo Piccione · 2007
Cited alongside, same era.
Bridging game theory and cryptography: Recent results and future directions
Jonathan Katz · 2008
Cited alongside, same era.
Cooperative security in distributed networks
Secure multiparty computation (mpc)
Yehuda Lindell · 2020
Later among the works it cites.
Game Theory
Michael Maschler, Shmuel Zamir, and Eilon Solan · 2020
Later among the works it cites.
Learning to deceive in multi-agent hidden role games
Matthew Aitchison, Lyndon Benke, and Penny Sweetser · 2021
Later among the works it cites.
Faster game solving via predictive Blackwell approachability: Connecting regret matching and mirror descent
Gabriele Farina, Christian Kroer, and Tuomas Sandholm · 2021
Later among the works it cites.
Chess as a testing grounds for the oracle approach to AI safety
James D. Miller, Roman Yampolskiy, Olle Häggström, and Stuart Armstrong · 2021
Later among the works it cites.
A survey on security and privacy of federated learning
Viraaji Mothukuri, Reza M Parizi, Seyedamin Pouriyeh, Yan Huang, Ali Dehghantanha, and Gautam Srivastava · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oscar Garcia-Morchon, Dmitriy Kuptsov, Andrei Gurtov, and Klaus Wehrle · 2013
Cited alongside, same era.
Team-maxmin equilibrium: efficiency bounds and algorithms
Nicola Basilico, Andrea Celli, Giuseppe De Nittis, and Nicola Gatti · 2017
Cited alongside, same era.
Correlation and unmediated cheap talk in repeated games with imperfect monitoring
Heng Liu · 2017
Cited alongside, same era.
Computational results for extensive-form adversarial team games
Andrea Celli and Nicola Gatti · 2018
Cited alongside, same era.
Solving Avalon with whispering
Paul F. Christiano · 2018
Cited alongside, same era.
Simplifying Avalon
Paul F. Christiano · 2018
Cited alongside, same era.
Practical exact algorithm for trembling-hand equilibrium refinements in games
Gabriele Farina, Nicola Gatti, and Tuomas Sandholm · 2018
Cited alongside, same era.
Hidden agenda: a social deduction game with diverse learned equilibria
Kavya Kopparapu, Edgar A. Duéñez-Guzmán, Jayd Matyas, Alexander Sasha Vezhnevets, John P. Agapiou, Kevin R. McKee, Richard Everett, Janusz Marecki, Joel Z. Leibo, and Thore Graepel · 2022
Later among the works it cites.
An integrated approach of designing functionality with security for distributed cyber-physical systems
Dipty Tripathi, Amit Biswas, Anil Kumar Tripathi, Lalit Kumar Singh, and Amrita Chaturvedi · 2022
Later among the works it cites.
Polynomial-time optimal equilibria with a mediator in extensive-form games
Brian Hu Zhang and Tuomas Sandholm · 2022
Later among the works it cites.
AI alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao · 2023
Closest in time.
Hoodwinked: Deception and cooperation in a text-based game for language models
Aidan O’Gara · 2023
Closest in time.
AI deception: A survey of examples, risks, and potential solutions
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2023
Closest in time.
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn · 2023
Closest in time.
Language agents with reinforcement learning for strategic play in the werewolf game
Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu · 2023
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.