Fetching the paper…
Reading the bibliography…
As AI systems increasingly assume roles where trust and alignment with human values are essential, understanding when and why they engage in deception has become a critical research priority.
Signal detection theory and psychophysics
D. M. Green and J. A. Swets · 1966
Earlier work this paper cites.
The market for "lemons": Quality uncertainty and the market mechanism
G. A. Akerlof · 1970
Earlier work this paper cites.
Job market signaling
M. Spence · 1973
Earlier work this paper cites.
Reexamination of the perfectness concept for equilibrium points in extensive games
R. Selten · 1975
Earlier work this paper cites.
Mate selection-a selection for a handicap
A. Zahavi · 1975
Earlier work this paper cites.
The selfish gene
R. Dawkins · 1976
Earlier work this paper cites.
The false consensus effect: An egocentric bias in social perception and attribution processes
L. Ross, D. Greene, and P. House · 1977
Earlier work this paper cites.
Metacognition and cognitive monitoring: A new area of cognitive–developmental inquiry
J. H. Flavell · 1979
Earlier work this paper cites.
Strategic information transmission
V. P. Crawford and J. Sobel · 1982
Earlier work this paper cites.
Evolution and the theory of games
J. M. Smith · 1982
Earlier work this paper cites.
Telling ingratiating lies: Effects of target sex and target attractiveness on verbal and nonverbal deceptive success
B. M. DePaulo, J. I. Stone, and G. D. Lassiter · 1985
Earlier work this paper cites.
Social structures: A network approach
B. Wellman and S. D. Berkowitz · 1988
Earlier work this paper cites.
The electronic mail game: Strategic behavior under almost common knowledge
A. Rubinstein · 1989
Earlier work this paper cites.
The coherence theory of truth: Realism, anti-realism, idealism
R. C. Walker · 1989
Earlier work this paper cites.
Game theory: analysis of conflict
R. B. Myerson · 1991
Earlier work this paper cites.
A theory of fads, fashion, custom, and cultural change as informational cascades
S. Bikhchandani, D. Hirshleifer, and I. Welch · 1992
Earlier work this paper cites.
Interpersonal deception theory
D. B. Buller and J. K. Burgoon · 1996
Earlier work this paper cites.
Cheap talk
J. Farrell and M. Rabin · 1996
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Social dilemmas: The anatomy of cooperation
P. Kollock · 1998
Earlier work this paper cites.
Paranoid cognition in social systems: thinking and acting in the shadow of doubt
R. M. Kramer · 1998
Earlier work this paper cites.
Managerial incentive problems: A dynamic perspective
B. Holmstrom · 1999
Earlier work this paper cites.
Detecting deceit via analysis of verbal and nonverbal behavior
A. Vrij, K. Edward, K. Roberts, and R. Bull · 2000
Earlier work this paper cites.
Information and competition in us forest service timber auctions
S. Athey and J. Levin · 2001
Cited alongside, same era.
Deception: The role of consequences
U. Gneezy · 2005
Cited alongside, same era.
Accuracy of deception judgments
C. F. Bond Jr and B. M. DePaulo · 2006
Cited alongside, same era.
Models of interpersonal trust development: Theoretical approaches, empirical evidence, and future directions
R. J. Lewicki, E. C. Tomlinson, and N. Gillespie · 2006
Cited alongside, same era.
A survey of trust and reputation systems for online service provision
A. Jøsang, R. Ismail, and C. Boyd · 2007
Cited alongside, same era.
Games and information
E. Rasmusen · 2007
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Later among the works it cites.
Decentralised learning in systems with many, many strategic agents
D. Mguni, J. Jennings, and E. M. de Cote · 2018
Later among the works it cites.
On the utility of learning about humans for human-ai coordination
M. Carroll, R. Shah, M. K. Ho, T. Griffiths, S. Seshia, P. Abbeel, and A. Dragan · 2019
Later among the works it cites.
Risks from learned optimization in advanced machine learning systems
E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant · 2019
Later among the works it cites.
The alignment problem: Machine learning and human values
B. Christian · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. M. Omohundro · 2008
Cited alongside, same era.
Political Institutions and Generalized Trust
D. Stolle and B. Rothstein · 2008
Cited alongside, same era.
Network analysis in the social sciences
S. P. Borgatti, A. Mehra, D. J. Brass, and G. Labianca · 2009
Cited alongside, same era.
Hidden incentives for auto-induced distributional shift
D. Krueger, T. Maharaj, and J. Leike · 2009
Cited alongside, same era.
Causality
J. Pearl · 2009
Cited alongside, same era.
Information manipulation theory 2: A propositional theory of deceptive discourse production
S. A. M. Jack, K. Morrison, J. E. Paik, A. M. Wisner, and X. Zhu · 2010
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg · 2020
Later among the works it cites.
Cooperative ai: machines must learn to find common ground
A. Dafoe, Y. Bachrach, G. Hadfield, E. Horvitz, K. Larson, and T. Graepel · 2021
Later among the works it cites.
Artificial intelligence act
European Commission · 2021
Later among the works it cites.
Truthful ai: Developing and governing ai that does not lie
O. Evans, O. Cotton-Barratt, L. Finnveden, A. Bales, A. Balwit, P. Wills, L. Righetti, and W. Saunders · 2021
Later among the works it cites.
Language models as agent models
J. Andreas · 2022
Later among the works it cites.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, A. P. Jacob, et al · 2022
Later among the works it cites.
X-risk analysis for ai research
D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt · 2022
Later among the works it cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
J. Menick, M. Trebacz, V. Mikulik, J. Aslanides, F. Song, M. Chadwick, M. Glaese, S. Young, L. Campbell-Gillingham, G. Irving, and N. McAleese · 2022
Later among the works it cites.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Meta Fundamental AI Research Diplomacy Team (FAIR), A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, A. P. Jacob, M. Komeili, K. Konath, M. Kwon, A. Lerer, M. Lewis, A. H. Miller, S. Mitts, A. Renduchintala, S. Roller, D. Rowe, W. Shi, J. Spisak, A. Wei, D. Wu, H. Zhang, and M. Zijlstra · 2022
Later among the works it cites.
Avalonbench: Evaluating llms playing the game of avalon, 2023
J. Light, M. Cai, S. Shen, and Z. Hu · 2023
Later among the works it cites.
Avalon’s game of thoughts: Battle against deception through recursive contemplation
S. Wang, C. Liu, Z. Zheng, S. Qi, S. Chen, Q. Yang, A. Zhao, C. Wang, S. Song, and G. Huang · 2023
Later among the works it cites.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica · 2024
Later among the works it cites.
Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents
G. Piatti, Z. Jin, M. Kleiman-Weiner, B. Schölkopf, M. Sachan, and R. Mihalcea · 2024
Later among the works it cites.
Persuasion games using large language models, 2024
G. P. Ramani, S. Karande, S. V, and Y. Bhatia · 2024
Later among the works it cites.
Among them: A game-based framework for assessing persuasion capabilities of llms
M. Idziejczak, V. Korzavatykh, M. Stawicki, A. Chmutov, M. Korcz, I. Błądek, and D. Brzezinski · 2025
Closest in time.
Racing through a minefield: the ai deployment problem, 2022
H. Karnofsky · 2025
Closest in time.