Fetching the paper…
Reading the bibliography…
We introduce a dataset of natural-language questions in the decision theory of so-called Newcomb-like problems.
Newcomb’s problem and two principles of choice
Robert Nozick · 1969
Earlier work this paper cites.
Newcomb’s problem and prisoners’ dilemma
Steven J. Brams · 1975
Earlier work this paper cites.
Counterfactuals and two kinds of expected utility
Allan Gibbard and William L Harper · 1978
Earlier work this paper cites.
Prisoners’ dilemma is a Newcomb problem
David Lewis · 1979
Earlier work this paper cites.
Counterfactuals and Two Kinds of Expected Utility , pages 153–190
Allan Gibbard and William L. Harper · 1981
Earlier work this paper cites.
Rational decision and causality
Ellery Eells · 1982
Earlier work this paper cites.
Causal decision theory
Brian Skyrms · 1982
Earlier work this paper cites.
In the neighbourhood of the Newcomb-predictor (reflections on rationality)
David Gauthier · 1989
Earlier work this paper cites.
Ensuring two bird deaths with one throw
John Leslie · 1991
Earlier work this paper cites.
On the interpretation of decision problems with imperfect recall
Michele Piccione and Ariel Rubinstein · 1997
Earlier work this paper cites.
The foundations of causal decision theory
James M Joyce · 1999
Earlier work this paper cites.
Self-locating belief and the sleeping beauty problem
Adam Elga · 2000
Earlier work this paper cites.
An indirect-evolution approach to Newcomb’s problem, 2001
Max Albert and Ronald Asher Heiner · 2001
Earlier work this paper cites.
Thinking inside the boxes
Jim Holt · 2002
Earlier work this paper cites.
Program equilibrium
Moshe Tennenholtz · 2004
Earlier work this paper cites.
Good and real: Demystifying paradoxes from physics to ethics
Gary L Drescher · 2006
Earlier work this paper cites.
No regrets, or: Edith piaf revamps decision theory
Frank Arntzenius · 2008
Earlier work this paper cites.
Towards a new decision theory, 08 2009
Wei Dai · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Putting a value on Beauty
Rachael Briggs · 2010
Earlier work this paper cites.
Causation, decision theory, and Bell’s theorem: A quantum analogue of the Newcomb problem
Eric G. Cavalcanti · 2010
Earlier work this paper cites.
Binding and its consequences
Christopher J. G. Meacham · 2010
Earlier work this paper cites.
Causal decision theory and determinism, 2010
Alexander Pruss · 2010
Earlier work this paper cites.
Infinite ethics
Nick Bostrom · 2011
Earlier work this paper cites.
Evidence, decision and causality
Arif Ahmed · 2014
Earlier work this paper cites.
Robust cooperation in the prisoner’s dilemma: Program equilibrium via provability logic, 1 2014
Mihaly Barasz, Paul Christiano, Benja Fallenstein, Marcello Herreshoff, Patrick LaVictoire, and Eliezer Yudkowsky · 2014
Earlier work this paper cites.
A dutch book against sleeping beauties who are evidential decision theorists
Vincent Conitzer · 2015
Earlier work this paper cites.
Lost memories and useless coins: revisiting the absentminded driver
Wolfgang Schwarz · 2015
Earlier work this paper cites.
Toward idealized decision theory
Nate Soares and Benja Fallenstein · 2015
Earlier work this paper cites.
Causation and correlation part 1: Evidential decision theory is correct, 09 2010b
Paul Almond · 2016
Earlier work this paper cites.
Smokers, psychos, and decision-theoretic uncertainty
William MacAskill · 2016
Earlier work this paper cites.
Reinforcement learning in conflicting environments for autonomous vehicles
Dominik Mayer, Johannes Feldmaier, and Hao Shen · 2016
Earlier work this paper cites.
Newcomb’s problem, 2016
Eliezer Yudkowsky · 2016
Earlier work this paper cites.
Wei Dai’s updateless decision theory
Tyrrell McAllister · 2017
Earlier work this paper cites.
Multiverse-wide cooperation via correlated decision making
Caspar Oesterheld · 2017
Earlier work this paper cites.
An introduction to decision theory
Martin Peterson · 2017
Cited alongside, same era.
Agent foundations for aligning machine intelligence with human interests: a technical research agenda
Nate Soares and Benya Fallenstein · 2017
Cited alongside, same era.
“Betting on the past” by Arif Ahmed, 02 2017
Johannes Treutlein · 2017
Cited alongside, same era.
EDT vs CDT, 2018a
Paul Christiano · 2018
Cited alongside, same era.
EDT vs CDT 2: conditioning on the impossible, 2018b
Paul Christiano · 2018
Cited alongside, same era.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Cited alongside, same era.
AI in the gray: Exploring moderation policies in dialogic large language models vs. human answers in controversial topics
Vahid Ghafouri, Vibhor Agarwal, Yong Zhang, Nishanth Sastry, Jose Such, and Guillermo Suarez-Tangil · 2023
Later among the works it cites.
Conditioning predictive models: Risks and strategies
Evan Hubinger, Adam Jermyn, Johannes Treutlein, Rubi Hudson, and Kate Woolverton · 2023
Later among the works it cites.
Reconciling evidential and causal decision theory
Simon Huttegger · 2023
Later among the works it cites.
Co-writing with opinionated language models affects users’ views
Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman · 2023
Later among the works it cites.
Large language models as superpositions of cultural perspectives
Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding the tickle defense in decision theory
Caspar Oesterheld · 2018
Cited alongside, same era.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston · 2018
Cited alongside, same era.
Causal decision theory and decision instability
Brad Armendt · 2019
Cited alongside, same era.
Designing preferences, beliefs, and identities for artificial intelligence
Vincent Conitzer · 2019
Cited alongside, same era.
Robust program equilibrium
Caspar Oesterheld · 2019
Cited alongside, same era.
Summary of evidence, decision, and causality, 2020
Dawn Drescher · 2020
Cited alongside, same era.
Later among the works it cites.
Game theory with simulation of other players
Vojtěch Kovařík, Caspar Oesterheld, and Vincent Conitzer · 2023
Later among the works it cites.
Moca: Measuring human-language model alignment on causal and moral judgment tasks
Allen Nie, Yuhui Zhang, Atharva Shailesh Amdekar, Chris Piech, Tatsunori B Hashimoto, and Tobias Gerstenberg · 2023
Later among the works it cites.
Capabilities of GPT-4 on medical challenge problems
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz · 2023
Later among the works it cites.
A theory of bounded inductive rationality
Caspar Oesterheld, Abram Demski, and Vincent Conitzer · 2023
Later among the works it cites.
GPQA: A graduate-level Google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2023
Later among the works it cites.
Preventing language models from hiding their reasoning
Fabien Roger and Ryan Greenblatt · 2023
Later among the works it cites.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Later among the works it cites.
The computational complexity of single-player imperfect-recall games
Emanuel Tewolde, Caspar Oesterheld, Vincent Conitzer, and Paul W Goldberg · 2023
Later among the works it cites.
Google chatbot refused to say whether Elon Musk is better than Adolf Hitler, 2 2024
Noor Al-Sibai · 2024
Closest in time.
Understanding the interplay of scale, data, and bias in language models: A case study with BERT
Muhammad Ali, Swetasudha Panda, Qinlan Shen, Michael Wick, and Ari Kobren · 2024
Closest in time.
Claude 3.5 Sonnet, 6 2024
Anthropic · 2024
Closest in time.
Why the unexpected? Dissecting the political and economic bias in Persian small and large language models
Ehsan Barkhordar, Surendrabikram Thapa, Ashwarya Maratha, and Usman Naseem · 2024
Closest in time.
Chatbot arena: An open platform for evaluating LLMs by human preference, 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
Can cdt rationalise the ex ante optimal policy via modified anthropics?
Emery Cooper, Caspar Oesterheld, and Vincent Conitzer · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2024
Closest in time.
AI control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger · 2024
Closest in time.
GPT-4 passes the bar exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2024
Closest in time.
Recursive joint simulation in games
Vojtech Kovarik, Caspar Oesterheld, and Vincent Conitzer · 2024
Closest in time.
ExpertQA: Expert-curated questions and attributed answers
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth · 2024
Closest in time.
Manuel Mondal, Ljiljana Dolamic, Gérôme Bovet, and Philippe Cudré-Mauroux · 2024
Closest in time.
Are large language models consistent over value-laden questions?
Jared Moore, Tanvi Deshpande, and Diyi Yang · 2024
Closest in time.
Similarity-based cooperative equilibrium
Caspar Oesterheld, Johannes Treutlein, Roger B Grosse, Vincent Conitzer, and Jakob Foerster · 2024
Closest in time.
Cooperate or collapse: Emergence of sustainability behaviors in a society of LLM agents
Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Schölkopf, Mrinmaya Sachan, and Rada Mihalcea · 2024
Closest in time.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, et al · 2024
Closest in time.
You don’t need a personality test to know these models are unreliable: Assessing the reliability of large language models on psychometric instruments
Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, and David Jurgens · 2024
Closest in time.
Imperfect-recall games: Equilibrium concepts and their complexity
Emanuel Tewolde, Brian Zhang, Caspar Oesterheld, Manolis Zampetakis, Tuomas Sandholm, Paul Goldberg, and Vincent Conitzer · 2024
Closest in time.
Do LLMs exhibit human-like response biases? a case study in survey design
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Closest in time.
Evaluating llm reasoning in the operations research domain with orqa
Mahdi Mostajabdaveh, Timothy T Yu, Samarendra Chandan Bindu Dash, Rindranirina Ramamonjison, Jabo Serge Byusa, Giuseppe Carenini, Zirui Zhou, and Yong Zhang · 2025
Closest in time.