Fetching the paper…
Reading the bibliography…
Numerous decision-making tasks require estimating causal effects under interventions on different parts of a system.
Seeing versus doing: two modes of accessing causal knowledge
Michael R Waldmann and York Hagmayer · 2005
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Distinguishing cause from effect using observational data: methods and benchmarks
Joris M Mooij, Jonas Peters, Dominik Janzing, Jakob Zscheischler, and Bernhard Schölkopf · 2016
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Earlier work this paper cites.
Using cognitive psychology to understand GPT-3
Marcel Binz and Eric Schulz · 2023
Cited alongside, same era.
Is knowledge all large language models needed for causal reasoning?
Hengrui Cai, Shengjie Liu, and Rui Song · 2023
Cited alongside, same era.
Can large language models infer causation from correlation?
Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona Diab, and Bernhard Schölkopf · 2023
Cited alongside, same era.
Gpt-4 technical report. arxiv 2303.08774
R OpenAI · 2023
Cited alongside, same era.
Arkil Patel, Satwik Bhattamishra, Siva Reddy, and Dzmitry Bahdanau · 2023
Cited alongside, same era.
Workarena: How capable are web agents at solving common knowledge work tasks?
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste · 2024
Closest in time.
Cladder: A benchmark to assess causal reasoning capabilities of language models
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng Lyu, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, et al · 2024
Closest in time.
Gpt-4 passes the bar exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2024
Closest in time.
Causal reasoning and large language models: Opening a new frontier for causality
Emre Kiciman, Robert Ness, Amit Sharma, and Chenhao Tan · 2024
Closest in time.
Xiao Liu, Zirui Wu, Xueqing Wu, Pan Lu, Kai-Wei Chang, and Yansong Feng · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.