CRASS: A novel data set and benchmark to test counterfactual reasoning of large language models
Jörg Frohberg and Frank Binder · 2022
Later among the works it cites.
Double prevention, causal judgments, and counterfactuals
Paul Henne and Kevin O’Neill · 2022
Later among the works it cites.
Investigating causal understanding in llms
Marius Hobbhahn, Tom Lieberum, and David Seiler · 2022
Later among the works it cites.
Unsuitability of notears for causal graph discovery when dealing with dimensional quantities
Marcus Kaiser and Maksim Sipos · 2022
Later among the works it cites.
Towards understanding how machines can learn causal overhypotheses
Original
Eliza Kosoy, David M Chan, Adrian Liu, Jasmine Collins, Bryanna Kaufmann, Sandy Han Huang, Jessica B Hamrick, John Canny, Nan Rosemary Ke, and Alison Gopnik · 2022
Later among the works it cites.
How come gpt can seem so brilliant one minute and so breathtakingly dumb the next?, 2022
Gary Marcus · 2022
Later among the works it cites.
An empirical evaluation of github copilot’s code suggestions
Nhan Nguyen and Sarah Nadi · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
The counterfactual-shapley value: Attributing change in system metrics
Original
Amit Sharma, Hua Li, and Jian Jiao · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Original
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Original
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, , and Jason Wei · 2022
Later among the works it cites.
Can foundation models talk causality?
Original
Moritz Willig, Matej Zečević, Devendra Singh Dhami, and Kristian Kersting · 2022
Later among the works it cites.
Can large language models distinguish cause from effect?
LYU Zhiheng, Zhijing Jin, Rada Mihalcea, Mrinmaya Sachan, and Bernhard Schölkopf · 2022
Later among the works it cites.
Large language models are human-level prompt engineers
Original
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2022
Later among the works it cites.
Causal modelling agents: Causal graph discovery through synergising metadata-and data-driven reasoning
Ahmed Abdulaal, Nina Montana-Brown, Tiantian He, Ayodeji Ijishakin, Ivana Drobnjak, Daniel C Castro, Daniel C Alexander, et al · 2023
Closest in time.
Aml-babel benchmark: Causal judgement
AML-BABEL · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Original
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
Reasoning with language model is planning with world model
Original
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Closest in time.
Gpt-4 passes the bar exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2023
Closest in time.
Counterfactual reasoning: Testing language models’ understanding of hypothetical scenarios
Jiaxuan Li, Lang Yu, and Allyson Ettinger · 2023
Closest in time.
Evaluating the logical reasoning ability of chatgpt and gpt-4, 2023
Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, and Yue Zhang · 2023
Closest in time.
Can large language models build causal graphs?
Original
Stephanie Long, Tibor Schuster, Alexandre Piché, ServiceNow Research, et al · 2023
Closest in time.
Modeling covid-19 disease processes by remote elicitation of causal bayesian networks from medical experts
Steven Mascaro, Yue Wu, Owen Woodberry, Erik P Nyberg, Ross Pearson, Jessica A Ramsay, Ariel O Mace, David A Foley, Thomas L Snelling, Ann E Nicholson, et al · 2023
Closest in time.
Capabilities of gpt-4 on medical challenge problems
Original
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Original
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Closest in time.
Causal-discovery performance of chatgpt in the context of neuropathic pain diagnosis
Original
Ruibo Tu, Chao Ma, and Cheng Zhang · 2023
Closest in time.
Causal parrots: Large language models may talk causality but are not causal
MORITZ Willig, MATEJ ZEČEVIĆ, DEVENDRA SINGH DHAMI, and KRISTIAN KERSTING · 2023
Closest in time.
Understanding causality with large language models: Feasibility and opportunities, 2023
Cheng Zhang, Stefan Bauer, Paul Bennett, Jiangfeng Gao, Wenbo Gong, Agrin Hilmkil, Joel Jennings, Chao Ma, Tom Minka, Nick Pawlowski, and James Vaughan · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2023
Closest in time.
Elephants never forget: Memorization and learning of tabular data in large language models, 2024
Original
Sebastian Bordt, Harsha Nori, Vanessa Rodrigues, Besmira Nushi, and Rich Caruana · 2024
Closest in time.