Fetching the paper…
Reading the bibliography…
The existing methods for evaluating the inference abilities of Large Language Models (LLMs) have been predominantly results-centric, making it challenging to assess the inference process comprehensively.
A Logical Calculus of the Ideas Immanent in Nervous Activity
Warren S McCulloch and Walter Pitts. 1943 · 1943
Earlier work this paper cites.
Computing Machinery and Intelligence
Alan Turing. 1950 · 1950
Earlier work this paper cites.
Cybernetics
Norbert Wiener. 1950 · 1950
Earlier work this paper cites.
The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain
Frank Rosenblatt. 1958 · 1958
Earlier work this paper cites.
The Language of Thought . Vol. 5
Jerry A Fodor. 1975 · 1975
Earlier work this paper cites.
Connectionism and Cognitive Architecture: A Critical Analysis
Jerry A Fodor and Zenon W Pylyshyn. 1988 · 1988
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart J Russell and Peter Norvig. 1995 · 1995
Earlier work this paper cites.
Lack of Combinatorial Productivity in Language Processing with Simple Recurrent Networks
Frank van der Velde, Gwendid T van der Voort van der Kleij, and Marc de Kamps. 2004 · 2004
Earlier work this paper cites.
Frames of Mind: The Theory of Multiple Intelligences
Howard E Gardner. 2011 · 2011
Earlier work this paper cites.
MovieQA: Understanding Stories in Movies through Question-Answering. In CVPR . 4631–4640
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016 · 2016
Earlier work this paper cites.
Explanation and Justification in Machine Learning: A Survey. In IJCAI Workshop
Or Biran and Courtenay Cotton. 2017 · 2017
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences. In NeurIPS
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes. In CVPR . 5828–5839
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. 2017 · 2017
Earlier work this paper cites.
TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering. In CVPR . 2758–2766
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Earlier work this paper cites.
Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks. In ICML . PMLR, 2873–2882
Brenden Lake and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
TVQA: Localized, Compositional Video Question Answering. In EMNLP
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018 · 2018
Earlier work this paper cites.
On the Measure of Intelligence
François Chollet. 2019 · 2019
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In NAACL
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs. In NAACL
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Explain Yourself! Leveraging Language Models for Commonsense Reasoning. In ACL
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In NAACL
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Do Multi-Hop Readers Dream of Reasoning Chains?. In EMNLP Workshop on MRQA
Haoyu Wang, Mo Yu, Xiaoxiao Guo, Rajarshi Das, Wenhan Xiong, and Tian Gao. 2019 · 2019
Earlier work this paper cites.
RAVEN: A Dataset for Relational and Analogical Visual Reasoning. In CVPR
Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu. 2019 · 2019
Earlier work this paper cites.
Compositionality Decomposed: How Do Neural Networks Generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020 · 2020
Earlier work this paper cites.
K-BERT: Enabling Language Representation with Knowledge Graphs. In AAAI
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang. 2020 · 2020
Earlier work this paper cites.
Bongard-Logo: A New Benchmark for Human-Level Concept Learning and Reasoning. In NeurIPS
Weili Nie, Zhiding Yu, Lei Mao, Ankit B Patel, Yuke Zhu, and Anima Anandkumar. 2020 · 2020
Cited alongside, same era.
ARC-Solution
Johan Sokrates Wind. 2020 · 2020
Cited alongside, same era.
ARC-Game
Alexey Borsky. 2021 · 2021
Cited alongside, same era.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Cited alongside, same era.
Fast and Flexible: Human Program Induction in Abstract Reasoning Tasks. In CogSci
Aysja Johnson, Wai Keen Vong, Brenden M. Lake, and Todd M. Gureckis. 2021 · 2021
Cited alongside, same era.
STAR: A Benchmark for Situated Reasoning in Real-World Videos. In NeurIPS
Unraveling the ARC Puzzle: Mimicking Human Solutions with Object-Centric Decision Transformer
Jaehyun Park, Jaegyun Im, Sanha Hwang, Mintaek Lim, Sabina Ualibekova, Sejin Kim, and Sundong Kim. 2023 · 2023
Later among the works it cites.
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
Zhaowei Wang, Haochen Shi, Weiqi Wang, Tianqing Fang, Hongming Zhang, Sehyun Choi, Xin Liu, and Yangqiu Song. 2023b · 2023
Later among the works it cites.
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations
Yudong Xu, Wenhao Li, Pashootan Vaezipoor, Scott Sanner, and Elias B Khalil. 2023b · 2023
Later among the works it cites.
Phy-Q as a Measure for Physical Reasoning Intelligence
Cheng Xue, Vimukthini Pinto, Chathura Gamage, Ekaterina Nikonova, Peng Zhang, and Jochen Renz. 2023 · 2023
Later among the works it cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In NeurIPS
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan. 2021 · 2021
Cited alongside, same era.
Communicating Natural Programs to Humans and Machines. In NeurIPS
Samuel Acquaviva, Yewen Pu, Marta Kryven, Theodoros Sechopoulos, Catherine Wong, Gabrielle Ecanow, Maxwell Nye, Michael Tessler, and Joshua B. Tenenbaum. 2022 · 2022
Cited alongside, same era.
Playgrounds for Abstraction and Reasoning. In NeurIPS Workshop on nCSI
Subin Kim, Prin Phunyaphibarn, Donghyun Ahn, and Sundong Kim. 2022 · 2022
Cited alongside, same era.
Solving Quantitative Reasoning Problems with Language Models. In NeurIPS
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022 · 2022
Cited alongside, same era.
Training Language Models to Follow Instructions with Human Feedback. In NeurIPS
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simense, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
FLAVA: A Foundational Language and Vision Alignment Model. In CVPR
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022 · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In NeurIPS
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. In ICLR
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. 2023 · 2023
Later among the works it cites.
A Survey on Evaluation of Large Language Models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024 · 2024
Closest in time.
Show, Don’t Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
Gonçalo Hora de Carvalho, Robert Pollice, and Oscar Knap. 2024 · 2024
Closest in time.
Large Language Models are Not Strong Abstract Reasoners
Gaël Gendron, Qiming Bao, Michael Witbrock, and Gillian Dobbie. 2024 · 2024
Closest in time.
Addressing the Abstraction and Reasoning Corpus via Procedural Example Generation
Michael Hodel. 2024 · 2024
Closest in time.
Generating Images with Multimodal Language Models. In NeurIPS , Vol. 36
Jing Yu Koh, Daniel Fried, and Russ R Salakhutdinov. 2024 · 2024
Closest in time.
ARC-PRIZE Competition
Lab42. 2024 · 2024
Closest in time.
Understanding and Patching Compositional Reasoning in LLMs
Zhaoyi Li, Gangwei Jiang, Hong Xie, Linqi Song, Defu Lian, and Ying Wei. 2024 · 2024
Closest in time.
LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
Mushui Liu, Yuhang Ma, Xinfeng Zhang, Yang Zhen, Zeng Zhao, Zhipeng Hu, Bai Liu, and Changjie Fan. 2024 · 2024
Closest in time.
MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal Models. In ICLR
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2024 · 2024
Closest in time.
An Examination of the Compositionality of Large Generative Vision-Language Models. In NAACL
Teli Ma, Rong Li, and Junwei Liang. 2024 · 2024
Closest in time.
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement. In ICLR
Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi, Nouha Dziri, and Xiang Ren. 2024 · 2024
Closest in time.
From Generation to Selection: Findings of Converting Analogical Problem-Solving into Multiple-Choice Questions. In EMNLP Findings
Donghyeon Shin, Seungpil Lee, Klea Lena Kovacec, and Sundong Kim. 2024 · 2024
Closest in time.
A Survey on Compositional Learning of AI Models: Theoretical and Experimental Practices
Sania Sinha, Tanawan Premsri, and Parisa Kordjamshidi. 2024 · 2024
Closest in time.
LLMs Cannot Find Reasoning Errors, but Can Correct Them Given the Error Location. In ACL Findings
Gladys Tyen, Hassan Mansoor, Victor Cărbune, Yuanzhu Peter Chen, and Tony Mak. 2024 · 2024
Closest in time.
PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change. In NeurIPS
Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati. 2024 · 2024
Closest in time.
Hypothesis Search: Inductive Reasoning with Language Models. In ICLR
Ruocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu, Nick Haber, and Noah D Goodman. 2024 · 2024
Closest in time.
The Generative AI Paradox: “What It Can Create, It May Not Understand”. In ICLR
Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, and Yejin Choi. 2024 · 2024
Closest in time.
Exploring Compositional Generalization of Large Language Models. In NAACL Workshop
Haoran Yang, Hongyuan Lu, Wai Lam, and Deng Cai. 2024 · 2024
Closest in time.
RATT: A Thought Structure for Coherent and Correct LLM Reasoning
Jinghan Zhang, Xiting Wang, Weijieying Ren, Lu Jiang, Dongjie Wang, and Kunpeng Liu. 2024b · 2024
Closest in time.