Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, yet their performance is highly dependent on the prompting strategy and model scale.
Hybridqa: A dataset of multi-hop question answering over tabular and textual data, 2021
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Wang · 2004
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Earlier work this paper cites.
Gpt-4 technical report
OpenAI · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2023
Earlier work this paper cites.
Explaining black box text modules in natural language with language models, 2023
Chandan Singh, Aliyah R. Hsu, Richard Antonello, Shailee Jain, Alexander G. Huth, Bin Yu, and Jianfeng Gao · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Earlier work this paper cites.
multi-agent-llm-0.1.3, 2024
Github AgnostiqHQ · 2024
Earlier work this paper cites.
Llm-generated black-box explanations can be adversarially helpful, 2024
Rohan Ajwani, Shashidhar Reddy Javaji, Frank Rudzicz, and Zining Zhu · 2024
Earlier work this paper cites.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al · 2024
Earlier work this paper cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Meta reasoning for large language models, 2024
Peizhong Gao, Ao Xie, Shaoguang Mao, Wenshan Wu, Yan Xia, Haipeng Mi, and Furu Wei · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Cited alongside, same era.
Understanding the effects of iterative prompting on truthfulness, 2024
Satyapriya Krishna, Chirag Agarwal, and Himabindu Lakkaraju · 2024
Cited alongside, same era.
Iteration of thought: Leveraging inner dialogue for autonomous large language model reasoning, 2024
Santosh Kumar Radha, Yasamin Nouri Jelyani, Ara Ghukasyan, and Oktay Goktas · 2024
Later among the works it cites.
Morehopqa: More than multi-hop reasoning, 2024
Julian Schnitzler, Xanh Ho, Jiahao Huang, Florian Boudin, Saku Sugawara, and Akiko Aizawa · 2024
Later among the works it cites.
BBox-adapter: Lightweight adapting for black-box large language models
Haotian Sun, Yuchen Zhuang, Wei Wei, Chao Zhang, and Bo Dai · 2024
Later among the works it cites.
A comprehensive survey of hallucination mitigation techniques in large language models, 2024
S. M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, and Deheng Ye · 2024
Cited alongside, same era.
Dgot: Dynamic graph of thoughts for scientific abstract generation, 2024
Xinyu Ning, Yutong Zhao, Yitong Liu, and Hongwen Yang · 2024
Cited alongside, same era.
OpenAI · 2024
Cited alongside, same era.
Introducing openai o1-preview
OpenAI · 2024
Cited alongside, same era.
Adapt: As-needed decomposition and planning with language models, 2024
Archiki Prasad, Alexander Koller, Mareike Hartmann, Peter Clark, Ashish Sabharwal, Mohit Bansal, and Tushar Khot · 2024
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Rewon Child Jeffrey Wu, Dario Amodei David Luan, and Ilya Sutskever · 2024
Cited alongside, same era.
Santosh Kumar Radha and Oktay Goktas · 2024
Cited alongside, same era.
Reinforcement learning enhanced llms: A survey, 2024a
Shuhe Wang, Shengyu Zhang, Jie Zhang, Runyi Hu, Xiaoya Li, Tianwei Zhang, Jiwei Li, Fei Wu, Guoyin Wang, and Eduard Hovy
Cited in the paper.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Redel: A toolkit for llm-powered recursive multi-agent systems, 2024
Andrew Zhu, Liam Dugan, and Chris Callison-Burch · 2024
Later among the works it cites.
Hydra: Model factorization framework for black-box llm personalization, 2024
Yuchen Zhuang, Haotian Sun, Yue Yu, Rushi Qiang, Qifan Wang, Chao Zhang, and Bo Dai · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, and Peiyi Wang et al · 2025
Closest in time.
Search-o1: Agentic search-enhanced large reasoning models, 2025
Xiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang, Yujia Zhou, Yutao Zhu, Peitian Zhang, and Zhicheng Dou · 2025
Closest in time.
On the reasoning capacity of ai models and how to quantify it
Santosh Kumar Radha and Oktay Goktas · 2025
Closest in time.