Fetching the paper…
Reading the bibliography…
Evaluating the reasoning abilities of large language models (LLMs) is challenging.
Labeling images with a computer game
Luis Von Ahn and Laura Dabbish · 2004
Earlier work this paper cites.
Games with a purpose
Luis Von Ahn · 2006
Earlier work this paper cites.
Designing games with a purpose
Luis Von Ahn and Laura Dabbish · 2008
Earlier work this paper cites.
Einstein’s riddle: Riddles, paradoxes, and conundrums to stretch your mind
S Jeremy · 2009
Earlier work this paper cites.
Generalized distances between rankings
Ravi Kumar and Sergei Vassilvitskii · 2010
Earlier work this paper cites.
Encyclopedia of the Sciences of Learning
Norbert M Seel · 2011
Earlier work this paper cites.
Spearman’s rank correlation coefficient
Philip Sedgwick · 2014
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Benchmarking the spectrum of agent capabilities
Danijar Hafner · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in nlp
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al · 2021
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Antonia Creswell, Murray Shanahan, and Irina Higgins · 2022
Earlier work this paper cites.
Human-like property induction is a challenge for large language models
Simon Jerome Han, Keith James Ransom, Andrew Perfors, and Charles Kemp · 2022
Earlier work this paper cites.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi · 2022
Earlier work this paper cites.
A property induction framework for neural language models
Kanishka Misra, Julia Taylor Rayz, and Allyson Ettinger · 2022
Earlier work this paper cites.
A comparison of the pearson, spearman rank and kendall tau correlation coefficients using quantitative variables
Rega Hassan Ali Shiekh and Essam F El-Hashash · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them, 2022
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Language models as inductive reasoners
Zonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, and Furu Wei · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
BIG bench authors · 2023
Cited alongside, same era.
Introducing connect by cloudresearch: Advancing online participant recruitment in the digital age
Rachel Hartman, Aaron J Moss, Shalom Noach Jaffe, Cheskie Rosenzweig, Leib Litman, and Jonathan Robinson · 2023
Cited alongside, same era.
Towards reasoning in large language models: A survey, 2023
Jie Huang and Kevin Chen-Chuan Chang · 2023
Cited alongside, same era.
Let’s verify step by step, 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Self-playing adversarial language game enhances llm reasoning
Pengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang, Yong Dai, Lei Han, and Nan Du · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al · 2024
Closest in time.
Shibo Hao, Yi Gu, Haotian Luo, Tianyang Liu, Xiyan Shao, Xinyuan Wang, Shuhua Xie, Haodi Ma, Adithya Samavedhi, Qiyue Gao, Zhen Wang, and Zhiting Hu · 2024
Closest in time.
Dspy: Compiling declarative language model calls into state-of-the-art pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, Heather Miller, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Agentbench: Evaluating llms as agents, 2023
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang · 2023
Cited alongside, same era.
Eye into ai: Evaluating the interpretability of explainable ai techniques through a game with a purpose
Katelyn Morrison, Mayank Jain, Jessica Hammer, and Adam Perer · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman · 2023
Cited alongside, same era.
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre · 2023
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought, 2023
Abulhair Saparov and He He · 2023
Cited alongside, same era.
A mirror to human question asking: Analyzing the akinator online question game
Gal Sasson and Yoed N Kenett · 2023
Cited alongside, same era.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Cited alongside, same era.
Does style matter? disentangling style and substance in chatbot arena, Aug. 2024
Tianle Li, Anastasios Angelopoulos, and Wei-Lin Chiang · 2024
Closest in time.
Improve mathematical reasoning in language models by automated process supervision, 2024
Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, and Abhinav Rastogi · 2024
Closest in time.
A llm benchmark based on the minecraft builder dialog agent task
Chris Madge and Massimo Poesio · 2024
Closest in time.
Introducing llama 3.1, 2024
MetaAI · 2024
Closest in time.
Mistral-large: Unveiling our latest model, 2024
Mistral · 2024
Closest in time.
Openai o1 system card, Sep. 2024a
OpenAI · 2024
Closest in time.
Gpt-4o system card, 2024b
OpenAI · 2024
Closest in time.
Benchmark agreement testing done right: A guide for llm benchmark evaluation
Yotam Perlitz, Ariel Gera, Ofir Arviv, Asaf Yehudai, Elron Bandel, Eyal Shnarch, Michal Shmueli-Scheuer, and Leshem Choshen · 2024
Closest in time.
Truce: Private benchmarking to prevent contamination and improve comparative evaluation of llms
Tanmay Rajore, Nishanth Chandran, Sunayana Sitaram, Divya Gupta, Rahul Sharma, Kashish Mittal, and Manohar Swaminathan · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-Baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Oguzhan Topsakal, Colby Jacob Edell, and Jackson Bailey Harper · 2024
Closest in time.
Livebench: A challenging, contamination-free llm benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Siddartha Naidu, et al · 2024
Closest in time.
Smartplay : A benchmark for LLMs as intelligent agents
Yue Wu, Xuan Tang, Tom Mitchell, and Yuanzhi Li · 2024
Closest in time.
Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Deli Zhao, and Lidong Bing · 2024
Closest in time.