Fetching the paper…
Reading the bibliography…
Predicting the performance of LLMs on individual task instances is essential to ensure their reliability in high-stakes applications.
The varimax criterion for analytic rotation in factor analysis
Henry F Kaiser. 1958 · 1958
Earlier work this paper cites.
Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning
Andrew S. Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2011 · 2011
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space. In International Conference on Learning Representations
Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Abductive Commonsense Reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Yih, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning. In Conference on Empirical Methods in Natural Language Processing
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Reasoning Over Paragraph Effects in Situations. In Conference on Empirical Methods in Natural Language Processing
Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Adversarial NLI: A New Benchmark for Natural Language Understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019 · 2019
Earlier work this paper cites.
ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020 · 2020
Earlier work this paper cites.
Training on the test set: Mapping the system-problem space in AI. In Proceedings of the AAAI conference on artificial intelligence , Vol. 36. 12256–12261
José Hernández-Orallo, Wout Schellaert, and Fernando Martínez-Plumed. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Wanli: Worker and ai collaboration for natural language inference dataset creation
Alisa Liu, Swabha Swayamdipta, Noah A Smith, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Earlier work this paper cites.
Reject Before You Run: Small Assessors Anticipate Big Language Models.. In EBeM@ IJCAI
Lexin Zhou, Fernando Martínez-Plumed, José Hernández-Orallo, Cèsar Ferri, and Wout Schellaert. 2022 · 2022
Cited alongside, same era.
SpaceNLI: Evaluating the Consistency of Predicting Inferences In Space
Lasha Abzianidze, Joost Zwarts, and Yoad Winter. 2023 · 2023
Cited alongside, same era.
Proposal for a Regulation of the European Parliament and of the Council on Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act)
Brando Benifei and Ioan-Dragos Tudorache. 2023 · 2023
Cited alongside, same era.
Revealing the structure of language model capabilities
Ryan Burnell, Han Hao, Andrew RA Conway, and Jose Hernandez Orallo. 2023a · 2023
Cited alongside, same era.
Rethink reporting of evaluation results in AI
Ryan Burnell, Wout Schellaert, John Burden, Tomer D Ullman, Fernando Martinez-Plumed, Joshua B Tenenbaum, Danaja Rutar, Lucy G Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, et al · 2023
Large Language Models for UAVs: Current State and Pathways to the Future
Shumaila Javaid, Nasir Saeed, and Bin He. 2024 · 2024
Closest in time.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2024 · 2024
Closest in time.
Alex Kipnis, Konstantinos Voudouris, Luca M Schulze Buschoff, and Eric Schulz. 2024 · 2024
Closest in time.
HELM Lite: Lightweight and Broad Capabilities Evaluation
Percy Liang, Yifan Mai, Josselin Somerville, Farzaan Kaiyom, Tony Lee, and Rishi Bommasani. [n. d.] · 2024
Closest in time.
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models–A Survey
Philipp Mondorf and Barbara Plank. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Calm-bench: A multi-task benchmark for evaluating causality-aware language models. In Findings of the Association for Computational Linguistics: EACL 2023 . 296–311
Dhairya Dalal, Paul Buitelaar, and Mihael Arcan. 2023 · 2023
Cited alongside, same era.
Towards Reasoning in Large Language Models: A Survey. In Findings of the Association for Computational Linguistics: ACL 2023 , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 1049–1065
Jie Huang and Kevin Chen-Chuan Chang. 2023 · 2023
Cited alongside, same era.
LogiQA 2.0—An Improved Dataset for Logical Reasoning in Natural Language Understanding
Hanmeng Liu, Jian Liu, Leyang Cui, Zhiyang Teng, Nan Duan, Ming Zhou, and Yue Zhang. 2023 · 2023
Cited alongside, same era.
Man Luo, Shrinidhi Kumbhar, Mihir Parmar, Neeraj Varshney, Pratyay Banerjee, Somak Aditya, Chitta Baral, et al · 2023
Cited alongside, same era.
GLoRE: Evaluating Logical Reasoning of Large Language Models
Zhiyang Teng, Ruoxi Ning, Jian Liu, Qiji Zhou, Yue Zhang, et al · 2023
Cited alongside, same era.
A Challenging Benchmark for Low-Resource Learning
Yudong Wang, Chang Ma, Qingxiu Dong, Lingpeng Kong, and Jingjing Xu. 2023 · 2023
Cited alongside, same era.
How Predictable Are Large Language Model Capabilities? A Case Study on BIG-bench. In The 2023 Conference on Empirical Methods in Natural Language Processing
Qinyuan Ye, Harvey Yiyun Fu, Xiang Ren, and Robin Jia. 2023 · 2023
Cited alongside, same era.
Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, and Yang You. 2024 · 2024
Closest in time.
Why the API output is inconsistent even after the temperature is set to 0
OpenAI. 2023 · 2024
Closest in time.
Deprecation Information
OpenAI. 2024a · 2024
Closest in time.
New embedding models and API updates
OpenAI. 2024b · 2024
Closest in time.
How predictable is language model benchmark performance?
David Owen. 2024 · 2024
Closest in time.
tinyBenchmarks: evaluating LLMs with fewer examples. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024 · 2024
Closest in time.
Observational Scaling Laws and the Predictability of Language Model Performance
Yangjun Ruan, Chris J Maddison, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 10406–10421
Charlotte Siska, Katerina Marazopoulou, Melissa Ailem, and James Bono. 2024 · 2024
Closest in time.
Anchor Points: Benchmarking Models with Much Fewer Examples. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian’s, Malta, 1576–1601
Rajan Vivek, Kawin Ethayarajh, Diyi Yang, and Douwe Kiela. 2024 · 2024
Closest in time.