Fetching the paper…
Reading the bibliography…
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable.
“The winograd schema challenge”
Hector Levesque, Ernest Davis and Leora Morgenstern · 2012
Earlier work this paper cites.
“Intriguing properties of neural networks”
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow and Rob Fergus · 2014
Earlier work this paper cites.
“Vqa: Visual question answering”
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Zitnick and Devi Parikh · 2015
Earlier work this paper cites.
“Parsing algebraic word problems into equations”
Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni and Siena Ang · 2015
Earlier work this paper cites.
“Reasoning about quantities in natural language”
Subhro Roy, Tim Vieira and Dan Roth · 2015
Earlier work this paper cites.
“The Limitations of Deep Learning in Adversarial Settings”
Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z. Celik and Ananthram Swami · 2016
Earlier work this paper cites.
“Solving general arithmetic word problems”
Subhro Roy and Dan Roth · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Towards evaluating the robustness of neural networks”
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
“Making the v in vqa matter: Elevating the role of image understanding in visual question answering”
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra and Devi Parikh · 2017
Earlier work this paper cites.
“Training verified learners with learned verifiers”
Krishnamurthy Dvijotham, Sven Gowal, Robert Stanforth, Relja Arandjelovic, Brendan O’Donoghue, Jonathan Uesato and Pushmeet Kohli · 2018
Earlier work this paper cites.
“Towards deep learning models resistant to adversarial attacks”
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras and Adrian Vladu · 2018
Earlier work this paper cites.
“Know what you don’t know: Unanswerable questions for SQuAD”
Pranav Rajpurkar, Robin Jia and Percy Liang · 2018
Earlier work this paper cites.
“GLUE: A multi-task benchmark and analysis platform for natural language understanding”
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy and Samuel Bowman · 2018
Earlier work this paper cites.
“HotpotQA: A dataset for diverse, explainable multi-hop question answering”
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov and Christopher Manning · 2018
Earlier work this paper cites.
“Tabfact: A large-scale dataset for table-based fact verification”
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou and William Wang · 2019
Earlier work this paper cites.
“DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs”
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh and Matt Gardner · 2019
Earlier work this paper cites.
“Measuring massive multitask language understanding”
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song and Jacob Steinhardt · 2020
Earlier work this paper cites.
“From ImageNet to Image Classification: Contextualizing Progress on Benchmarks”
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas and Aleksander Madry · 2020
Cited alongside, same era.
“What will it take to fix benchmarking in natural language understanding?”
Samuel Bowman and George Dahl · 2021
Cited alongside, same era.
“Training verifiers to solve math word problems”
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton and Reiichiro Nakano · 2021
Cited alongside, same era.
“Measuring Mathematical Problem Solving With the MATH Dataset”
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song and Jacob Steinhardt · 2021
Cited alongside, same era.
“Pervasive label errors in test sets destabilize machine learning benchmarks”
“Automatic model selection with large language models for reasoning”
James Zhao, Yuxi Xie, Kenji Kawaguchi, Junxian He and Michael Xie · 2023
Later among the works it cites.
“Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku”, 2024
Anthropic · 2024
Later among the works it cites.
“Moffatt v. Air Canada” 2024 BCCRT 149, Small Claims Decisions, File No. SC-2023-005609, Final Decision, 2024
Civil Resolution Tribunal · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang and Angela Fan · 2024
Later among the works it cites.
“DeepSeek-V3 Technical Report”, 2024
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai and Daya Guo · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Curtis Northcutt, Anish Athalye and Jonas Mueller · 2021
Cited alongside, same era.
“Are NLP models really able to solve simple math word problems?”
Arkil Patel, Satwik Bhattamishra and Navin Goyal · 2021
Cited alongside, same era.
“Beyond the imitation game: Quantifying and extrapolating the capabilities of language models”
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Shoeb, Abubakar Abid, Adam Fisch, Adam Brown, Adam Santoro, Aditya Gupta and Adrià Garriga-Alonso · 2022
Cited alongside, same era.
“Challenging big-bench tasks and whether chain-of-thought can solve them”
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi and Denny Zhou · 2022
Cited alongside, same era.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc Le and Denny Zhou · 2022
Cited alongside, same era.
“Claude 3.5 Sonnet Model Card Addendum”, 2023
Anthropic · 2023
Cited alongside, same era.
“Retrieval-augmented generation for large language models: A survey”
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun and Haofen Wang · 2023
Cited alongside, same era.
“Reasoning with language model is planning with world model”
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang and Zhiting Hu · 2023
Cited alongside, same era.
Later among the works it cites.
Aryo Gema, Joshua Leang, Giwon Hong, Alessio Devoto, Alberto Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du and Mohammad Madani · 2024
Later among the works it cites.
“DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence”
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu and YK Li · 2024
Later among the works it cites.
“Has anyone analyzed what’s the ground truth label error rate for GSM8k? It’s possible we entered data leakage and overfitting territory a while ago.”, 2024
Peter Henderson · 2024
Later among the works it cites.
“SWE-bench: Can Language Models Resolve Real-world Github Issues?”
Carlos Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press and Karthik Narasimhan · 2024
Later among the works it cites.
Marianna Nezhurina, Lucia Cipolina-Kun, Mehdi Cherti and Jenia Jitsev · 2024
Later among the works it cites.
“OpenAI o1 System Card”, 2024
OpenAI · 2024
Later among the works it cites.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat and Julian Schrittwieser · 2024
Later among the works it cites.
“Why AI Can’t Spell ’Strawberry”’
Amanda Silberling · 2024
Later among the works it cites.
“Qwen2.5: A Party of Foundation Models”, 2024
Qwen Team · 2024
Later among the works it cites.
“Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?”
Zhe Yang, Yichang Zhang, Tianyu Liu, Jian Yang, Junyang Lin, Chang Zhou and Zhifang Sui · 2024
Later among the works it cites.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang and Xiao Bi · 2025
Closest in time.
“OpenAI o3-mini System Card”, 2025
OpenAI · 2025
Closest in time.