Fetching the paper…
Reading the bibliography…
Performance evaluation plays a crucial role in the development life cycle of large language models (LLMs).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Who belongs in the family?
Robert L Thorndike. 1953 · 1953
Earlier work this paper cites.
Kolmogorov-Smirnov two-sample tests
John W Pratt, Jean D Gibbons, John W Pratt, and Jean D Gibbons. 1981 · 1981
Earlier work this paper cites.
The dip test of unimodality
John A Hartigan and Pamela M Hartigan. 1985 · 1985
Earlier work this paper cites.
Auction algorithms for network flow problems: A tutorial introduction
Dimitri P Bertsekas. 1992 · 1992
Earlier work this paper cites.
Generalization of the Kolmogorov-Smirnov test
Erhard Reschenhofer. 1997 · 1997
Earlier work this paper cites.
Approximate nearest neighbors: towards removing the curse of dimensionality. In Proceedings of the thirtieth annual ACM symposium on Theory of computing . 604–613
Piotr Indyk and Rajeev Motwani. 1998 · 1998
Earlier work this paper cites.
Stratified Sampling Types
Garrett Glasgow. 2005 · 2005
Earlier work this paper cites.
Compressed sensing
David L Donoho. 2006 · 2006
Earlier work this paper cites.
Finding a" kneedle" in a haystack: Detecting knee points in system behavior. In 2011 31st international conference on distributed computing systems workshops . IEEE, 166–171
Ville Satopaa, Jeannie Albrecht, David Irwin, and Barath Raghavan. 2011 · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams. 2012 · 2012
Earlier work this paper cites.
Better mixing via deep representations. In International conference on machine learning . PMLR, 552–560
Yoshua Bengio, Grégoire Mesnil, Yann Dauphin, and Salah Rifai. 2013 · 2013
Earlier work this paper cites.
Balanced k-means for clustering. In Structural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshop, S+ SSPR 2014, Joensuu, Finland, August 20-22, 2014. Proceedings . Springer, 32–41
Mikko I Malinen and Pasi Fränti. 2014 · 2014
Earlier work this paper cites.
Adaptive strategy for stratified Monte Carlo sampling
Alexandra Carpentier, Remi Munos, and András Antos. 2015 · 2015
Earlier work this paper cites.
On the rate of convergence in Wasserstein distance of the empirical measure
Nicolas Fournier and Arnaud Guillin. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Carreras (Eds.). Association for Computational Linguistics, Austin, Texas, 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Martin Arjovsky, Soumith Chintala, and Léon Bottou. 2017 · 2017
Earlier work this paper cites.
Detecting adversarial samples from artifacts
Reuben Feinman, Ryan R Curtin, Saurabh Shintre, and Andrew B Gardner. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan (Eds.). Association for Computational Linguistics, Vancouver, Canada, 1601–1611
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Why is random testing effective for partition tolerance bugs?
Rupak Majumdar and Filip Niksic. 2017 · 2017
Earlier work this paper cites.
Deepxplore: Automated whitebox testing of deep learning systems. In proceedings of the 26th Symposium on Operating Systems Principles . 1–18
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Randomized testing of distributed systems with probabilistic guarantees
Burcu Kulahcioglu Ozkan, Rupak Majumdar, Filip Niksic, Mitra Tabaei Befrouei, and Georg Weissenbacher. 2018 · 2018
Earlier work this paper cites.
Active Learning for Convolutional Neural Networks: A Core-Set Approach. In International Conference on Learning Representations
Ozan Sener and Silvio Savarese. 2018 · 2018
Earlier work this paper cites.
Guiding deep learning system testing using surprise adequacy. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) . IEEE, 1039–1049
Jinhan Kim, Robert Feldt, and Shin Yoo. 2019 · 2019
Earlier work this paper cites.
Hierarchical optimal transport for multimodal distribution alignment
John Lee, Max Dabagia, Eva Dyer, and Christopher Rozell. 2019b · 2019
Earlier work this paper cites.
Boosting operational dnn testing efficiency through conditioning. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 499–509
Zenan Li, Xiaoxing Ma, Chang Xu, Chun Cao, Jingwei Xu, and Jian Lü. 2019 · 2019
Earlier work this paper cites.
Adaptive: Parallel active learning of mathematical functions
Bas Nijholt, Joseph Weston, Jorn Hoofwijk, and Anton Akhmerov. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Practical accuracy estimation for efficient deep neural network testing
Junjie Chen, Zhuo Wu, Zan Wang, Hanmo You, Lingming Zhang, and Ming Yan. 2020 · 2020
Earlier work this paper cites.
Deepgini: prioritizing massive tests to enhance the robustness of deep neural networks. In Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis . 177–188
Yang Feng, Qingkai Shi, Xinyu Gao, Jun Wan, Chunrong Fang, and Zhenyu Chen. 2020 · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Cited alongside, same era.
Are labels always necessary for classifier accuracy evaluation?. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 15069–15078
Weijian Deng and Liang Zheng. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Cited alongside, same era.
Predicting with confidence on unseen distributions. In Proceedings of the IEEE/CVF international conference on computer vision . 1134–1144
Devin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell, and Ludwig Schmidt. 2021 · 2021
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Later among the works it cites.
CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 5067–5089
Weixiang Yan, Yuchen Tian, Yunzhe Li, Qian Chen, and Wen Wang. 2023 · 2023
Later among the works it cites.
Foundations of human spatial problem solving
Noah Zarr and Joshua W Brown. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q 2 Q^{2} : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 7856–7870
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Cited alongside, same era.
Active testing: Sample-efficient model evaluation. In International Conference on Machine Learning . PMLR, 5753–5763
Jannik Kossen, Sebastian Farquhar, Yarin Gal, and Tom Rainforth. 2021 · 2021
Cited alongside, same era.
Sashank Santhanam, Behnam Hedayatnia, Spandana Gella, Aishwarya Padmakumar, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. 2021 · 2021
Cited alongside, same era.
Prioritizing test inputs for deep neural networks via mutation analysis. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 397–409
Zan Wang, Hanmo You, Junjie Chen, Yingyi Zhang, Xuyuan Dong, and Wenbin Zhang. 2021 · 2021
Cited alongside, same era.
Active surrogate estimators: An active learning approach to label-efficient model evaluation
Jannik Kossen, Sebastian Farquhar, Yarin Gal, and Thomas Rainforth. 2022 · 2022
Cited alongside, same era.
A new generation of perspective api: Efficient multilingual character-level transformers. In Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining . 3197–3207
Alyssa Lees, Vinh Q Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022 · 2022
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 3214–3252
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
An Exploratory Study of AI System Risk Assessment from the Lens of Data Distribution and Uncertainty
Zhijie Wang, Yuheng Huang, Lei Ma, Haruki Yokoyama, Susumu Tokumoto, and Kazuki Munakata. 2022 · 2022
Cited alongside, same era.
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al · 2024
Closest in time.
Label-Efficient Model Selection for Text Generation
Shir Ashury-Tahan, Benjamin Sznajder, Leshem Choshen, Liat Ein-Dor, Eyal Shnarch, and Ariel Gera. 2024 · 2024
Closest in time.
Humans or LLMs as the Judge? A Study on Judgement Bias. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 8301–8327
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024b · 2024
Closest in time.
NVIDIA cuVS
NVIDIA Corporation. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Patchscope: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024 · 2024
Closest in time.
DeepSample: DNN sampling-based testing for operational accuracy assessment. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–12
Antonio Guerriero, Roberto Pietrantuono, and Stefano Russo. 2024 · 2024
Closest in time.
Language Models Represent Space and Time. In The Twelfth International Conference on Learning Representations
Wes Gurnee and Max Tegmark. 2024 · 2024
Closest in time.
Large language models for software engineering: A systematic literature review
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024 · 2024
Closest in time.
Exploring Concept Depth: How Large Language Models Acquire Knowledge at Different Layers?
Mingyu Jin, Qinkai Yu, Jingyuan Huang, Qingcheng Zeng, Zhenting Wang, Wenyue Hua, Haiyan Zhao, Kai Mei, Yanda Meng, Kaize Ding, et al · 2024
Closest in time.
Benchmarking Cognitive Biases in Large Language Models as Evaluators. In Findings of the Association for Computational Linguistics: ACL 2024 , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 517–545
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2024 · 2024
Closest in time.
Evaluating Language Models for Efficient Code Generation. In First Conference on Language Modeling
Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. 2024 · 2024
Closest in time.
Characterizing out-of-distribution error via optimal transport
Yuzhe Lu, Yilong Qin, Runtian Zhai, Andrew Shen, Ketong Chen, Zhenlin Wang, Soheil Kolouri, Simon Stepputtis, Joseph Campbell, and Katia Sycara. 2024 · 2024
Closest in time.
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. In Proceedings of the Conference on Language Modeling (COLM) . 123–134
Samuel Marks and Max Tegmark. 2024 · 2024
Closest in time.
Predicting the Performance of Foundation Models via Agreement-on-the-Line
Aman Mehra, Rahul Saxena, Taeyoun Kim, Christina Baek, Zico Kolter, and Aditi Raghunathan. 2024 · 2024
Closest in time.
Divide-and-Aggregate Learning for Evaluating Performance on Unlabeled Data. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21395–21402
Shuyu Miao, Jian Liu, Lin Zheng, and Hong Jin. 2024 · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R Bowman, and Shi Feng. 2024 · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
Language-specific neurons: The key to multilingual capabilities in large language models
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024 · 2024
Closest in time.
Where Do Large Language Models Fail When Generating Code?
Zhijie Wang, Zijie Zhou, Da Song, Yuheng Huang, Shengmai Chen, Lei Ma, and Tianyi Zhang. 2024b · 2024
Closest in time.
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. In The Twelfth International Conference on Learning Representations
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Closest in time.
Large language models as markov chains
Oussama Zekri, Ambroise Odonnat, Abdelhakim Benechehab, Linus Bleistein, Nicolas Boullé, and Ievgen Redko. 2024 · 2024
Closest in time.
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
Lin Zhao, Tianchen Zhao, Zinan Lin, Xuefei Ning, Guohao Dai, Huazhong Yang, and Yu Wang. 2024 · 2024
Closest in time.
Towards fully exploiting llm internal states to enhance knowledge boundary perception
Shiyu Ni, Keping Bi, Jiafeng Guo, Lulu Yu, Baolong Bi, and Xueqi Cheng. 2025 · 2025
Closest in time.
JudgeBench: A Benchmark for Evaluating LLM-Based Judges. In The Thirteenth International Conference on Learning Representations
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang, Alejandro Cuadron, Chenguang Wang, Raluca Popa, and Ion Stoica. 2025 · 2025
Closest in time.