Fetching the paper…
Reading the bibliography…
Evaluating LLMs and text-to-image models is a computationally intensive task often overlooked.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
A new readability yardstick
R. Flesch. 1948 · 1948
Earlier work this paper cites.
The technique of clear writing
R. Gunning. 1952 · 1952
Earlier work this paper cites.
The Automatic Creation of Literature Abstracts
H. P. Luhn. 1958 · 1958
Earlier work this paper cites.
Readability revisited: The new Dale-Chall readability formula
J. S. Chall and E. Dale. 1995 · 1995
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2009
Earlier work this paper cites.
Research on Text Clustering Algorithms. In 2010 2nd International Workshop on Database Technology and Applications . 1–3
Qun Li and Xinyuan Huang. 2010 · 2010
Earlier work this paper cites.
Average word length dynamics as indicator of cultural changes in society
Vladimir Bochkarev, Anna Shevlyakova, and Valery Solovyev. 2012 · 2012
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Evaluation of vector embedding models in clustering of text documents. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019) , Ruslan Mitkov and Galia Angelova (Eds.). INCOMA Ltd., Varna, Bulgaria, 1304–1311
Tomasz Walkowiak and Mateusz Gniewkowski. 2019 · 2019
Earlier work this paper cites.
Strategies for Difficulty Sampling Providing Diversity in Datasets
John Smith and Lisa Johnson. 2020 · 2020
Earlier work this paper cites.
Investigating minimum text lengths for lexical diversity indices
Fred Zenker and Kristopher Kyle. 2021 · 2020
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Cited alongside, same era.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Cited alongside, same era.
Improved K-Means Text Clustering Algorithm Based on BERT and Density Peak. In 2021 2nd Information Communication Technologies Conference (ICTC) . 260–264
Wenhao Hu, Dong Xu, and Zhihua Niu. 2021 · 2021
Cited alongside, same era.
AnglE-optimized Text Embeddings
Xianming Li and Jing Li. 2023 · 2023
Later among the works it cites.
Holistic Evaluation of Language Models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023 · 2023
Later among the works it cites.
When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
Max Marion, Ahmet Üstün, Luiza Pozzobon, Alex Wang, Marzieh Fadaee, and Sara Hooker. 2023 · 2023
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Mostly Basic Python Problems Dataset
Google Research. 2022 · 2022
Cited alongside, same era.
DeepCore: A Comprehensive Library for Coreset Selection in Deep Learning
Chengcheng Guo, Bo Zhao, and Yanbing Bai. 2022 · 2022
Cited alongside, same era.
Open LLM Leaderboard
Hugging Face. 2022b · 2022
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022 · 2022
Cited alongside, same era.
Framework for Topic Modeling using BERT, LDA and K-Means. In 2022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering (ICACITE) . 2204–2208
Kashi Sethia, Madhur Saxena, Mukul Goyal, and R.K. Yadav. 2022 · 2022
Cited alongside, same era.
The Efficiency Spectrum of Large Language Models: An Algorithmic Survey
Tianyu Ding, Tianyi Chen, Haidong Zhu, Jiachen Jiang, Yiqi Zhong, Jinxin Zhou, Guangzhi Wang, Zhihui Zhu, Ilya Zharkov, and Luming Liang. 2023 · 2023
Cited alongside, same era.
HEIM Leaderboard
Hugging Face. 2022a
Cited in the paper.
Later among the works it cites.
Data Selection for Language Models via Importance Resampling. In Thirty-seventh Conference on Neural Information Processing Systems
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. 2023 · 2023
Later among the works it cites.
Deep learning on a healthy data diet: finding important examples for fairness. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence (AAAI’23/IAAI’23/EAAI’23) . AAAI Press, Article 1637, 9 pages
Abdelrahman Zayed, Prasanna Parthasarathi, Gonçalo Mordido, Hamid Palangi, Samira Shabanian, and Sarath Chandar. 2023 · 2023
Later among the works it cites.
Misspelling Correction with Pre-trained Contextual Language Model
Yifei Hu, Xiaonan Jing, Youlim Ko, and Julia Taylor Rayz. 2024 · 2024
Closest in time.
Efficient Benchmarking of Language Models
Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, and Leshem Choshen. 2024 · 2024
Closest in time.
tinyBenchmarks: evaluating LLMs with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024 · 2024
Closest in time.
Lifelong Benchmarks: Efficient Model Evaluation in an Era of Rapid Progress
Ameya Prabhu, Vishaal Udandarao, Philip Torr, Matthias Bethge, Adel Bibi, and Samuel Albanie. 2024 · 2024
Closest in time.
Anchor Points: Benchmarking Models with Much Fewer Examples. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian’s, Malta, 1576–1601
Rajan Vivek, Kawin Ethayarajh, Diyi Yang, and Douwe Kiela. 2024 · 2024
Closest in time.