Fetching the paper…
Reading the bibliography…
Model selection for a given target task can be costly, as it may entail extensive annotation of the quality of outputs of different models.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 1909
Earlier work this paper cites.
Statistical theories of mental test scores
FM Lord, MR Novick, and Allan Birnbaum. 1968 · 1968
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
(meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Modern hierarchical, agglomerative clustering algorithms
Daniel Müllner. 2011 · 2011
Earlier work this paper cites.
Active evaluation of classifiers on large datasets
Namit Katariya, Arun Iyer, and Sunita Sarawagi. 2012 · 2012
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gulçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
QuAC: Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese. 2018 · 2018
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
ChatEval: A tool for chatbot evaluation
João Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani, Lyle Ungar, and Chris Callison-Burch. 2019 · 2019
Cited alongside, same era.
Best practices for the human evaluation of automatically generated text
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel J. Krahmer. 2019 · 2019
Cited alongside, same era.
Corpus wide argument mining—a working solution
Liat Ein-Dor, Eyal Shnarch, Lena Dankin, Alon Halfon, Benjamin Sznajder, Ariel Gera, Carlos Alzate, Martin Gleize, Leshem Choshen, Yufang Hou, et al. 2020 · 2020
Cited alongside, same era.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022 · 2022
Later among the works it cites.
Active evaluation: Efficient NLG evaluation with few pairwise comparisons
Akash Kumar Mohankumar and Mitesh Khapra. 2022 · 2022
Later among the works it cites.
Active learning for abstractive text summarization
Akim Tsvigun, Ivan Lysenko, Danila Sedashov, Ivan Lazichny, Eldar Damirov, Vladimir Karlov, Artemy Belousov, Leonid Sanochkin, Maxim Panov, Alexander Panchenko, Mikhail Burtsev, and Artem Shelmanov. 2022 · 2022
Later among the works it cites.
A survey of active learning for natural language processing
Zhisong Zhang, Emma Strubell, and Eduard Hovy. 2022 · 2022
Later among the works it cites.
Emergent and predictable memorization in large language models
Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raf. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond user self-reported Likert scale ratings: A comparison model for automatic dialog evaluation
Weixin Liang, James Zou, and Zhou Yu. 2020 · 2020
Cited alongside, same era.
ALT-MAS: A data-efficient framework for active testing of machine learning algorithms
Huong Ha, Sunil Gupta, Santu Rana, and Svetha Venkatesh. 2021 · 2021
Cited alongside, same era.
Active bayesian assessment of black-box classifiers
Disi Ji, Robert L. Logan, Padhraic Smyth, and Mark Steyvers. 2021 · 2021
Cited alongside, same era.
Active testing: Sample-efficient model evaluation
Jannik Kossen, Sebastian Farquhar, Yarin Gal, and Tom Rainforth. 2021 · 2021
Cited alongside, same era.
Evaluation examples are not equally informative: How should that change NLP leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Cited alongside, same era.
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021 · 2021
Cited alongside, same era.
Comparing test sets with item response theory
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, and Samuel R. Bowman. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Benchmarking large language model capabilities for conditional generation
Joshua Maynez, Priyanka Agrawal, and Sebastian Gehrmann. 2023 · 2023
Later among the works it cites.
State of what art? a call for multi-prompt LLM evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2023 · 2023
Later among the works it cites.
Efficient benchmarking (of language models)
Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, and Leshem Choshen. 2023 · 2023
Later among the works it cites.
Anchor points: Benchmarking models with much fewer examples
Rajan Vivek, Kawin Ethayarajh, Diyi Yang, and Douwe Kiela. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
Efficient multi-prompt evaluation of LLMs
Felipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. 2024 · 2024
Closest in time.