Fetching the paper…
Reading the bibliography…
Quantitative evaluation metrics have traditionally been pivotal in gauging the advancements of artificial intelligence systems, including large language models (LLMs).
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Correlation between rouge and human evaluation of extractive meeting summaries
Feifan Liu and Yang Liu. 2008 · 2008
Earlier work this paper cites.
An investigation into the validity of some metrics for automatically evaluating natural language generation systems
Ehud Reiter and Anja Belz. 2009 · 2009
Earlier work this paper cites.
Chia-Wei Liu, Ryan Lowe, Iulian V Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondrej Dusek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
Manifold: A model-agnostic framework for interpretation and diagnosis of machine learning models
Jiawei Zhang, Yang Wang, Piero Molino, Lezhi Li, and David S Ebert. 2018 · 2018
Earlier work this paper cites.
GEval: Tool for debugging NLP datasets and models
Filip Graliński, Anna Wróblewska, Tomasz Stanisławek, Kamil Grabowski, and Tomasz Górecki. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of nlp models with checklist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Cited alongside, same era.
Dialogsum: A real-life scenario dialogue summarization dataset
Yulong Chen, Yang Liu, Liang Chen, and Yue Zhang. 2021 · 2021
Cited alongside, same era.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022 · 2022
Later among the works it cites.
Agro: Adversarial discovery of error-prone groups for robust optimization
Bhargavi Paranjape, Pradeep Dasigi, Vivek Srikumar, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Closest in time.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Closest in time.
Introducing chatgpt
OpenAI. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automl: A survey of the state-of-the-art
Xin He, Kaiyong Zhao, and Xiaowen Chu. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021 · 2021
Cited alongside, same era.
Explanation-based human debugging of NLP models: A survey
Piyawat Lertvittayakumjorn and Francesca Toni. 2021 · 2021
Cited alongside, same era.
Meaningfully debugging model mistakes using conceptual counterfactual explanations
Abubakar Abid, Mert Yuksekgonul, and James Zou. 2022 · 2022
Cited alongside, same era.
Automl in the age of large language models: Current challenges, future opportunities and risks
Alexander Tornede, Difan Deng, Theresa Eimer, Joseph Giovanelli, Aditya Mohan, Tim Ruhkopf, Sarah Segel, Daphne Theodorakopoulos, Tanja Tornede, Henning Wachsmuth, et al. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Codebertscore: Evaluating code generation with pretrained models of code
Shuyan Zhou, Uri Alon, Sumit Agarwal, and Graham Neubig. 2023 · 2023
Closest in time.