Fetching the paper…
Reading the bibliography…
Evaluating natural language generation (NLG) is a vital but challenging problem in natural language processing.
Chatgpt to replace crowdsourcing of paraphrases for intent classification: Higher diversity and comparable model robustness
Cegin, Ján, Jakub Simko, and Peter Brusilovsky. 2023 · 1905
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, Kishore, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Interpreting BLEU/NIST scores: How much improvement do we need to have a better system?
Zhang, Ying, Stephan Vogel, and Alex Waibel. 2004 · 2004
Earlier work this paper cites.
Evaluating evaluation methods for generation in the presence of variation
Stent, Amanda, Matthew Marge, and Mohit Singhai. 2005 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Celikyilmaz, Asli, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Challenges in data-to-document generation
Wiseman, Sam, Stuart M. Shieber, and Alexander M. Rush. 2017 · 2017
Earlier work this paper cites.
Language (technology) is power: A critical survey of "bias" in NLP
Blodgett, Su Lin, Solon Barocas, Hal Daumé III, and Hanna M. Wallach. 2020 · 2020
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Durmus, Esin, He He, and Mona T. Diab. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, Mark, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Earlier work this paper cites.
Survey on evaluation methods for dialogue systems
Deriu, Jan, Álvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Fabbri, Alexander R., Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2021 · 2021
Earlier work this paper cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Gehrmann, Sebastian, Tosin P. Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondrej Dusek, Chris Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Rubungo Andre Niyongabo, Salomey Osei, Ankur P. Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Earlier work this paper cites.
Human evaluation of creative NLG systems: An interdisciplinary survey on recent papers
Hämäläinen, Mika and Khalid Al-Najjar. 2021 · 2021
Earlier work this paper cites.
Human evaluation of automatically generated text: Current trends and best practice guidelines
van der Lee, Chris, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, Reiichiro, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Perturbation checklists for evaluating NLG evaluation metrics
Sai, Ananya B., Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M. Khapra. 2021 · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation
Yuan, Weizhe, Graham Neubig, and Pengfei Liu. 2021 · 2021
Earlier work this paper cites.
A human-machine collaborative framework for evaluating malevolence in dialogues
Zhang, Yangjun, Pengjie Ren, and Maarten de Rijke. 2021 · 2021
Earlier work this paper cites.
Frugalscore: Learning cheaper, lighter and faster evaluation metrics for automatic text generation
Eddine, Moussa Kamal, Guokan Shang, Antoine J.-P. Tixier, and Michalis Vazirgiannis. 2022 · 2022
Earlier work this paper cites.
Results of WMT22 metrics shared task: Stop using BLEU - neural metrics are better and more robust
Freitag, Markus, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George F. Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Earlier work this paper cites.
Capturing failures of large language models via human cognitive biases
Jones, Erik and Jacob Steinhardt. 2022 · 2022
Earlier work this paper cites.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jung, Jaehun, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Yujia, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022 · 2022
Earlier work this paper cites.
Teaching language models to support answers with verified quotes
Menick, Jacob, Maja Trebacz, Vladimir Mikulik, John Aslanides, H. Francis Song, Martin J. Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and Nat McAleese. 2022 · 2022
Earlier work this paper cites.
Adaptive testing and debugging of NLP models
Ribeiro, Marco Túlio and Scott M. Lundberg. 2022 · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
Saunders, William, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022 · 2022
Earlier work this paper cites.
Bertscore is unfair: On social bias in language model-based metrics for text generation
Sun, Tianxiang, Junliang He, Xipeng Qiu, and Xuanjing Huang. 2022 · 2022
Earlier work this paper cites.
Language models with image descriptors are strong few-shot video-language learners
Wang, Zhenhailong, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei, Xudong Lin, Shuohang Wang, Ziyi Yang, Chenguang Zhu, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, and Heng Ji. 2022 · 2022
Earlier work this paper cites.
Interpreting language models with contrastive explanations
Yin, Kayo and Graham Neubig. 2022 · 2022
Earlier work this paper cites.
OPT: open pre-trained transformer language models
Zhang, Susan, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications
Zhou, Kaitlyn, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, and Alexandra Olteanu. 2022 · 2022
Earlier work this paper cites.
Palm 2 technical report
Anil, Rohan, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Ábrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, and et al. 2023 · 2023
Earlier work this paper cites.
Benchmarking foundation models with language-model-as-an-examiner
Bai, Yushi, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. 2023 · 2023
Earlier work this paper cites.
Can large language models be an alternative to human evaluations?
Chiang, David Cheng-Han and Hung-yi Lee. 2023a · 2023
Earlier work this paper cites.
A closer look into using large language models for automatic evaluation
Chiang, David Cheng-Han and Hung-yi Lee. 2023b · 2023
Earlier work this paper cites.
LM vs LM: detecting factual errors via cross examination
Cohen, Roi, May Hamri, Mor Geva, and Amir Globerson. 2023 · 2023
Earlier work this paper cites.
Are large language models reliable judges? A study on the factuality evaluation capabilities of llms
Fu, Xue-Yong, Md. Tahmid Rahman Laskar, Cheng Chen, and Shashi Bhushan TN. 2023 · 2023
Earlier work this paper cites.
Human-like summarization evaluation with chatgpt
Gao, Mingqi, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Earlier work this paper cites.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Gehrmann, Sebastian, Elizabeth Clark, and Thibault Sellam. 2023 · 2023
Earlier work this paper cites.
Chatgpt outperforms crowd-workers for text-annotation tasks
Gilardi, Fabrizio, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Earlier work this paper cites.
Coascore: Chain-of-aspects prompting for NLG evaluation
Gong, Peiyuan and Jiaxin Mao. 2023 · 2023
Earlier work this paper cites.
ALLURE: auditing and improving llm-based evaluation of text using iterative in-context-learning
Hasanbeig, Hosein, Hiteshi Sharma, Leo Betthauser, Felipe Vieira Frujeri, and Ida Momennejad. 2023 · 2023
Cited alongside, same era.
On the blind spots of model-based evaluation metrics for text generation
He, Tianxing, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James R. Glass, and Yulia Tsvetkov. 2023a · 2023
Cited alongside, same era.
On the blind spots of model-based evaluation metrics for text generation
He, Tianxing, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James R. Glass, and Yulia Tsvetkov. 2023b · 2023
Cited alongside, same era.
Decipherpref: Analyzing influential factors in human preference judgments via GPT-4
Hu, Yebowen, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, and Fei Liu. 2023 · 2023
Cited alongside, same era.
Multi-dimensional evaluation of text summarization with in-context learning
Jain, Sameer, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023 · 2023
Through the lens of core competency: Survey on evaluation of large language models
Zhuang, Ziyu, Qiguang Chen, Longxuan Ma, Mingda Li, Yi Han, Yushan Qian, Haopeng Bai, Zixian Feng, Weinan Zhang, and Ting Liu. 2023 · 2023
Later among the works it cites.
Ashktorab, Zahra, Michael Desmond, Qian Pan, James M. Johnson, Martin Santillan Cooper, Elizabeth M. Daly, Rahul Nair, Tejaswini Pedapati, Swapnaja Achintalwar, and Werner Geyer. 2024 · 2024
Closest in time.
Llms instead of human judges? A large scale empirical study across 20 NLP evaluation tasks
Bavaresco, Anna, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K. Surikuchi, Ece Takmaz, and Alberto Testoni. 2024 · 2024
Closest in time.
Compassjudger-1: All-in-one judge model helps model evaluation and evolution
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ji, Yunjie, Yan Gong, Yiping Peng, Chao Ni, Peiyan Sun, Dongyu Pan, Baochang Ma, and Xiangang Li. 2023 · 2023
Cited alongside, same era.
Zero-shot faithfulness evaluation for text summarization with foundation language model
Jia, Qi, Siyu Ren, Yizhu Liu, and Kenny Q. Zhu. 2023 · 2023
Cited alongside, same era.
Ke, Pei, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. 2023 · 2023
Cited alongside, same era.
Which is better? exploring prompting strategy for llm-based metrics
Kim, Joonghoon, Sangmin Lee, Seung Hun Han, Saeran Park, Jiyoon Lee, Kiyoon Jeong, and Pilsung Kang. 2023 · 2023
Cited alongside, same era.
GEMBA-MQM: detecting translation quality error spans with GPT-4
Kocmi, Tom and Christian Federmann. 2023a · 2023
Cited alongside, same era.
Large language models are state-of-the-art evaluators of translation quality
Kocmi, Tom and Christian Federmann. 2023b · 2023
Cited alongside, same era.
Little giants: Exploring the potential of small llms as evaluation metrics in summarization in the eval4nlp 2023 shared task
Kotonya, Neema, Saran Krishnasamy, Joel R. Tetreault, and Alejandro Jaimes. 2023 · 2023
Cited alongside, same era.
Cao, Maosong, Alexander Lam, Haodong Duan, Hongwei Liu, Songyang Zhang, and Kai Chen. 2024 · 2024
Closest in time.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, Chi-Min, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2024 · 2024
Closest in time.
Booookscore: A systematic exploration of book-length summarization in the era of llms
Chang, Yapei, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024 · 2024
Closest in time.
Scaling instruction-finetuned language models
Chung, Hyung Won, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2024 · 2024
Closest in time.
Ragas: Automated evaluation of retrieval augmented generation
ES, Shahul, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024 · 2024
Closest in time.
Gptscore: Evaluate as you desire
Fu, Jinlan, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2024 · 2024
Closest in time.
Language models hallucinate, but may excel at fact verification
Guan, Jian, Jesse Dodge, David Wadden, Minlie Huang, and Hao Peng. 2024 · 2024
Closest in time.
METAL: towards multilingual meta-evaluation
Hada, Rishav, Varun Gumma, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2024a · 2024
Closest in time.
Are large language model-based evaluators the solution to scaling up multilingual evaluation?
Hada, Rishav, Varun Gumma, Adrian de Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram. 2024b · 2024
Closest in time.
Socreval: Large language models with the socratic method for reference-free reasoning evaluation
He, Hangfeng, Hongming Zhang, and Dan Roth. 2024 · 2024
Closest in time.
Themis: A reference-free NLG evaluation language model with flexibility and interpretability
Hu, Xinyu, Li Lin, Mingqi Gao, Xunjian Yin, and Xiaojun Wan. 2024 · 2024
Closest in time.
Tigerscore: Towards building explainable metric for all text generation tasks
Jiang, Dongfu, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2024 · 2024
Closest in time.
Prometheus: Inducing fine-grained evaluation capability in language models
Kim, Seungone, Jamin Shin, Yejin Choi, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2024a · 2024
Closest in time.
Evallm: Interactive evaluation of large language model prompts on user-defined criteria
Kim, Tae Soo, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024d · 2024
Closest in time.
Instructpatentgpt: Training patent language models to follow instructions with human feedback
Lee, Jieh-Sheng. 2024 · 2024
Closest in time.
Prexme! large scale prompt exploration of open source llms for machine translation and summarization evaluation
Leiter, Christoph and Steffen Eger. 2024 · 2024
Closest in time.
Generative judge for evaluating alignment
Li, Junlong, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2024 · 2024
Closest in time.
Benchmarking generation and evaluation capabilities of large language models for instruction controllable summarization
Liu, Yixin, Alexander R. Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan. 2024a · 2024
Closest in time.
On learning to summarize with large language models as references
Liu, Yixin, Kejian Shi, Katherine He, Longtian Ye, Alexander R. Fabbri, Pengfei Liu, Dragomir Radev, and Arman Cohan. 2024b · 2024
Closest in time.
Calibrating llm-based evaluator
Liu, Yuxuan, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024c · 2024
Closest in time.
LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Liusie, Adian, Potsawee Manakul, and Mark J. F. Gales. 2024 · 2024
Closest in time.
Large language models meet user interfaces: The case of provisioning feedback
Pozdniakov, Stanislav, Jonathan Brazil, Solmaz Abdi, Aneesha Bakharia, Shazia Sadiq, Dragan Gasevic, Paul Denny, and Hassan Khosravi. 2024 · 2024
Closest in time.
Concrete problems in AI safety, revisited
Raji, Inioluwa Deborah and Roel Dobbe. 2024 · 2024
Closest in time.
Branch-solve-merge improves large language model evaluation and generation
Saha, Swarnadeep, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, and Xian Li. 2024 · 2024
Closest in time.
Who validates the validators? aligning llm-assisted evaluation of LLM outputs with human preferences
Shankar, Shreya, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran, and Ian Arawjo. 2024 · 2024
Closest in time.
Repeval: Effective text evaluation with LLM representation
Sheng, Shuqian, Yi Xu, Tianhang Zhang, Zanwei Shen, Luoyi Fu, Jiaxin Ding, Lei Zhou, Xiaoying Gan, Xinbing Wang, and Chenghu Zhou. 2024 · 2024
Closest in time.
Fusion-eval: Integrating assistant evaluators with llms
Shu, Lei, Nevan Wichers, Liangchen Luo, Yun Zhu, Yinxiao Liu, Jindong Chen, and Lei Meng. 2024 · 2024
Closest in time.
Large language models are not fair evaluators
Wang, Peiyi, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024a · 2024
Closest in time.
Pandalm: An automatic evaluation benchmark for LLM instruction tuning optimization
Wang, Yidong, Zhuohao Yu, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang. 2024c · 2024
Closest in time.
FLASK: fine-grained language model evaluation based on alignment skill sets
Ye, Seonghyeon, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo. 2024 · 2024
Closest in time.
Batcheval: Towards human-like text evaluation
Yuan, Peiwen, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu, and Kan Li. 2024 · 2024
Closest in time.
A comprehensive analysis of the effectiveness of large language models as automatic dialogue evaluators
Zhang, Chen, Luis Fernando D’Haro, Yiming Chen, Malu Zhang, and Haizhou Li. 2024 · 2024
Closest in time.
Ai-assisted human evaluation of machine translation
Zouhar, Vilém, Tom Kocmi, and Mrinmaya Sachan. 2024 · 2024
Closest in time.
Think together and work better: Combining humans’ and llms’ think-aloud outcomes for effective text evaluation
Chu, SeongYeub, Jong Woo Kim, and Mun Yong Yi. 2025 · 2025
Closest in time.
Ma, Bolei, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, and Barbara Plank. 2025 · 2025
Closest in time.
Evaluating the evaluator: Measuring llms’ adherence to task evaluation instructions
Murugadoss, Bhuvanashree, Christian Pölitz, Ian Drosos, Vu Le, Nick McKenna, Carina Suzana Negreanu, Chris Parnin, and Advait Sarkar. 2025 · 2025
Closest in time.
Style over substance: Evaluation biases for large language models
Wu, Minghao and Alham Fikri Aji. 2025 · 2025
Closest in time.