Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have been reported to outperform existing automatic evaluation metrics in some tasks, such as text summarization and machine translation.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Better evaluation for grammatical error correction
Daniel Dahlmeier and Hwee Tou Ng. 2012 · 2012
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013 · 2013
Earlier work this paper cites.
Efficient elicitation of annotations for human evaluation of machine translation
Keisuke Sakaguchi, Matt Post, and Benjamin Van Durme. 2014 · 2014
Earlier work this paper cites.
Human evaluation of grammatical error correction systems
Roman Grundkiewicz, Marcin Junczys-Dowmunt, and Edward Gillian. 2015 · 2015
Earlier work this paper cites.
Ground truth for grammatical error correction metrics
Courtney Napoles, Keisuke Sakaguchi, Matt Post, and Joel Tetreault. 2015 · 2015
Earlier work this paper cites.
Courtney Napoles, Keisuke Sakaguchi, Matt Post, and Joel Tetreault. 2016 · 2016
Earlier work this paper cites.
Reference-based metrics can be replaced with reference-less metrics in evaluating grammatical error correction systems
Hiroki Asano, Tomoya Mizumoto, and Kentaro Inui. 2017 · 2017
Earlier work this paper cites.
Automatic annotation and evaluation of error types for grammatical error correction
Christopher Bryant, Mariano Felice, and Ted Briscoe. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
SOME: Reference-less sub-metrics optimized for manual evaluations of grammatical error correction
Ryoma Yoshimura, Masahiro Kaneko, Tomoyuki Kajiwara, and Mamoru Komachi. 2020 · 2020
Earlier work this paper cites.
Is this the end of the gold standard? a straightforward reference-less grammatical error correction metric
Md Asadul Islam and Enrico Magnani. 2021 · 2021
Earlier work this paper cites.
Pathways: Asynchronous distributed dataflow for ML
Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Dan Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, Brennan Saeta, Parker Schuh, Ryan Sepassi, Laurent El Shafey, Chandramohan A. Thekkath, and Yonghui Wu. 2022 · 2022
Cited alongside, same era.
EditEval: An instruction-based benchmark for text improvements
Jane Dwivedi-Yu, Timo Schick, Zhengbao Jiang, Maria Lomeli, Patrick Lewis, Gautier Izacard, Edouard Grave, Sebastian Riedel, and Fabio Petroni. 2022 · 2022
Cited alongside, same era.
Revisiting grammatical error correction evaluation and beyond
Peiyuan Gong, Xuebo Liu, Heyan Huang, and Min Zhang. 2022 · 2022
Cited alongside, same era.
IMPARA: Impact-based metric for GEC using parallel data
Koki Maeda, Masahiro Kaneko, and Naoaki Okazaki. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Large language models understand and can be enhanced by emotional stimuli
Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, and Xing Xie. 2023 · 2023
Later among the works it cites.
G-eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a · 2023
Later among the works it cites.
Exploring effectiveness of GPT-3 in grammatical error correction: A study on performance and controllability in prompt-based methods
Mengsay Loem, Masahiro Kaneko, Sho Takase, and Naoaki Okazaki. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Peer: A collaborative language model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel. 2022 · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023 · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
Analyzing the performance of GPT-3.5 and GPT-4 in grammatical error correction
Steven Coyne, Keisuke Sakaguchi, Diana Galvan-Sosa, Michael Zock, and Kentaro Inui. 2023 · 2023
Cited alongside, same era.
Is ChatGPT a highly fluent grammatical error correction system? a comprehensive evaluation
Tao Fang, Shu Yang, Kaixin Lan, Derek F. Wong, Jinpeng Hu, Lidia S. Chao, and Yue Zhang. 2023 · 2023
Cited alongside, same era.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023 · 2023
Cited alongside, same era.
Yixiao Song, Kalpesh Krishna, Rajesh Bhatt, Kevin Gimpel, and Mohit Iyyer. 2023 · 2023
Later among the works it cites.
Evaluation metrics in the era of GPT-4: Reliably evaluating large language models on sequence to sequence tasks
Andrea Sottana, Bin Liang, Kai Zou, and Zheng Yuan. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Rating short L2 essays on the CEFR scale with GPT-4
Kevin P. Yancey, Geoffrey Laflair, Anthony Verardi, and Jill Burstein. 2023 · 2023
Later among the works it cites.
A comprehensive capability analysis of GPT-3 and GPT-3.5 series models
Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, Jie Zhou, Siming Chen, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023 · 2023
Later among the works it cites.
Revisiting meta-evaluation for grammatical error correction
Masamune Kobayashi, Masato Mita, and Mamoru Komachi. 2024 · 2024
Closest in time.
Taking the correction difficulty into account in grammatical error correction evaluation
Takumi Gotou, Ryo Nagata, Masato Mita, and Kazuaki Hanawa. 2020 · 2095
Closest in time.