Fetching the paper…
Reading the bibliography…
Metrics are the foundation for automatic evaluation in grammatical error correction (GEC), with their evaluation of the metrics (meta-evaluation) relying on their correlation with human judgments.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
A coefficient of agreement for nominal scales
J. Cohen. 1960 · 1960
Earlier work this paper cites.
Binary Codes Capable of Correcting Deletions, Insertions and Reversals
VI Levenshtein. 1966 · 1966
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Appraise: An open-source toolkit for manual phrase-based evaluation of translations
Christian Federmann. 2010 · 2010
Earlier work this paper cites.
Evaluating performance of grammatical error detection to maximize learning effect
Ryo Nagata and Kazuhide Nakatani. 2010 · 2010
Earlier work this paper cites.
Helping our own: The HOO 2011 pilot shared task
Robert Dale and Adam Kilgarriff. 2011 · 2011
Earlier work this paper cites.
Better evaluation for grammatical error correction
Daniel Dahlmeier and Hwee Tou Ng. 2012 · 2012
Earlier work this paper cites.
HOO 2012: A report on the preposition and determiner error correction shared task
Robert Dale, Ilya Anisimoff, and George Narroway. 2012 · 2012
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013 · 2013
Earlier work this paper cites.
The CoNLL-2013 shared task on grammatical error correction
Hwee Tou Ng, Siew Mei Wu, Yuanbin Wu, Christian Hadiwinoto, and Joel Tetreault. 2013 · 2013
Earlier work this paper cites.
The AMU system in the CoNLL-2014 shared task: Grammatical error correction by data-intensive and feature-rich statistical machine translation
Marcin Junczys-Dowmunt and Roman Grundkiewicz. 2014 · 2014
Earlier work this paper cites.
Multidimensional Quality Metrics (MQM): A framework for declaring and describing translation quality metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
The CoNLL-2014 shared task on grammatical error correction
Hwee Tou Ng, Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and Christopher Bryant. 2014 · 2014
Earlier work this paper cites.
The Illinois-Columbia system in the CoNLL-2014 shared task
Alla Rozovskaya, Kai-Wei Chang, Mark Sammons, Dan Roth, and Nizar Habash. 2014 · 2014
Earlier work this paper cites.
Efficient elicitation of annotations for human evaluation of machine translation
Keisuke Sakaguchi, Matt Post, and Benjamin Van Durme. 2014 · 2014
Earlier work this paper cites.
Towards a standard evaluation method for grammatical error detection and correction
Mariano Felice and Ted Briscoe. 2015 · 2015
Earlier work this paper cites.
Human evaluation of grammatical error correction systems
Roman Grundkiewicz, Marcin Junczys-Dowmunt, and Edward Gillian. 2015 · 2015
Earlier work this paper cites.
Ground truth for grammatical error correction metrics
Courtney Napoles, Keisuke Sakaguchi, Matt Post, and Joel Tetreault. 2015 · 2015
Earlier work this paper cites.
There’s no comparison: Reference-less evaluation metrics in grammatical error correction
Courtney Napoles, Keisuke Sakaguchi, and Joel Tetreault. 2016b · 2016
Cited alongside, same era.
Reassessing the goals of grammatical error correction: Fluency instead of grammaticality
Keisuke Sakaguchi, Courtney Napoles, Matt Post, and Joel Tetreault. 2016 · 2016
Cited alongside, same era.
Reference-based metrics can be replaced with reference-less metrics in evaluating grammatical error correction systems
Hiroki Asano, Tomoya Mizumoto, and Kentaro Inui. 2017 · 2017
Cited alongside, same era.
Automatic annotation and evaluation of error types for grammatical error correction
Christopher Bryant, Mariano Felice, and Ted Briscoe. 2017 · 2017
Cited alongside, same era.
JFLEG: A fluency corpus and benchmark for grammatical error correction
Courtney Napoles, Keisuke Sakaguchi, and Joel Tetreault. 2017 · 2017
Cited alongside, same era.
A statistical analysis of summarization evaluation metrics using resampling methods
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021 · 2021
Later among the works it cites.
Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021 · 2021
Later among the works it cites.
Is this the end of the gold standard? a straightforward reference-less grammatical error correction metric
Md Asadul Islam and Enrico Magnani. 2021 · 2021
Later among the works it cites.
Neural quality estimation with multiple hypotheses for grammatical error correction
Zhenghao Liu, Xiaoyuan Yi, Maosong Sun, Liner Yang, and Tat-Seng Chua. 2021 · 2021
Later among the works it cites.
A simple recipe for multilingual grammatical error correction
Sascha Rothe, Jonathan Mallinson, Eric Malmi, Sebastian Krause, and Aliaksei Severyn. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Neural quality estimation of grammatical error correction
Shamil Chollampatt and Hwee Tou Ng. 2018a · 2018
Cited alongside, same era.
Reference-less measure of faithfulness for grammatical error correction
Leshem Choshen and Omri Abend. 2018b · 2018
Cited alongside, same era.
doccano: Text annotation tool for human
Hiroki Nakayama, Takahiro Kubo, Junya Kamura, Yasufumi Taniguchi, and Xu Liang. 2018 · 2018
Cited alongside, same era.
Parallel iterative edit models for local sequence transduction
Abhijeet Awasthi, Sunita Sarawagi, Rasna Goyal, Sabyasachi Ghosh, and Vihari Piratla. 2019 · 2019
Cited alongside, same era.
The BEA-2019 shared task on grammatical error correction
Christopher Bryant, Mariano Felice, Øistein E. Andersen, and Ted Briscoe. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
LM-critic: Language models for unsupervised grammatical error correction
Michihiro Yasunaga, Jure Leskovec, and Percy Liang. 2021 · 2021
Later among the works it cites.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Revisiting grammatical error correction evaluation and beyond
Peiyuan Gong, Xuebo Liu, Heyan Huang, and Min Zhang. 2022 · 2022
Later among the works it cites.
IMPARA: Impact-based metric for GEC using parallel data
Koki Maeda, Masahiro Kaneko, and Naoaki Okazaki. 2022 · 2022
Later among the works it cites.
Ensembling and knowledge distilling of large sequence taggers for grammatical error correction
Maksym Tarnavskyi, Artem Chernodub, and Kostiantyn Omelianchuk. 2022 · 2022
Later among the works it cites.
Grammatical error correction: A survey of the state of the art
Christopher Bryant, Zheng Yuan, Muhammad Reza Qorib, Hannan Cao, Hwee Tou Ng, and Ted Briscoe. 2023 · 2023
Later among the works it cites.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Later among the works it cites.
Analyzing the performance of GPT-3.5 and GPT-4 in grammatical error correction
Steven Coyne, Keisuke Sakaguchi, Diana Galvan-Sosa, Michael Zock, and Kentaro Inui. 2023 · 2023
Later among the works it cites.
TransGEC: Improving grammatical error correction with translationese
Tao Fang, Xuebo Liu, Derek F. Wong, Runzhe Zhan, Liang Ding, Lidia S. Chao, Dacheng Tao, and Min Zhang. 2023 · 2023
Later among the works it cites.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Later among the works it cites.
TemplateGEC: Improving grammatical error correction with detection template
Yinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong, Derek F. Wong, Yang Gao, Heyan Huang, and Min Zhang. 2023 · 2023
Later among the works it cites.
Yixin Liu and Alexander R. Fabbri. 2023 · 2023
Later among the works it cites.
Benchmarking large language model capabilities for conditional generation
Joshua Maynez, Priyanka Agrawal, and Sebastian Gehrmann. 2023 · 2023
Later among the works it cites.
Taking the correction difficulty into account in grammatical error correction evaluation
Takumi Gotou, Ryo Nagata, Masato Mita, and Kazuaki Hanawa. 2020 · 2095
Closest in time.