Fetching the paper…
Reading the bibliography…
In this work, we explore a useful but often neglected methodology for robustness analysis of text generation evaluation metrics: stress tests with synthetic data.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Analyzing and modeling rank data
John I Marden. 1995 · 1995
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
The CoNLL-2014 shared task on grammatical error correction
Hwee Tou Ng, Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and Christopher Bryant. 2014 · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Seq2sick: Evaluating the robustness of sequence-to-sequence models with adversarial examples
Minhao Cheng, Jinfeng Yi, Huan Zhang, Pin-Yu Chen, and Cho-Jui Hsieh. 2018 · 2018
Earlier work this paper cites.
The multitarget ted talks task
Kevin Duh. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Earlier work this paper cites.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Generating informative and diverse conversational responses via adversarial information maximization
Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018 · 2018
Earlier work this paper cites.
Analysis Methods in Neural Language Processing: A Survey
Yonatan Belinkov and James Glass. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Detecting egregious responses in neural sequence-to-sequence models
Tianxing He and James Glass. 2019 · 2019
Earlier work this paper cites.
Large-scale, diverse, paraphrastic bitexts via sampling and clustering
J. Edward Hu, Abhinav Singh, Nils Holzenberger, Matt Post, and Benjamin Van Durme. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Earlier work this paper cites.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Earlier work this paper cites.
Perturbation sensitivity analysis to detect unintended model biases
Vinodkumar Prabhakaran, Ben Hutchinson, and Margaret Mitchell. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Are red roses red? evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
What’s in a name? are BERT named entity representations just as good for any other name?
Sriram Balasubramanian, Naman Jain, Gaurav Jindal, Abhijeet Awasthi, and Sunita Sarawagi. 2020 · 2020
Earlier work this paper cites.
Curious case of language generation evaluation metrics: A cautionary tale
Ozan Caglayan, Pranava Madhyastha, and Lucia Specia. 2020 · 2020
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Cited alongside, same era.
BERT-ATTACK: Adversarial attack against BERT using BERT
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020 · 2020
Cited alongside, same era.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Exposure bias versus self-recovery: Are distortions really incremental for autoregressive text generation?
Tianxing He, Jingzhao Zhang, Zhiming Zhou, and James Glass. 2021 · 2021
Later among the works it cites.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Global explainability of BERT-based evaluation metrics by disentangling along linguistic factors
Marvin Kaster, Wei Zhao, and Steffen Eger. 2021 · 2021
Later among the works it cites.
DExperts: Decoding-time controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Cited alongside, same era.
USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020 · 2020
Cited alongside, same era.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020 · 2020
Cited alongside, same era.
Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks?
Thang Pham, Trung Bui, Long Mai, and Anh Nguyen. 2021 · 2021
Later among the works it cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaïd Harchaoui. 2021 · 2021
Later among the works it cites.
XTREME-R: Towards more challenging and nuanced multilingual evaluation
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021 · 2021
Later among the works it cites.
Findings of the WMT 2021 shared task on quality estimation
Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. 2021 · 2021
Later among the works it cites.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Probing Classifiers: Promises, Shortcomings, and Advances
Yonatan Belinkov. 2022 · 2022
Closest in time.
Survey of low-resource machine translation
Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindřich Helcl, and Alexandra Birch. 2022 · 2022
Closest in time.
Controlling the focus of pretrained language generation models
Jiabao Ji, Yoon Kim, James Glass, and Tianxing He. 2022 · 2022
Closest in time.
Transparent human evaluation for image captioning
Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras, Yejin Choi, and Noah A. Smith. 2022a · 2022
Closest in time.
Bidimensional leaderboards: Generate and evaluate language hand in hand
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, and Noah A. Smith. 2022b · 2022
Closest in time.
CTRLEval: An unsupervised reference-free metric for evaluating controlled text generation
Pei Ke, Hao Zhou, Yankai Lin, Peng Li, Jie Zhou, Xiaoyan Zhu, and Minlie Huang. 2022 · 2022
Closest in time.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022 · 2022
Closest in time.
On the evaluation metrics for paraphrase generation
Lingfeng Shen, Lemao Liu, Haiyun Jiang, and Shuming Shi. 2022 · 2022
Closest in time.
Dialect-robust evaluation of generated text
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, and Sebastian Gehrmann. 2022 · 2022
Closest in time.
Layer or representation space: What makes BERT-based evaluation metrics robust?
Doan Nam Long Vu, Nafise Sadat Moosavi, and Steffen Eger. 2022 · 2022
Closest in time.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022 · 2022
Closest in time.
Not all errors are equal: Learning text generation metrics using stratified error synthesis
Wenda Xu, Yi-lin Tuan, Yujie Lu, Michael Saxon, Lei Li, and William Yang Wang. 2022 · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022 · 2022
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2022
Closest in time.
Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization
Yiran Chen, Pengfei Liu, and Xipeng Qiu. 2021 · 2095
Closest in time.