Fetching the paper…
Reading the bibliography…
This paper reviews and summarizes human evaluation practices described in 97 style transfer papers with respect to three main evaluation aspects: style transfer, meaning preservation, and fluency.
Absolute identification by relative judgment
Neil Stewart, Gordon D. A. Brown, and Nick Chater. 2005 · 2005
Earlier work this paper cites.
(Meta-) Evaluation of Machine Translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
A survey on text simplification
Punardeep Sikka, Manmeet Singh, Allen Pink, and Vijay Mago. 2020 · 2008
Earlier work this paper cites.
Learning with annotation noise
Eyal Beigman and Beata Beigman Klebanov. 2009 · 2009
Earlier work this paper cites.
Continuous Measurement Scales in Human Evaluation of Machine Translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Predicting grammaticality on an ordinal scale
Michael Heilman, Aoife Cahill, Nitin Madnani, Melissa Lopez, Matthew Mulholland, and Joel Tetreault. 2014 · 2014
Earlier work this paper cites.
SemEval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation
Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2016 · 2016
Earlier work this paper cites.
A persona-based neural conversation model
Jiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
An empirical analysis of formality in online communication
Ellie Pavlick and Joel Tetreault. 2016 · 2016
Earlier work this paper cites.
Optimizing statistical machine translation for text simplification
Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen, and Chris Callison-Burch. 2016 · 2016
Earlier work this paper cites.
Style transfer from non-parallel text by cross-alignment
Tianxiao Shen, Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2017 · 2017
Earlier work this paper cites.
Delete, retrieve, generate: a simple approach to sentiment and style transfer
Juncen Li, Robin Jia, He He, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer
Sudha Rao and Joel Tetreault. 2018 · 2018
Cited alongside, same era.
Making better use of the crowd: How crowdsourcing can advance machine learning research
Jennifer Wortman Vaughan. 2018 · 2018
Cited alongside, same era.
Imat: Unsupervised text attribute transfer via iterative matching and translation
Z. Jin, D. Jin, J. Mueller, N. Matthews, and E. Santus. 2019 · 2019
Cited alongside, same era.
Domain adaptive text style transfer
Dianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, and Ming-Ting Sun. 2019 · 2019
Cited alongside, same era.
Evaluating style transfer for text
Remi Mir, Bjarke Felbo, Nick Obradovich, and Iyad Rahwan. 2019 · 2019
Cited alongside, same era.
Unsupervised evaluation metrics and learning criteria for non-parallel textual transfer
Annotation-based semantics
Kiyong Lee. 2020 · 2020
Later among the works it cites.
PowerTransformer: Unsupervised controllable revision for biased language correction
Xinyao Ma, Maarten Sap, Hannah Rashkin, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Automatically neutralizing subjective bias in text
Reid Pryzant, Richard Diehl Martinez, Nathan Dass, Sadao Kurohashi, Dan Jurafsky, and Diyi Yang. 2020 · 2020
Later among the works it cites.
“This is a problem, don’t you agree?” framing and bias in human evaluation for natural language generation
Stephanie Schoch, Diyi Yang, and Yangfeng Ji. 2020 · 2020
Later among the works it cites.
A systematic review of reproducibility research in natural language processing
Anya Belz, Shubham Agarwal, Anastasia Shimorina, and Ehud Reiter. 2021 · 2021
Closest in time.
Xformal: A benchmark for multilingual formality style transfer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richard Yuanzhe Pang and Kevin Gimpel. 2019 · 2019
Cited alongside, same era.
Data-driven sentence simplification: Survey and benchmark
Fernando Alva-Manchego, Carolina Scarton, and Lucia Specia. 2020 · 2020
Cited alongside, same era.
Disentangling the properties of human evaluation methods: A classification system to support comparability, meta-evaluation and reproducibility testing
Anya Belz, Simon Mille, and David M. Howcroft. 2020 · 2020
Cited alongside, same era.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Reformulating unsupervised style transfer as paraphrase generation
Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020 · 2020
Cited alongside, same era.
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2020 · 2020
Cited alongside, same era.
Eleftheria Briakou, Di Lu, Ke Zhang, and Joel Tetreault. 2021 · 2021
Closest in time.
The great misalignment problem in human evaluation of NLP methods
Mika Hämäläinen and Khalid Alnajjar. 2021 · 2021
Closest in time.
Deep learning for text style transfer: A survey
Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, and Rada Mihalcea. 2021 · 2021
Closest in time.
Interrater disagreement resolution: A systematic procedure to reach consensus in annotation tasks
Yvette Oortwijn, Thijs Ossenkoppele, and Arianna Betti. 2021 · 2021
Closest in time.
Anastasia Shimorina and Anya Belz. 2021 · 2021
Closest in time.
Beyond fair pay: Ethical implications of nlp crowdsourcing
Boaz Shmueli, Jan Fell, Soumya Ray, and Lun-Wei Ku. 2021 · 2021
Closest in time.