Fetching the paper…
Reading the bibliography…
Human evaluation plays a crucial role in Natural Language Processing (NLP) as it assesses the quality and relevance of developed systems, thereby facilitating their enhancement.
Trading off diversity and quality in natural language generation
Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan. 2020 · 2004
Earlier work this paper cites.
Principles of evaluation in natural language processing
Patrick Paroubek, Stéphane Chaudiron, and Lynette Hirschman. 2007 · 2007
Earlier work this paper cites.
Intrinsic vs. extrinsic evaluation measures for referring expression generation
Anja Belz and Albert Gatt. 2008 · 2008
Earlier work this paper cites.
Speech and language processing
Constituency Parsing. 2009 · 2009
Earlier work this paper cites.
The Joanna Briggs Institute reviewers’ manual 2015: methodology for JBI scoping reviews , chapter 11. The Joanna Briggs Institute
Micah DJ Peters, Christina M Godfrey, Patricia McInerney, Cassia Baldini Soares, Hanan Khalil, and Deborah Parker. 2015 · 2015
Earlier work this paper cites.
Evaluation methodologies in automatic question generation 2013-2018
Jacopo Amidei, Paul Piwek, and Alistair Willis. 2018a · 2018
Earlier work this paper cites.
Set to ordered text: Generating discharge instructions from medical billing codes
Litton J Kurisinkel and Nancy Chen. 2019 · 2019
Earlier work this paper cites.
Best practices for the human evaluation of automatically generated text
Chris Van Der Lee, Albert Gatt, Emiel Van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019 · 2019
Earlier work this paper cites.
Twenty years of confusion in human evaluation: Nlg needs evaluation sheets and standardised definition
David Howcroft, Anya Belz, Miruna Clinciu, Dimitra Gkatzia, Sadid A Hasan, Saad Mahamood, Simon Mille, Emiel Van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Counterfactual off-policy training for neural dialogue generation
Qingfu Zhu, Weinan Zhang, Ting Liu, and William Yang Wang. 2020 · 2020
Cited alongside, same era.
Towards document-level human mt evaluation: On the issues of annotator agreement, effort and misevaluation
Sheila Castilho. 2021 · 2021
Cited alongside, same era.
Reliability of human evaluation for text summarization: Lessons learned and challenges ahead
Neslihan Iskender, Tim Polzehl, and Sebastian Möller. 2021 · 2021
Cited alongside, same era.
Towards standard criteria for human evaluation of chatbots: A survey
Hongru Liang and Huaqing Li. 2021 · 2021
Cited alongside, same era.
Sleepqa: A health coaching dataset on sleep for extractive question answering
Iva Bojic, Qi Chwen Ong, Megh Thakkar, Esha Kamran, Irving Yu Le Shua, Jaime Rei Ern Pang, Jessica Chen, Vaaruni Nayak, Shafiq Joty, and Josip Car. 2022 · 2022
Later among the works it cites.
Nareor: The narrative reordering problem
Varun Gangal, Steven Y Feng, Malihe Alikhani, Teruko Mitamura, and Eduard Hovy. 2022 · 2022
Later among the works it cites.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Yixin Liu, Alexander R Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, et al. 2022 · 2022
Later among the works it cites.
Flexible generation from fragmentary linguistic input
Peng Qian and Roger Levy. 2022 · 2022
Later among the works it cites.
Translating hanja historical documents to contemporary korean and english
Juhee Son, Jiho Jin, Haneul Yoo, JinYeong Bak, Kyunghyun Cho, and Alice Oh. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is this translation error critical?: Classification-based human and automatic machine translation evaluation focusing on critical errors
Katsuhito Sudoh, Kosuke Takahashi, and Satoshi Nakamura. 2021 · 2021
Cited alongside, same era.
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021 · 2021
Cited alongside, same era.
An investigation of suitability of pre-trained language models for dialogue generation–avoiding discrepancies
Yan Zeng and Jian-Yun Nie. 2021 · 2021
Cited alongside, same era.
Rethinking the agreement in human evaluation tasks
Jacopo Amidei, Paul Piwek, and Alistair Willis. 2018b
Cited in the paper.
A data-centric framework for improving domain-specific machine reading comprehension datasets
Iva Bojic, Josef Halim, Verena Suharman, Sreeja Tar, Qi Chwen Ong, Duy Phung, Mathieu Ravaut, Shafiq Joty, and Josip Car. 2023a
Cited in the paper.
Iva Bojic, Qi Chwen Ong, Shafiq Joty, and Josip Car. 2023b
Cited in the paper.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2023 · 2023
Closest in time.
Toward verifiable and reproducible human evaluation for text-to-image generation
Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, and Shin’ichi Satoh. 2023 · 2023
Closest in time.