Fetching the paper…
Reading the bibliography…
Modern neural language models can produce remarkably fluent and grammatical text.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Logic and conversation
Herbert P Grice. 1975 · 1975
Earlier work this paper cites.
Scripts, plans, goals, and understanding: An inquiry into human knowledge structures
Roger C Schank and Robert P Abelson. 1977 · 1977
Earlier work this paper cites.
Light-weight entailment checking for computational semantics
Christof Monz and Maarten de Rijke. 2001 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Conceptnet—a practical commonsense reasoning tool-kit
Hugo Liu and Push Singh. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
A survey of evaluation metrics used for nlg systems
Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2020 · 2008
Earlier work this paper cites.
Roft: A tool for evaluating human detection of machine-generated text
Liam Dugan, Daphne Ippolito, Arun Kirubarajan, and Chris Callison-Burch. 2020 · 2010
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2020 · 2012
Earlier work this paper cites.
Mctest: A challenge dataset for the open-domain machine comprehension of text
Matthew Richardson, Christopher JC Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
The unified and holistic method gamma ( γ \gamma ) for inter-annotator agreement measure and alignment
Yann Mathet, Antoine Widlöcher, and Jean-Philippe Métivier. 2015 · 2015
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Learning to write with cooperative discriminators
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Content analysis: An introduction to its methodology
Klaus Krippendorff. 2018 · 2018
Cited alongside, same era.
Rankme: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondrej Dusek, and Verena Rieser. 2018 · 2018
Cited alongside, same era.
Rethinking engagement with online news through social and visual co-annotation
Gavin Wood, Kiel Long, Tom Feltwell, Scarlett Rowland, Phillip Brooker, Jamie Mahoney, John Vines, Julie Barnett, and Shaun Lawson. 2018 · 2018
Cited alongside, same era.
Unifying human and statistical evaluation for natural language generation
Tatsunori Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Later among the works it cites.
Stanza: A python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020 · 2020
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2021 · 2021
Closest in time.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Entity, relation, and event extraction with contextualized span representations
David Wadden, Ulme Wennberg, Yi Luan, and Hannaneh Hajishirzi. 2019 · 2019
Cited alongside, same era.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Gpt-3 creative fiction
Gwern Branwen. 2020 · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin. 2020 · 2020
Cited alongside, same era.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020 · 2020
Cited alongside, same era.
Closest in time.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021 · 2021
Closest in time.
Choose your own adventure: Paired suggestions in collaborative writing for evaluating story generation models
Elizabeth Clark and Noah A. Smith. 2021 · 2021
Closest in time.
Perception score: A learned metric for open-ended text generation evaluation
Jing Gu, Qing yang Wu, and Zhou Yu. 2021 · 2021
Closest in time.
Genie: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, and Daniel S. Weld. 2021 · 2021
Closest in time.
Mauve: Human-machine divergence curves for evaluating open-ended text generation
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Closest in time.
Prompt programming for large language models: Beyond the few-shot paradigm
Laria Reynolds and Kyle McDonell. 2021 · 2021
Closest in time.
pygamma-agreement: Gamma γ \gamma measure for inter/intra-annotator agreement in python
Hadrien Titeux and Rachid Riad. 2021 · 2021
Closest in time.
TuringAdvice: A generative and dynamic evaluation of language use
Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi. 2021 · 2021
Closest in time.