Fetching the paper…
Reading the bibliography…
Evaluation in machine learning is usually informed by past choices, for example which datasets or metrics to use.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Mlsum: The multilingual summarization corpus
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2020 · 2004
Earlier work this paper cites.
The OPUS corpus - parallel and free: http://logos.uio.no/opus
Jörg Tiedemann and Lars Nygaard. 2004 · 2004
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Simpitiki: a simplification corpus for italian
Sara Tonelli, Alessio Palmero Aprosio, and Francesca Saltori. 2016 · 2016
Earlier work this paper cites.
Optimizing statistical machine translation for text simplification
Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen, and Chris Callison-Burch. 2016 · 2016
Earlier work this paper cites.
The E2E dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
Challenges in data-to-document generation
Sam Wiseman, Stuart Shieber, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Open subtitles paraphrase corpus for six languages
Mathias Creutz. 2018 · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
BLEU is not suitable for the evaluation of text simplification
Elior Sulem, Omri Abend, and Ari Rappoport. 2018 · 2018
Earlier work this paper cites.
Constrained decoding for neural NLG from compositional representations in task-oriented dialogue
Anusha Balakrishnan, Jinfeng Rao, Kartikeya Upasani, Michael White, and Rajen Subba. 2019 · 2019
Earlier work this paper cites.
Taskmaster-1: Toward a realistic and diverse dialog dataset
Bill Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, and Andy Cedilnik. 2019 · 2019
Earlier work this paper cites.
Semantic noise matters for neural natural language generation
Ondřej Dušek, David M. Howcroft, and Verena Rieser. 2019 · 2019
Earlier work this paper cites.
Neural generation for Czech: Data and baselines
Ondřej Dušek and Filip Jurčíček. 2019 · 2019
Earlier work this paper cites.
Findings of the third workshop on neural generation and translation
Hiroaki Hayashi, Yusuke Oda, Alexandra Birch, Ioannis Konstas, Andrew Finch, Minh-Thang Luong, Graham Neubig, and Katsuhito Sudoh. 2019 · 2019
Earlier work this paper cites.
ViGGO: A video game corpus for data-to-text generation in open-domain conversation
Juraj Juraska, Kevin Bowden, and Marilyn Walker. 2019 · 2019
Earlier work this paper cites.
Template-free data-to-text generation of Finnish sports news
Jenna Kanerva, Samuel Rönnqvist, Riina Kekki, Tapio Salakoski, and Filip Ginter. 2019 · 2019
Earlier work this paper cites.
Generating summaries with topic templates and structured convolutional decoders
Laura Perez-Beltrachini, Yang Liu, and Mirella Lapata. 2019 · 2019
Earlier work this paper cites.
Naver labs Europe’s systems for the document-level generation and translation task at WNGT 2019
Fahimeh Saleh, Alexandre Berard, Ioan Calapodescu, and Laurent Besacier. 2019 · 2019
Earlier work this paper cites.
Best practices for the human evaluation of automatically generated text
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019 · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
ASSET: A dataset for tuning and evaluation of sentence simplification models with multiple rewriting transformations
Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, and Lucia Specia. 2020 · 2020
Cited alongside, same era.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Evaluating the state-of-the-art of end-to-end natural language generation: The e2e nlg challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Utility is in the eye of the user: A critique of NLP leaderboards
Kawin Ethayarajh and Dan Jurafsky. 2020 · 2020
Cited alongside, same era.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Paragraph-level simplification of medical texts
Ashwin Devaraj, Iain Marshall, Byron Wallace, and Junyi Jessy Li. 2021 · 2021
Later among the works it cites.
Nl-augmenter: A framework for task-sensitive natural language augmentation
Kaustubh D Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Srivastava, Samson Tan, et al. 2021 · 2021
Later among the works it cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Later among the works it cites.
XL-sum: Large-scale multilingual abstractive summarization for 44 languages
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural CRF model for sentence alignment in text simplification
Chao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong, and Wei Xu. 2020 · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Text-to-text pre-training for data-to-text tasks
Mihir Kale and Abhinav Rastogi. 2020 · 2020
Cited alongside, same era.
Turku enhanced parser pipeline: From raw text to enhanced graphs in the IWPT 2020 shared task
Jenna Kanerva, Filip Ginter, and Sampo Pyysalo. 2020 · 2020
Cited alongside, same era.
WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020 · 2020
Cited alongside, same era.
CommonGen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020 · 2020
Cited alongside, same era.
The third multilingual surface realisation shared task (SR’20): Overview and evaluation results
Simon Mille, Anya Belz, Bernd Bohnet, Thiago Castro Ferreira, Yvette Graham, and Leo Wanner. 2020 · 2020
Cited alongside, same era.
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Later among the works it cites.
BARThez: a skilled pretrained French sequence-to-sequence model
Moussa Kamal Eddine, Antoine Tixier, and Michalis Vazirgiannis. 2021 · 2021
Later among the works it cites.
Finnish paraphrase corpus
Jenna Kanerva, Filip Ginter, Li-Hsin Chang, Iiro Rastas, Valtteri Skantsi, Jemina Kilpeläinen, Hanna-Mari Kupari, Jenna Saarni, Maija Sevón, and Otto Tarkka. 2021 · 2021
Later among the works it cites.
BiSECT: Learning to split and rephrase sentences with bitexts
Joongwon Kim, Mounica Maddela, Reno Kriz, Wei Xu, and Chris Callison-Burch. 2021a · 2021
Later among the works it cites.
“how robust r u?”: Evaluating task-oriented dialogue systems on spoken conversations
Seokhwan Kim, Yang Liu, Di Jin, Alexandros Papangelis, Karthik Gopalakrishnan, Behnam Hedayatnia, and Dilek Z. Hakkani-Tür. 2021b · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
Reusable templates and guides for documenting datasets and models for natural language processing and generation: A case study of the HuggingFace and GEM data and model cards
Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez, Pawan Sasanka Ammanamanchi, Sebastian Gehrmann, and Yacine Jernite. 2021 · 2021
Later among the works it cites.
Automatic construction of evaluation suites for natural language generation datasets
Simon Mille, Kaustubh Dhole, Saad Mahamood, Laura Perez-Beltrachini, Varun Gangal, Mihir Kale, Emiel van Miltenburg, and Sebastian Gehrmann. 2021 · 2021
Later among the works it cites.
DART: Open-domain structured data record to text generation
Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani. 2021 · 2021
Later among the works it cites.
Models and datasets for cross-lingual summarisation
Laura Perez-Beltrachini and Mirella Lapata. 2021 · 2021
Later among the works it cites.
Data-to-text generation with macro planning
Ratish Puduppully and Mirella Lapata. 2021 · 2021
Later among the works it cites.
AI and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada. 2021 · 2021
Later among the works it cites.
D2S: Document-to-slide generation via query-based text summarization
Edward Sun, Yufang Hou, Dakuo Wang, Yunfeng Zhang, and Nancy X. R. Wang. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
Daniel Deutsch and Dan Roth. 2022 · 2022
Closest in time.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Closest in time.
Bidimensional leaderboards: Generate and evaluate language hand in hand
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander R. Fabbri, Yejin Choi, and Noah A. Smith. 2022 · 2022
Closest in time.
Data cards: Purposeful and transparent dataset documentation for responsible ai
Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. 2022 · 2022
Closest in time.
SQuALITY: Building a long-document summarization dataset the hard way
Alex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang, and Samuel R. Bowman. 2022 · 2022
Closest in time.
Fantastic questions and where to find them: FairytaleQA – an authentic dataset for narrative comprehension
Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, and Mark Warschauer. 2022 · 2022
Closest in time.
Data-to-text generation with entity modeling
Ratish Puduppully, Li Dong, and Mirella Lapata. 2019a · 2035
Closest in time.