Fetching the paper…
Reading the bibliography…
NLP researchers need more, higher-quality text datasets.
Sticking to the facts: Confident decoding for faithful data-to-text generation
Ran Tian, Shashi Narayan, Thibault Sellam, and Ankur P. Parikh · 1910
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Controlled hallucinations: Learning to generate faithfully from noisy data
Katja Filippova · 2010
Earlier work this paper cites.
First women, second sex: Gender bias in wikipedia
Eduardo Graells-Garrido, Mounia Lalmas, and Filippo Menczer · 2015
Earlier work this paper cites.
It’s a man’s wikipedia? assessing gender inequality in an online encyclopedia
Claudia Wagner, David García, Mohsen Jadidi, and Markus Strohmaier · 2015
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Remi Lebret, David Grangier, and Michael Auli · 2016
Earlier work this paper cites.
Hafez: an interactive poetry generation system
Marjan Ghazvininejad, Xing Shi, Jay Priyadarshi, and Kevin Knight · 2017
Earlier work this paper cites.
The e2e dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser · 2017
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré · 2017
Earlier work this paper cites.
Challenges in data-to-document generation
Sam Wiseman, Stuart Shieber, and Alexander Rush · 2017
Earlier work this paper cites.
United States Census Bureau , 2018
Educational attainment in the united states: 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Ai can be sexist and racist — it’s time to make it fair
James Zou and Londa Schiebinger · 2018
Earlier work this paper cites.
Handling divergent reference texts when evaluating table-to-text generation
Bhuwan Dhingra, Manaal Faruqui, Ankur Parikh, Ming-Wei Chang, Dipanjan Das, and William Cohen · 2019
Earlier work this paper cites.
Visual interaction with deep learning models through collaborative semantic inference
Sebastian Gehrmann, Hendrik Strobelt, Robert Krüger, Hanspeter Pfister, and Alexander M. Rush · 2019
Earlier work this paper cites.
Metaphoria: An algorithmic companion for metaphor creation
Katy Ilonka Gero and Lydia B Chilton · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov · 2019
Earlier work this paper cites.
Assessing social and intersectional biases in contextualized word representations
Yi Chern Tan and L. Elisa Celis · 2019
Earlier work this paper cites.
Gender bias in contextualized word embeddings
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Topic-preserving synthetic news generation: An adversarial deep reinforcement learning approach
Huan Liu Ahmadreza Mosallanezhad, Kai Shu · 2020
Cited alongside, same era.
ASSET: A dataset for tuning and evaluation of sentence simplification models with multiple rewriting transformations
Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoıt Sagot, and Lucia Specia · 2020
Cited alongside, same era.
Do not have enough data? deep learning to the rescue!
Esther Goldbraich Amir Kantor George Kour Segev Shlomov Naama Tepper Naama Zwerdling Ateret Anaby-Tavor, Boaz Carmeli · 2020
Cited alongside, same era.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell · 2020
Cited alongside, same era.
Text-to-text pre-training for data-to-text tasks
Mihir Kale and Abhinav Rastogi · 2020
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George Dahl · 2021
Closest in time.
The impact of multiple parallel phrase suggestions on email input and composition behaviour of native and non-native english writers
Daniel Buschek, Martin Zürn, and Malin Eiband · 2021
Closest in time.
Neural data-to-text generation with lm-based text augmentation
Ernie Chang, Xiaoyu Shen, Dawei Zhu, Vera Demberg, and Hui Su · 2021
Closest in time.
Medically aware gpt-3 as a data generator for medical dialogue summarization
Bharath Chintagunta, Namit Katariya, Xavier Amatriain, and Anitha Kannan · 2021
Closest in time.
Nine potential pitfalls when designing human-ai co-creative systems
Florian Lehmann Hai Dang Daniel Buschek, Lukas Mecke · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Susan Leavy, Barry O’Sullivan, and Eugenia Siapera · 2020
Cited alongside, same era.
ToTTo: A controlled table-to-text generation dataset
Ankur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das · 2020
Cited alongside, same era.
Training question answering models from synthetic data
Raul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary, and Bryan Catanzaro · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan · 2020
Cited alongside, same era.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh · 2020
Cited alongside, same era.
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner · 2021
Closest in time.
Human-in-the-loop for data collection: a multi-target counter narrative dataset to fight online hate speech
Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroğlu, and Marco Guerini · 2021
Closest in time.
Controlled analyses of social biases in wikipedia bios
Anjalie Field, Chan Young Park, and Yulia Tsvetkov · 2021
Closest in time.
Are synthetic clinical notes useful for real natural language processing tasks: A case study on clinical entity recognition
Xiaoqian Jiang Karthik Natarajan Serguei Vs Pakhomov Hongfang Liu Hua Xu Jianfu Li, Yujia Zhou · 2021
Closest in time.
Generative models can help writers without writing for them
Noah Madrid Kenneth Arnold, April Volzer · 2021
Closest in time.
Reusable templates and guides for documenting datasets and models for natural language processing and generation: A case study of the HuggingFace and GEM data and model cards
Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez, Pawan Sasanka Ammanamanchi, Sebastian Gehrmann, and Yacine Jernite · 2021
Closest in time.
What ingredients make for an effective crowdsourcing protocol for difficult NLU data collection tasks?
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, and Samuel R. Bowman · 2021
Closest in time.
Mitigating dataset harms requires stewardship: Lessons from 1000 papers
Kenneth L Peng, Arunesh Mathur, and Arvind Narayanan · 2021
Closest in time.
Learning compact metrics for mt, 2021
Amy Pu, Hyung Won Chung, Ankur P. Parikh, Sebastian Gehrmann, and Thibault Sellam · 2021
Closest in time.
Changing the world by changing the data
Anna Rogers · 2021
Closest in time.
We need to talk about random splits
Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, and Katja Filippova · 2021
Closest in time.
Process for adapting language models to society (PALMS) with values-targeted datasets
Irene Solaiman and Christy Dennison · 2021
Closest in time.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S. Weld · 2021
Closest in time.
Generating synthetic text data to evaluate causal inference methods
Mark Dredze Zach Wood-Doughty, Ilya Shpitser · 2021
Closest in time.