Fetching the paper…
Reading the bibliography…
Existing data-to-text generation datasets are mostly limited to English.
Neural generation for czech: Data and baselines
Ondřej Dušek and Filip Jurčíček. 2019 · 1910
Earlier work this paper cites.
PostGraphe: A system for the generation of statistical graphics and text
Massimo Fasciano and Guy Lapalme. 1996 · 1996
Earlier work this paper cites.
Building applied natural language generation systems
Ehud Reiter and Robert Dale. 1997 · 1997
Earlier work this paper cites.
Describing complex charts in natural language: A caption generation system
Vibhu O. Mittal, Johanna D. Moore, Giuseppe Carenini, and Steven Roth. 1998 · 1998
Earlier work this paper cites.
Translationese—a myth or an empirical fact?: A study into the linguistic identifiability of translated language
Sonja Tirkkonen-Condit. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Lessons from deploying nlg technology for marine weather forecast text generation
Somayajulu G Sripada, Ehud Reiter, Ian Davy, and Kristian Nilssen. 2004 · 2004
Earlier work this paper cites.
Summarizing information graphics textually
Seniz Demir, Sandra Carberry, and Kathleen F. McCoy. 2012 · 2012
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Rémi Lebret, David Grangier, and Michael Auli. 2016 · 2016
Earlier work this paper cites.
The E2E dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
Challenges in data-to-document generation
Sam Wiseman, Stuart Shieber, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
Can neural generators for dialogue learn sentence planning and discourse structuring?
Lena Reed, Shereen Oraby, and Marilyn Walker. 2018 · 2018
Earlier work this paper cites.
Handling divergent reference texts when evaluating table-to-text generation
Bhuwan Dhingra, Manaal Faruqui, Ankur Parikh, Ming-Wei Chang, Dipanjan Das, and William Cohen. 2019 · 2019
Earlier work this paper cites.
Semantic noise matters for neural natural language generation
Ondřej Dušek, David M. Howcroft, and Verena Rieser. 2019 · 2019
Earlier work this paper cites.
Findings of the third workshop on neural generation and translation
Hiroaki Hayashi, Yusuke Oda, Alexandra Birch, Ioannis Konstas, Andrew Finch, Minh-Thang Luong, Graham Neubig, and Katsuhito Sudoh. 2019 · 2019
Earlier work this paper cites.
Template-free data-to-text generation of Finnish sports news
Jenna Kanerva, Samuel Rönnqvist, Riina Kekki, Tapio Salakoski, and Filip Ginter. 2019 · 2019
Earlier work this paper cites.
Massive vs. curated embeddings for low-resourced languages: the case of Yorùbá and Twi
Jesujoba Alabi, Kwabena Amponsah-Kaakyire, David Adelani, and Cristina España-Bonet. 2020 · 2020
Earlier work this paper cites.
Best practices for data-efficient modeling in NLG:how to train production-ready neural models with less data
Ankit Arun, Soumya Batra, Vikas Bhardwaj, Ashwini Challa, Pinar Donmez, Peyman Heidari, Hakan Inan, Shashank Jain, Anuj Kumar, Shawn Mei, Karthik Mohan, and Michael White. 2020 · 2020
Earlier work this paper cites.
How human is machine translationese? comparing human and machine translations of text and speech
Yuri Bizzoni, Tom S Juzek, Cristina España-Bonet, Koel Dutta Chowdhury, Josef van Genabith, and Elke Teich. 2020 · 2020
Cited alongside, same era.
Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. 2020 · 2020
Cited alongside, same era.
Controlled hallucinations: Learning to generate faithfully from noisy data
Katja Filippova. 2020 · 2020
Cited alongside, same era.
Statistical power and translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2020 · 2020
Cited alongside, same era.
Participatory research for low-resourced machine translation: A case study in African languages
Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. 2020 · 2020
AfroMT: Pretraining strategies and reproducible benchmarks for translation of 8 African languages
Machel Reid, Junjie Hu, Graham Neubig, and Yutaka Matsuo. 2021 · 2021
Later among the works it cites.
XTREME-R: Towards more challenging and nuanced multilingual evaluation
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021 · 2021
Later among the works it cites.
Towards table-to-text generation with numerical reasoning
Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura. 2021 · 2021
Later among the works it cites.
MassiveSumm: a very large-scale, very multilingual, news summarisation dataset
Daniel Varab and Natalie Schluter. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chart-to-text: Generating natural language descriptions for charts by adapting the transformer model
Jason Obeid and Enamul Hoque. 2020 · 2020
Cited alongside, same era.
ToTTo: A controlled table-to-text generation dataset
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020 · 2020
Cited alongside, same era.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
SportSett:basketball - a robust and maintainable data-set for natural language generation
Craig Thomson, Ehud Reiter, and Somayajulu Sripada. 2020 · 2020
Cited alongside, same era.
The washington post to debut ai-powered audio updates for 2020 election results
The Washington Post. 2020 · 2020
Cited alongside, same era.
Masakhaner: Named entity recognition for african languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, et al. 2021 · 2021
Cited alongside, same era.
Ann Yuan, Daphne Ippolito, Vitaly Nikolaev, Chris Callison-Burch, Andy Coenen, and Sebastian Gehrmann. 2021 · 2021
Later among the works it cites.
Towards afrocentric NLP for African languages: Where we are and where we can go
Ife Adebara and Muhammad Abdul-Mageed. 2022 · 2022
Closest in time.
A few thousand translations go a long way! leveraging pre-trained models for African news translation
David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, and Sam Manthalu. 2022 · 2022
Closest in time.
Multi task learning for zero shot performance prediction of multilingual models
Kabir Ahuja, Shanu Kumar, Sandipan Dandapat, and Monojit Choudhury. 2022 · 2022
Closest in time.
uFACT: Unfaithful alien-corpora training for semantically consistent data-to-text generation
Tisha Anders, Alexandru Coca, and Bill Byrne. 2022 · 2022
Closest in time.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang. 2022 · 2022
Closest in time.
Dataset geography: Mapping language data to language users
Fahim Faisal, Yinkai Wang, and Antonios Anastasopoulos. 2022 · 2022
Closest in time.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Closest in time.
Chart-to-text: A large-scale benchmark for chart summarization
Shankar Kantharaj, Rixie Tiffany Leong, Xiang Lin, Ahmed Masry, Megh Thakkar, Enamul Hoque, and Shafiq Joty. 2022 · 2022
Closest in time.
Report from the nsf future directions workshop on automatic evaluation of dialog: Research directions and challenges
Shikib Mehri, Jinho Choi, Luis Fernando D’Haro, Jan Deriu, Maxine Eskenazi, Milica Gasic, Kallirroi Georgila, Dilek Hakkani-Tur, Zekang Li, Verena Rieser, Samira Shaikh, David Traum, Yi-Ting Yeh, Zhou Yu, Yizhe Zhang, and Chen Zhang. 2022 · 2022
Closest in time.
Improving compositional generalization with self-training for data-to-text generation
Sanket Vaibhav Mehta, Jinfeng Rao, Yi Tay, Mihir Kale, Ankur Parikh, and Emma Strubell. 2022 · 2022
Closest in time.
NaijaSenti: A nigerian Twitter sentiment corpus for multilingual sentiment analysis
Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Sa’id Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Alípio Jorge, and Pavel Brazdil. 2022 · 2022
Closest in time.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed H. Chi, and Quoc Le. 2022 · 2022
Closest in time.
How do Seq2Seq models perform on end-to-end data-to-text generation?
Xunjian Yin and Xiaojun Wan. 2022 · 2022
Closest in time.