Fetching the paper…
Reading the bibliography…
We present Dolphin, a novel benchmark that addresses the need for a natural language generation (NLG) evaluation framework dedicated to the wide collection of Arabic languages and varieties.
Neural arabic question answering
Hussein Mozannar, Karl El Hajal, Elie Maamary, and Hazem Hajj. 2019 · 1906
Earlier work this paper cites.
Anetac: Arabic named entity transliteration and classification dataset
Mohamed Seghir Hadj Ameur, Farid Meziane, and Ahmed Guessoum. 2019 · 1907
Earlier work this paper cites.
Mlqa: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
From wer and ril to mer and wil: improved evaluation measures for connected speech recognition
Andrew Morris, Viktoria Maier, and Phil Green. 2004 · 2004
Earlier work this paper cites.
MultiUN: A multilingual corpus from united nation documents
Andreas Eisele and Yu Chen. 2010 · 2010
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020 · 2010
Earlier work this paper cites.
Better evaluation for grammatical error correction
Daniel Dahlmeier and Hwee Tou Ng. 2012 · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
Machine translation of Arabic dialects
Rabih Zbib, Erika Malchiodi, Jacob Devlin, David Stallard, Spyros Matsoukas, Richard Schwartz, John Makhoul, Omar Zaidan, and Chris Callison-Burch. 2012 · 2012
Earlier work this paper cites.
Arabizi detection and conversion to arabic
Kareem Darwish. 2013 · 2013
Earlier work this paper cites.
A Multidialectal Parallel Corpus of Arabic
Houda Bouamor, Nizar Habash, and Kemal Oflazer. 2014 · 2014
Earlier work this paper cites.
Development of a TV broadcasts speech recognition system for qatari Arabic
Mohamed Elmahdy, Mark Hasegawa-Johnson, and Eiman Mustafawi. 2014 · 2014
Earlier work this paper cites.
The first QALB shared task on automatic text correction for Arabic
Behrang Mohit, Alla Rozovskaya, Nizar Habash, Wajdi Zaghouani, and Ossama Obeid. 2014 · 2014
Earlier work this paper cites.
Collecting natural sms and chat conversations in multiple languages: The bolt phase 2 corpus
Zhiyi Song, Stephanie M Strassel, Haejoong Lee, Kevin Walker, Jonathan Wright, Jennifer Garland, Dana Fore, Brian Gainor, Preston Cabe, Thomas Thomas, et al. 2014 · 2014
Earlier work this paper cites.
The second QALB shared task on automatic text correction for Arabic
Alla Rozovskaya, Houda Bouamor, Nizar Habash, Wajdi Zaghouani, Ossama Obeid, and Behrang Mohit. 2015 · 2015
Earlier work this paper cites.
The united nations parallel corpus v1. 0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016 · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
The Madar Arabic Dialect Corpus and Lexicon
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, et al. 2018 · 2018
Earlier work this paper cites.
Dawqas: A dataset for arabic why question answering system
Walaa Ismail and Masun Nabhan Homsi. 2018 · 2018
Cited alongside, same era.
Design Challenges in Named Entity Transliteration
Yuval Merhav and Stephen Ash. 2018 · 2018
Cited alongside, same era.
Dial2msa: A tweets corpus for converting dialectal arabic to modern standard arabic
Hamdy Mubarak. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Towards building arabic paraphrasing benchmark
Marwah Alian, Arafat Awajan, Ahmad Al-Hasan, and Raeda Akuzhia. 2019 · 2019
Cited alongside, same era.
PhoMT: A high-quality and large-scale benchmark dataset for Vietnamese-English machine translation
Long Doan, Linh The Nguyen, Nguyen Luong Tran, Thai Hoang, and Dat Quoc Nguyen. 2021 · 2021
Later among the works it cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Later among the works it cites.
Xl-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ali Fadel, Ibraheem Tuffaha, Bara’ Al-Jawarneh, and Mahmoud Al-Ayyoub. 2019 · 2019
Cited alongside, same era.
Arabert: Transformer-based model for arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj. 2020 · 2020
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Cited alongside, same era.
Benchmarking multidomain English-Indonesian machine translation
Tri Wahyu Guntara, Alham Fikri Aji, and Radityo Eko Prasojo. 2020 · 2020
Cited alongside, same era.
EXAMS: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering
Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020 · 2020
Cited alongside, same era.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
GLGE: A new general language generation evaluation benchmark
Dayiheng Liu, Yu Yan, Yeyun Gong, Weizhen Qi, Hang Zhang, Jian Jiao, Weizhu Chen, Jie Fu, Linjun Shou, Ming Gong, Pengcheng Wang, Jiusheng Chen, Daxin Jiang, Jiancheng Lv, Ruofei Zhang, Winnie Wu, Ming Zhou, and Nan Duan. 2021 · 2021
Later among the works it cites.
Empathetic BERT2BERT conversational model: Learning Arabic language generation with little data
Tarek Naous, Wissam Antoun, Reem Mahmoud, and Hazem Hajj. 2021 · 2021
Later among the works it cites.
Moroccan dialect -darija- open dataset
Aissam Outchakoucht and Hamza Es-Samaali. 2021 · 2021
Later among the works it cites.
Atar: Attention-based lstm for arabizi transliteration
Bashar Talafha, Analle Abuammar, and Mahmoud Al-Ayyoub. 2021 · 2021
Later among the works it cites.
MassiveSumm: a very large-scale, very multilingual, news summarisation dataset
Daniel Varab and Natalie Schluter. 2021 · 2021
Later among the works it cites.
CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark
Yuan Yao, Qingxiu Dong, Jian Guan, Boxi Cao, Zhengyan Zhang, Chaojun Xiao, Xiaozhi Wang, Fanchao Qi, Junwei Bao, Jinran Nie, et al. 2021 · 2021
Later among the works it cites.
User-Centric Gender Rewriting
Bashar Alhafni, Nizar Habash, and Houda Bouamor. 2022 · 2022
Later among the works it cites.
MTG: A benchmark suite for multilingual text generation
Yiran Chen, Zhenqiao Song, Xianze Wu, Danqing Wang, Jingjing Xu, Jiaze Chen, Hao Zhou, and Lei Li. 2022 · 2022
Later among the works it cites.
Clse: Corpus of linguistically significant entities
Aleksandr Chuklin, Justin Zhao, and Mihir Kale. 2022 · 2022
Later among the works it cites.
Arabart: a pretrained arabic sequence-to-sequence model for abstractive summarization
Moussa Kamal Eddine, Nadi Tomeh, Nizar Habash, Joseph Le Roux, and Michalis Vazirgiannis. 2022 · 2022
Later among the works it cites.
Automatic Text Summarization for Moroccan Arabic Dialect Using an Artificial Intelligence Approach , pages 158–177
Kamel Gaanoun, Abdou Naira, Anass Allak, and Imade Benelallam. 2022 · 2022
Later among the works it cites.
GEMv2: Multilingual NLG benchmarking in a single line of code
Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina Mcmillan-major, Anna Shvets, Ashish Upadhyay, and Bernd Bohnet. 2022 · 2022
Later among the works it cites.
LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation
Jian Guan, Zhuoer Feng, Yamei Chen, Ruilin He, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2022 · 2022
Later among the works it cites.
ZAEBUC: An annotated Arabic-English bilingual writer corpus
Nizar Habash and David Palfreyman. 2022 · 2022
Later among the works it cites.
IndicNLG benchmark: Multilingual datasets for diverse NLG tasks in Indic languages
Aman Kumar, Himani Shrotriya, Prachi Sahu, Amogh Mishra, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M. Khapra, and Pratyush Kumar. 2022 · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2022 · 2022
Later among the works it cites.
BanglaNLG and BanglaT5: Benchmarks and resources for evaluating low-resource natural language generation in Bangla
Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, and Rifat Shahriyar. 2023 · 2023
Closest in time.
ORCA: A Challenging Benchmark for Arabic Language Understanding
AbdelRahim Elmadany, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023 · 2023
Closest in time.
Open-domain response generation in low-resource settings using self-supervised pre-training of warm-started transformers
Tarek Naous, Zahraa Bassyouni, Bassel Mousi, Hazem Hajj, Wassim El Hajj, and Khaled Shaban. 2023 · 2023
Closest in time.