Fetching the paper…
Reading the bibliography…
Cross-lingual summarization (CLS) has attracted increasing interest in recent years due to the availability of large-scale web-mined datasets and the advancements of multilingual language models.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Mlsum: The multilingual summarization corpus
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2020 · 2004
Earlier work this paper cites.
Part-of-Speech tagging for English-Spanish code-switched text
Thamar Solorio and Yang Liu. 2008 · 2008
Earlier work this paper cites.
Multilingual translation with extensible multilingual pretraining and finetuning
Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020 · 2008
Earlier work this paper cites.
Cross-language document summarization based on machine translation quality prediction
Xiaojun Wan, Huiying Li, and Jianguo Xiao. 2010 · 2010
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020 · 2010
Earlier work this paper cites.
Seame: a mandarin-english code-switching speech corpus in south-east asia
Lyu, Dau-Cheng and Tan, Tien-Ping and Chng, Eng Siong and Li, Haizhou. 2010 · 2010
Earlier work this paper cites.
A mandarin-english code-switching corpus
Li, Ying and Yu, Yue and Fung, Pascale. 2012 · 2012
Earlier work this paper cites.
Code mixing: A challenge for language identification in the language of social media
Utsab Barman, Amitava Das, Joachim Wagner, and Jennifer Foster. 2014 · 2014
Earlier work this paper cites.
Identifying languages at the word level in code-mixed Indian social media text
Amitava Das and Björn Gambäck. 2014 · 2014
Earlier work this paper cites.
On measuring the complexity of code-mixing
Björn Gambäck and Amitava Das. 2014 · 2014
Earlier work this paper cites.
Overview for the first shared task on language identification in code-switched data
Thamar Solorio, Elizabeth Blair, Suraj Maharjan, Steven Bethard, Mona Diab, Mahmoud Ghoneim, Abdelati Hawwari, Fahad AlGhamdi, Julia Hirschberg, Alison Chang, and Pascale Fung. 2014 · 2014
Earlier work this paper cites.
Comparing the level of code-switching in corpora
Björn Gambäck and Amitava Das. 2016 · 2016
Earlier work this paper cites.
Abstractive cross-language summarization via translation model enhanced predicate argument structure fusing
Jiajun Zhang, Yu Zhou, and Chengqing Zong. 2016 · 2016
Earlier work this paper cites.
Zero-shot cross-lingual neural headline generation
Ayana, Shi-qi Shen, Yun Chen, Cheng Yang, Zhi-yuan Liu, and Mao-song Sun. 2018 · 2018
Earlier work this paper cites.
Named entity recognition for Hindi-English code-mixed social media text
Vinay Singh, Deepanshu Vijay, Syed Sarfaraz Akhtar, and Manish Shrivastava. 2018 · 2018
Cited alongside, same era.
Joint part-of-speech and language ID tagging for code-switched data
Victor Soto and Julia Hirschberg. 2018 · 2018
Cited alongside, same era.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Grusky, Max and Naaman, Mor and Artzi, Yoav. 2018 · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Cited alongside, same era.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Gliwa, Bogdan and Mochol, Iwona and Biesek, Maciej and Wawer, Aleksander. 2019 · 2019
Cited alongside, same era.
From machine translation to code-switching: Generating high-quality code-switched text
Ishan Tarunesh, Syamantak Kumar, and Preethi Jyothi. 2021 · 2021
Later among the works it cites.
Are multilingual models effective in code-switching?
Genta Indra Winata, Samuel Cahyawijaya, Zihan Liu, Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2021 · 2021
Later among the works it cites.
Tahmid Hasan and Abhik Bhattacharjee and Wasi Uddin Ahmad and Yuan-Fang Li and Yong-bin Kang and Rifat Shahriyar. 2021 · 2021
Later among the works it cites.
GupShup: Summarizing open-domain code-switched conversations
Mehnaz, Laiba and Mahata, Debanjan and Gosangi, Rakesh and Gunturi, Uma Sushmitha and Jain, Riya and Gupta, Gauri and Kumar, Amardeep and Lee, Isabelle G and Acharya, Anish and Shah, Rajiv. 2021 · 2021
Later among the works it cites.
Models and Datasets for Cross-Lingual Summarisation
Laura Perez-Beltrachini and Mirella Lapata. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Global Voices: Crossing Borders in Automatic News Summarization
Nguyen, Khanh and Daumé III, Hal. 2019 · 2019
Cited alongside, same era.
NCLS: Neural Cross-Lingual Summarization
Zhu, Junnan and Wang, Qian and Wang, Yining and Zhou, Yu and Zhang, Jiajun and Wang, Shaonan and Zong, Chengqing. 2019 · 2019
Cited alongside, same era.
Statistical power and translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2020 · 2020
Cited alongside, same era.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation
Aguilar, Gustavo and Kar, Sudipta and Solorio, Thamar. 2020 · 2020
Cited alongside, same era.
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022 · 2022
Later among the works it cites.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022 · 2022
Later among the works it cites.
Cocoa: An encoder-decoder model for controllable code-switched generation
Sneha Mondal, Shreya Pathak, Preethi Jyothi, and Aravindan Raghuveer. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Later among the works it cites.
The decades progress on code-switching research in nlp: A systematic survey on trends and challenges
Genta Indra Winata, Alham Fikri Aji, Zheng-Xin Yong, and Thamar Solorio. 2022 · 2022
Later among the works it cites.
TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a Speech Recognition Baseline
Li, Chengfei and Deng, Shuhao and Wang, Yaoping and Wang, Guangjing and Gong, Yaguang and Chen, Changbin and Bai, Jinfeng. 2022 · 2022
Later among the works it cites.
ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation
Lovenia, Holy and Cahyawijaya, Samuel and Winata, Genta and Xu, Peng and Xu, Yan and Liu, Zihan and Frieske, Rita and Yu, Tiezheng and Dai, Wenliang and Barezi, Elham J. and Chen, Qifeng and Ma, Xiaojuan and Shi, Bertram and Fung, Pascale. 2022 · 2022
Later among the works it cites.
Clidsum: A benchmark dataset for cross-lingual dialogue summarization
Wang, Jiaan and Meng, Fandong and Lu, Ziyao and Zheng, Duo and Li, Zhixu and Qu, Jianfeng and Zhou, Jie. 2022 · 2022
Later among the works it cites.
Long-Document Cross-Lingual Summarization
Zheng, Shaohui and Li, Zhixu and Wang, Jiaan and Qu, Jianfeng and Liu, An and Zhao, Lei and Chen, Zhigang. 2022 · 2022
Later among the works it cites.
A survey of code-switching: Linguistic and social perspectives for language technologies
A Seza Doğruöz, Sunayana Sitaram, Barbara E Bullock, and Almeida Jacqueline Toribio. 2023 · 2023
Closest in time.
Prompting multilingual large language models to generate code-mixed texts: The case of south east asian languages
Zheng-Xin Yong, Ruochen Zhang, Jessica Zosa Forde, Skyler Wang, Samuel Cahyawijaya, Holy Lovenia, Genta Indra Winata, Lintang Sutawika, Jan Christian Blaise Cruz, Long Phan, et al. 2023 · 2023
Closest in time.