Fetching the paper…
Reading the bibliography…
While code-mixing is a common linguistic practice in many parts of the world, collecting high-quality and low-cost code-mixed data remains a challenge for natural language processing (NLP) research.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Commonwealth act no. 184: Govph
Commonwealth of the Philippines. 1936 · 1936
Earlier work this paper cites.
Undang-Undang Dasar Negara Republik Indonesia Tahun 1945
Republik Indonesia. 2002 · 1945
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J Richard Landis and Gary G Koch. 1977 · 1977
Earlier work this paper cites.
Syntactic structure and social function of code-switching , volume 2
Shana Poplack. 1978 · 1978
Earlier work this paper cites.
Life with two languages: An introduction to bilingualism
F. Grosjean. 1982 · 1982
Earlier work this paper cites.
Linguistic constraints on intrasentential code-switching: A study of spanish/hebrew bilingualism
Susan Berk-Seligson. 1986 · 1986
Earlier work this paper cites.
Practical statistics for medical research
Douglas G Altman. 1990 · 1990
Earlier work this paper cites.
The case of the nonce loan in tamil
David Sankoff, Shana Poplack, and Swathi Vanniarajan. 1990 · 1990
Earlier work this paper cites.
Phrase structure and grammatical relations in Tagalog
Paul Kroeger. 1993 · 1993
Earlier work this paper cites.
Code-switching as a verbal strategy among chinese in a campus setting in taiwan
Su-Chiao Chen. 1996 · 1996
Earlier work this paper cites.
Mainland southeast asia: A unique linguistic area
Brian Migliazza. 1996 · 1996
Earlier work this paper cites.
Duelling languages: Grammatical structure in codeswitching
Carol Myers-Scotton. 1997 · 1997
Earlier work this paper cites.
Malaysian tamils and tamil linguistic culture
Harold F. Schiffman. 1998 · 1998
Earlier work this paper cites.
Spanish-english code-switching among us latinos
A. J. Toribio. 2002 · 2002
Earlier work this paper cites.
The Indonesian language: Its history and role in modern society
James Neil Sneddon. 2003 · 2003
Earlier work this paper cites.
Social and psychological factors in language mixing
T. K. Bhatia and W. C. Ritchie. 2004 · 2004
Earlier work this paper cites.
Linguistic profiling
John Baugh. 2005 · 2005
Earlier work this paper cites.
The languages of East and Southeast Asia: an introduction
Cliff Goddard. 2005 · 2005
Earlier work this paper cites.
Southeast asian englishes
Maria Lourdes S Bautista and Andrew B Gonzalez. 2006 · 2006
Earlier work this paper cites.
A two-level morphological analyser for the Indonesian language
Femphy Pisceldo, Rahmad Mahendra, Ruli Manurung, and I Wayan Arka. 2008 · 2008
Earlier work this paper cites.
Learning to predict code-switching points
Thamar Solorio and Yang Liu. 2008 · 2008
Earlier work this paper cites.
Automatic recognition of cantonese-english code-mixing speech
Joyce YC Chan, Houwei Cao, PC Ching, and Tan Lee. 2009 · 2009
Earlier work this paper cites.
Mother tongue as bridge language of instruction: Policies and experiences in southeast Asia
Mont Redmond, Kimmo Kosonen, and Catherine Young. 2009 · 2009
Earlier work this paper cites.
Seame: a mandarin-english code-switching speech corpus in south-east asia
Dau-Cheng Lyu, Tien-Ping Tan, Eng Siong Chng, and Haizhou Li. 2010 · 2010
Earlier work this paper cites.
Antipassive and ergativity in tagalog
Edith Aldridge. 2012 · 2012
Earlier work this paper cites.
Pattern matching refinements to dictionary-based code-switching point detection
Nathaniel Oco and Rachel Edita Roxas. 2012 · 2012
Earlier work this paper cites.
Proceedings of the First Workshop on Computational Approaches to Code Switching . Association for Computational Linguistics, Doha, Qatar
Mona Diab, Julia Hirschberg, Pascale Fung, and Thamar Solorio, editors. 2014 · 2014
Earlier work this paper cites.
On measuring the complexity of code-mixing
Björn Gambäck and Amitava Das. 2014 · 2014
Cited alongside, same era.
English in southeast asia: Pedagogical and policy implications
Andy Kirkpatrick. 2014 · 2014
Cited alongside, same era.
The national university of singapore sms corpus
T Chen and Kan Min-Yen. 2015 · 2015
Cited alongside, same era.
Don’t play, play - singlish is studied around the globe
Yuen Sin. 2017 · 2017
Cited alongside, same era.
Language modeling for code-mixing: The role of linguistic theory based synthetic data
Adithya Pratapa, Gayatri Bhat, Monojit Choudhury, Sunayana Sitaram, Sandipan Dandapat, and Kalika Bali. 2018 · 2018
Cited alongside, same era.
Normalization of Indonesian-English code-mixed Twitter data
Anab Maulana Barik, Rahmad Mahendra, and Mirna Adriani. 2019 · 2019
Cited alongside, same era.
IndoRobusta: Towards robustness against diverse code-mixed Indonesian local languages
Muhammad Farid Adilazuarda, Samuel Cahyawijaya, Genta Indra Winata, Pascale Fung, and Ayu Purwarianti. 2022 · 2022
Later among the works it cites.
One country, 700+ languages: NLP challenges for underrepresented languages and dialects in Indonesia
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, and Sebastian Ruder. 2022 · 2022
Later among the works it cites.
Nusacrowd: Open source initiative for indonesian nlp resources
Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Indra Winata, Bryan Wilie, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, Fajri Koto, Jennifer Santoso, David Moeljadi, Cahya Wirawan, Frederikus Hudi, Ivan Halim Parmonangan, Ika Alfina, Muhammad Satrio Wicaksono, Ilham Firdausi Putra, Samsul Rahmadani, Yulianti Oenang, Ali Akbar Septiandri, James Jaya, Kaustubh D. Dhole, Arie Ardiyanti Suryani, Rifki Afina Putri, Dan Su, Keith Stevens, Made Nindyatama Nityasya, Muhammad Farid Adilazuarda, Ryan Ignatius, Ryandito Diandaru, Tiezheng Yu, Vito Ghifari, Wenliang Dai, Yan Xu, Dyah Damapuspita, Cuk Tho, Ichwanul Muslim Karo Karo, Tirana Noor Fatyanosa, Ziwei Ji, Pascale Fung, Graham Neubig, Timothy Baldwin, Sebastian Ruder, Herry Sujaini, Sakriani Sakti, and Ayu Purwarianti. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Code-switched language models using neural based synthetic data from parallel sentences
Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, and Pascale Fung. 2019 · 2019
Cited alongside, same era.
LinCE: A centralized benchmark for linguistic code-switching evaluation
Gustavo Aguilar, Sudipta Kar, and Thamar Solorio. 2020 · 2020
Cited alongside, same era.
Do multilingual users prefer chat-bots that code-mix? let’s nudge and find out!
Anshul Bawa, Pranav Khadpe, Pratik Joshi, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
GLUECoS: An evaluation benchmark for code-switched NLP
Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Attention-informed mixed-language training for zero-shot cross-lingual task-oriented dialogue systems
Zihan Liu, Genta Indra Winata, Zhaojiang Lin, Peng Xu, and Pascale Fung. 2020 · 2020
Cited alongside, same era.
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Later among the works it cites.
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology
Mark Dingemanse and Andreas Liesenfeld. 2022 · 2022
Later among the works it cites.
Multi-figurative language generation
Huiyuan Lai and Malvina Nissim. 2022 · 2022
Later among the works it cites.
The bigscience ROOTS corpus: A 1.6TB composite multilingual dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Šaško, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel, Leon Weber, Manuel Romero Muñoz, Jian Zhu, Daniel Van Strien, Zaid Alyafeai, Khalid Almubarak, Vu Minh Chien, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Luccioni, and Yacine Jernite. 2022 · 2022
Later among the works it cites.
ASCEND: A spontaneous Chinese-English dataset for code-switching in multi-turn conversation
Holy Lovenia, Samuel Cahyawijaya, Genta Winata, Peng Xu, Yan Xu, Zihan Liu, Rita Frieske, Tiezheng Yu, Wenliang Dai, Elham J. Barezi, Qifeng Chen, Xiaojuan Ma, Bertram Shi, and Pascale Fung. 2022 · 2022
Later among the works it cites.
CoCoa: An encoder-decoder model for controllable code-switched generation
Sneha Mondal, Ritika ., Shreya Pathak, Preethi Jyothi, and Aravindan Raghuveer. 2022 · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
M2D2: A massively multi-domain language modeling dataset
Machel Reid, Victor Zhong, Suchin Gururangan, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Later among the works it cites.
Data feedback loops: Model-driven amplification of dataset biases
Rohan Taori and Tatsunori B Hashimoto. 2022 · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022 · 2022
Later among the works it cites.
The decades progress on code-switching research in nlp: A systematic survey on trends and challenges
Genta Indra Winata, Alham Fikri Aji, Zheng-Xin Yong, and Thamar Solorio. 2022 · 2022
Later among the works it cites.
Crocosum: A benchmark dataset for cross-lingual code-switched summarization
Ruochen Zhang and Carsten Eickhoff. 2023 · 2022
Later among the works it cites.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023 · 2023
Closest in time.
Annollm: Making large language models to be better crowdsourced annotators
Xingwei He, Zhenghao Lin, Yeyun Gong, A Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. 2023 · 2023
Closest in time.
Multi-lingual and multi-cultural figurative language understanding
Anubha Kabra, Emmy Liu, Simran Khanuja, Alham Fikri Aji, Genta Indra Winata, Samuel Cahyawijaya, Anuoluwapo Aremu, Perez Ogayo, and Graham Neubig. 2023 · 2023
Closest in time.
Scaling speech technology to 1,000+ languages
Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, et al. 2023 · 2023
Closest in time.
BLOOMChat: a New Open Multilingual Chat LLM
Together Computer SambaNova Systems. 2023 · 2023
Closest in time.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. 2023 · 2023
Closest in time.
Does synthetic data generation of llms help clinical text mining?
Ruixiang Tang, Xiaotian Han, Xiaoqian Jiang, and Xia Hu. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Llm-powered data augmentation for enhanced crosslingual performance
Chenxi Whitehouse, Monojit Choudhury, and Alham Fikri Aji. 2023 · 2023
Closest in time.
NusaX: Multilingual parallel sentiment dataset for 10 Indonesian local languages
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, and Sebastian Ruder. 2023 · 2023
Closest in time.
Multilingual large language models are not (yet) code-switchers
Ruochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, and Alham Fikri Aji. 2023 · 2023
Closest in time.