Fetching the paper…
Reading the bibliography…
Improving multilingual language models capabilities in low-resource languages is generally difficult due to the scarcity of large-scale data in those languages.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Opinion observer: analyzing and comparing opinions on the web
Bing Liu, Minqing Hu, and Junsheng Cheng. 2005 · 2005
Earlier work this paper cites.
The spanish adaptation of anew (affective norms for english words)
Jaime Redondo, Isabel Fraga, Isabel Padrón, and Montserrat Comesaña. 2007 · 2007
Earlier work this paper cites.
Co-training for cross-lingual sentiment classification
Xiaojun Wan. 2009 · 2009
Earlier work this paper cites.
SentiWordNet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining
Stefano Baccianella, Andrea Esuli, and Fabrizio Sebastiani. 2010 · 2010
Earlier work this paper cites.
PanLex and LEXTRACT: Translating all words of all languages of the world
Timothy Baldwin, Jonathan Pool, and Susan Colowick. 2010 · 2010
Earlier work this paper cites.
A new anew: Evaluation of a word list for sentiment analysis in microblogs
Finn Årup Nielsen. 2011 · 2011
Earlier work this paper cites.
Cross-lingual mixture model for sentiment classification
Xinfan Meng, Furu Wei, Xiaohua Liu, Ming Zhou, Ge Xu, and Houfeng Wang. 2012 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Norms of valence, arousal, dominance, and age of acquisition for 4,300 dutch words
Agnes Moors, Jan De Houwer, Dirk Hermans, Sabine Wanmaker, Kevin Van Schie, Anne-Laura Van Harmelen, Maarten De Schryver, Jeffrey De Winne, and Marc Brysbaert. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Sentiment analysis of short informal texts
Svetlana Kiritchenko, Xiaodan Zhu, and Saif M Mohammad. 2014 · 2014
Earlier work this paper cites.
Cross lingual sentiment analysis using modified BRAE
Sarthak Jain and Shashank Batra. 2015 · 2015
Earlier work this paper cites.
A comparative study on Twitter sentiment analysis: Which features are good?
Fajri Koto and Mirna Adriani. 2015 · 2015
Earlier work this paper cites.
Aspect-level cross-lingual sentiment classification with constrained SMT
Patrik Lambert. 2015 · 2015
Earlier work this paper cites.
SentiRuEval: testing object-oriented sentiment analysis systems in Russian
Natalia Loukachevitch, Pavel Blinov, Evgeny Kotelnikov, Yulia Rubtsova, Vladimir Ivanov, and Elena Tutubalina. 2015 · 2015
Earlier work this paper cites.
PerSent: A freely available Persian sentiment lexicon
Kia Dashtipour, Amir Hussain, Qiang Zhou, Alexander Gelbukh, Ahmad YA Hawalah, and Erik Cambria. 2016 · 2016
Earlier work this paper cites.
SemEval-2016 task 7: Determining sentiment intensity of English and Arabic phrases
Svetlana Kiritchenko, Saif Mohammad, and Mohammad Salameh. 2016 · 2016
Earlier work this paper cites.
EN-ES-CS: An English-Spanish code-switching Twitter corpus for multilingual sentiment analysis
David Vilares, Miguel A. Alonso, and Carlos Gómez-Rodríguez. 2016 · 2016
Earlier work this paper cites.
Attention-based LSTM network for cross-lingual sentiment classification
Xinjie Zhou, Xiaojun Wan, and Jianguo Xiao. 2016a · 2016
Earlier work this paper cites.
Cross-lingual sentiment analysis without (good) translation
Mohamed Abdalla and Graeme Hirst. 2017 · 2017
Earlier work this paper cites.
EmoNet: Fine-grained emotion detection with gated recurrent neural networks
Muhammad Abdul-Mageed and Lyle Ungar. 2017 · 2017
Earlier work this paper cites.
Inset lexicon: Evaluation of a word list for Indonesian sentiment analysis in microblogs
Fajri Koto and Gemala Y Rahmaningtyas. 2017 · 2017
Cited alongside, same era.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand. 2017 · 2017
Cited alongside, same era.
Large scale crowdsourcing and characterization of Twitter abusive behavior
Antigoni Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Cited alongside, same era.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Cited alongside, same era.
Obtaining reliable human ratings of valence, arousal, and dominance for 20,000 English words
Saif Mohammad. 2018 · 2018
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Creating and evaluating resources for sentiment analysis in the low-resource language: Sindhi
Wazir Ali, Naveed Ali, Yong Dai, Jay Kumar, Saifullah Tumrani, and Zenglin Xu. 2021 · 2021
Later among the works it cites.
Introducing a large Tunisian Arabizi dialectal dataset for sentiment analysis
Chayma Fourati, Hatem Haddad, Abir Messaoudi, Moez BenHajhmida, Aymen Ben Elhaj Mabrouk, and Malek Naski. 2021 · 2021
Later among the works it cites.
Task-specific pre-training and cross lingual transfer for sentiment analysis in Dravidian code-switched languages
Akshat Gupta, Sai Krishna Rallabandi, and Alan W Black. 2021 · 2021
Later among the works it cites.
IndoBERTweet: A pretrained language model for Indonesian Twitter with effective domain-specific vocabulary initialization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018 · 2018
Cited alongside, same era.
An investigation of transfer learning-based sentiment analysis in Japanese
Enkhbold Bataa and Joshua Wu. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Parsing with multilingual BERT, a small corpus, and a small treebank
Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Will-they-won’t-they: A very large dataset for stance detection on Twitter
Costanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, and Nigel Collier. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
GoEmotions: A dataset of fine-grained emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020 · 2020
Cited alongside, same era.
Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021 · 2021
Later among the works it cites.
P-stance: A large dataset for stance detection in political domain
Yingjie Li, Tiberiu Sosea, Aditya Sawant, Ajith Jayaraman Nair, Diana Inkpen, and Cornelia Caragea. 2021 · 2021
Later among the works it cites.
Few-shot learning with multilingual language models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2021 · 2021
Later among the works it cites.
HateCheck: Functional tests for hate speech detection models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021 · 2021
Later among the works it cites.
Cross-cultural similarity features for cross-lingual transfer learning of pragmatically motivated tasks
Jimin Sun, Hwijeen Ahn, Chan Young Park, Yulia Tsvetkov, and David R. Mortensen. 2021 · 2021
Later among the works it cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
MetaXL: Meta representation transformation for low-resource cross-lingual learning
Mengzhou Xia, Guoqing Zheng, Subhabrata Mukherjee, Milad Shokouhi, Graham Neubig, and Ahmed Hassan Awadallah. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
Cross-lingual aspect-based sentiment analysis with aspect term code-switching
Wenxuan Zhang, Ruidan He, Haiyun Peng, Lidong Bing, and Wai Lam. 2021 · 2021
Later among the works it cites.
Mawqif: A multi-label Arabic dataset for target-specific stance detection
Nora Saleh Alturayeif, Hamzah Abdullah Luqman, and Moataz Aly Kamaleldin Ahmed. 2022 · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022 · 2022
Later among the works it cites.
NaijaSenti: A Nigerian Twitter sentiment corpus for multilingual sentiment analysis
Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Sa’id Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Alípio Jorge, and Pavel Brazdil. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Expanding pretrained models to thousands more languages via lexicon-based adaptation
Xinyi Wang, Sebastian Ruder, and Graham Neubig. 2022 · 2022
Later among the works it cites.
SemEval-2023 Task 12: Sentiment Analysis for African Languages (AfriSenti-SemEval)
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Seid Muhie Yimam, David Ifeoluwa Adelani, Ibrahim Sa’id Ahmad, Nedjma Ousidhoum, Abinew Ali Ayele, Saif M. Mohammad, Meriem Beloucif, and Sebastian Ruder. 2023 · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
NusaX: Multilingual parallel sentiment dataset for 10 Indonesian local languages
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, and Sebastian Ruder. 2023 · 2023
Later among the works it cites.