Fetching the paper…
Reading the bibliography…
The first step in any NLP pipeline is to split the text into individual tokens.
A comparative study between methods of arabic baseline detection
AL-Shatnawi Atallah and Khairuddin Omar · 2009
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
Labr: A large scale arabic book reviews dataset
Mohamed Aly and Amir Atiya · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Madamira: A fast, comprehensive tool for morphological analysis and disambiguation of arabic
Arfath Pasha, Mohamed Al-Badrashiny, Mona T Diab, Ahmed El Kholy, Ramy Eskander, Nizar Habash, Manoj Pooleery, Owen Rambow, and Ryan Roth · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Morfessor 2.0: Toolkit for statistical morphological segmentation
Peter Smit, Sami Virpioja, Stig-Arne Grönroos, and Mikko Kurimo · 2014
Earlier work this paper cites.
Variable-length word encodings for neural translation models
Rohan Chitnis and John DeNero · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Farasa: A fast and furious segmenter for arabic
Ahmed Abdelali, Kareem Darwish, Nadir Durrani, and Hamdy Mubarak · 2016
Earlier work this paper cites.
1.5 billion words arabic corpus
Ibrahim Abu El-Khair · 2016
Earlier work this paper cites.
Orthographic syllable as basic unit for smt between related languages
Anoop Kunchukuttan and Pushpak Bhattacharyya · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Arabic online handwriting recognition (aohr) a survey
Baligh M Al-Helali and Sabri A Mahmoud · 2017
Earlier work this paper cites.
Arabic tweets sentimental analysis using machine learning
Khaled Mohammad Alomari, Hatem M ElSherif, and Khaled Shaalan · 2017
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Cited alongside, same era.
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, Ahmed Abdelali, Yonatan Belinkov, and Stephan Vogel · 2017
Cited alongside, same era.
Dataset for arabic classification
mohamed BINIZ · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Later among the works it cites.
Waleed A Yousef, Omar M Ibrahime, Taha M Madbouly, and Moustafa A Mahmoud · 2019
Later among the works it cites.
Classifying and diacritizing arabic poems using deep recurrent neural networks
Gheith A Abandah, Mohammed Z Khedher, Mohammad R Abdel-Majeed, Hamdi M Mansour, Salma F Hulliel, and Lara M Bisharat · 2020
Later among the works it cites.
Arbert & marbert: deep bidirectional transformers for arabic
Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Cited alongside, same era.
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
A comprehensive survey of arabic sentiment analysis
Mahmoud Al-Ayyoub, Abed Allah Khamaiseh, Yaser Jararweh, and Mohammed N Al-Kabi · 2019
Cited alongside, same era.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones · 2019
Cited alongside, same era.
A survey of opinion mining in arabic: a comprehensive system perspective covering challenges and advances in tools, resources, models, applications, and visualizations
Gilbert Badaro, Ramy Baly, Hazem Hajj, Wassim El-Hajj, Khaled Bashir Shaban, Nizar Habash, Ahmad Al-Sallab, and Ali Hamdi · 2019
Cited alongside, same era.
hulmona: The universal language model in arabic
Obeida ElJundi, Wissam Antoun, Nour El Droubi, Hazem Hajj, Wassim El-Hajj, and Khaled Shaban · 2019
Cited alongside, same era.
Arabic sentiment analysis: studies, resources, and tools
Imane Guellil, Faical Azouaou, and Marcelo Mendoza · 2019
Cited alongside, same era.
Ibrahim Abu Farha and Walid Magdy · 2020
Later among the works it cites.
Meter classification of arabic poems using deep bidirectional recurrent neural networks
Maged S Al-shaibani, Zaid Alyafeai, and Irfan Ahmad · 2020
Later among the works it cites.
Metrec: A dataset for meter classification of arabic poetry
Maged S Al-Shaibani, Zaid Alyafeai, and Irfan Ahmad · 2020
Later among the works it cites.
On the importance of tokenization in arabic embedding models
Mohamed Alkaoud and Mairaj Syed · 2020
Later among the works it cites.
Arabert: Transformer-based model for arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj · 2020
Later among the works it cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett · 2020
Later among the works it cites.
Arabic optical characters recognition by neural network based arabic unicode
Mahdi Nsaif Jasim · 2020
Later among the works it cites.
Pre-training bert on arabic tweets: Practical considerations
Ahmed Abdelali, Sabit Hassan, Hamdy Mubarak, Kareem Darwish, and Younes Samih · 2021
Closest in time.
Charformer: Fast character transformers via gradient-based subword tokenization
Yi Tay, Vinh Q Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler · 2021
Closest in time.
Multi-view subword regularization
Xinyi Wang, Sebastian Ruder, and Graham Neubig · 2021
Closest in time.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel · 2021
Closest in time.