Fetching the paper…
Reading the bibliography…
Large pretrained language models (PLMs) typically tokenize the input string into contiguous subwords before any pretraining or inference.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Anersys: An arabic named entity recognition system based on maximum entropy
Yassine Benajiba, Paolo Rosso, and José Miguel Benedíruiz. 2007 · 2007
Earlier work this paper cites.
Overview of the SPMRL 2013 shared task: A cross-framework evaluation of parsing morphologically rich languages
Djamé Seddah, Reut Tsarfaty, Sandra Kübler, Marie Candito, Jinho D. Choi, Richárd Farkas, Jennifer Foster, Iakes Goenaga, Koldo Gojenola Galletebeitia, Yoav Goldberg, Spence Green, Nizar Habash, Marco Kuhlmann, Wolfgang Maier, Joakim Nivre, Adam Przepiórkowski, Ryan Roth, Wolfgang Seeker, Yannick Versley, Veronika Vincze, Marcin Woliński, Alina Wróblewska, and Eric Villemonte de la Clergerie. 2013 · 2013
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
H Bahadir Sahin, Caglar Tirkaz, Eray Yildiz, Mustafa Tolga Eren, and Ozan Sonmez. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
CoNLL-UL: Universal morphological lattices for Universal Dependency parsing
Amir More, Özlem Çetinoğlu, Çağrı Çöltekin, Nizar Habash, Benoît Sagot, Djamé Seddah, Dima Taji, and Reut Tsarfaty. 2018 · 2018
Earlier work this paper cites.
The Hebrew Universal Dependency treebank: Past present and future
Shoval Sade, Amit Seker, and Reut Tsarfaty. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Massively multilingual transfer for NER
Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019 · 2019
Cited alongside, same era.
AraBERT: Transformer-based model for Arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj. 2020 · 2020
Cited alongside, same era.
Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020 · 2020
Cited alongside, same era.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Later among the works it cites.
Berturk - bert models for turkish
Stefan Schweter. 2020 · 2020
Later among the works it cites.
A pointer network architecture for joint morphological segmentation and tagging
Amit Seker and Reut Tsarfaty. 2020 · 2020
Later among the works it cites.
From SPMRL to NMRL: What did we learn (and unlearn) in a decade of parsing morphologically-rich languages (MRLs)?
Reut Tsarfaty, Dan Bareket, Stav Klein, and Amit Seker. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
Getting the ##life out of living: How adequate are word-pieces for modelling complex morphology?
Stav Klein and Reut Tsarfaty. 2020 · 2020
Cited alongside, same era.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Cited alongside, same era.
CAMeL tools: An open source python toolkit for Arabic natural language processing
Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, and Nizar Habash. 2020 · 2020
Cited alongside, same era.
Neural Modeling for Named Entities and Morphology (NEMO2)
Dan Bareket and Reut Tsarfaty. 2021 · 2021
Later among the works it cites.
Hebert & hebemo: a hebrew bert model and a tool for polarity analysis and emotion recognition
Avihay Chriqui and Inbal Yahav. 2021 · 2021
Later among the works it cites.
ParaShoot: A Hebrew question answering dataset
Omri Keren and Omer Levy. 2021 · 2021
Later among the works it cites.
Alephbert: A hebrew large pre-trained language model to start-off your hebrew nlp application with
Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky, Refael Shaked Greenfeld, and Reut Tsarfaty. 2021 · 2021
Later among the works it cites.