Fetching the paper…
Reading the bibliography…
Deep learning-based language models pretrained on large unannotated text corpora have been demonstrated to allow efficient transfer learning for natural language processing, with recent approaches such as the transformer-based BERT model advancing the state of the art across a variety of tasks.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
How multilingual is multilingual bert?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 1906
Earlier work this paper cites.
The stanford typed dependencies representation
Marie-Catherine De Marneffe and Christopher D Manning. 2008 · 2008
Earlier work this paper cites.
Efficient web crawling for large text corpora
Vít Suchomel, Jan Pomikálek, et al. 2012 · 2012
Earlier work this paper cites.
Specifying treebanks, outsourcing parsebanks: Finntreebank 3
Atro Voutilainen, Kristiina Muhonen, Tanja Katariina Purtonen, Krister Lindén, et al. 2012 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Building the essential resources for Finnish: the Turku Dependency Treebank
Katri Haverinen, Jenna Nyblom, Timo Viljanen, Veronika Laippala, Samuel Kohonen, Anna Missilä, Stina Ojala, Tapio Salakoski, and Filip Ginter. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Towards universal web parsebanks
Juhani Luotolahti, Jenna Kanerva, Veronika Laippala, Sampo Pyysalo, and Filip Ginter. 2015 · 2015
Earlier work this paper cites.
Universal Dependencies for Finnish
Sampo Pyysalo, Jenna Kanerva, Anna Missilä, Veronika Laippala, and Filip Ginter. 2015 · 2015
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016 · 2016
Earlier work this paper cites.
Simple and accurate dependency parsing using bidirectional LSTM feature representations
Eliyahu Kiperwasser and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Universal Dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine De Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic, Christopher D Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, et al. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Stanford’s graph-based neural dependency parser at the CoNLL 2017 Shared Task
Timothy Dozat, Peng Qi, and Christopher D Manning. 2017 · 2017
Earlier work this paper cites.
CoNLL 2017 shared task - automatically annotated raw texts and word embeddings
Filip Ginter, Jan Hajič, Juhani Luotolahti, Milan Straka, and Daniel Zeman. 2017 · 2017
Cited alongside, same era.
Tagging named entities in 19th century and modern Finnish newspaper material with a Finnish semantic tagger
Kimmo Kettunen and Laura Löfberg. 2017 · 2017
Cited alongside, same era.
Tokenizing, pos tagging, lemmatizing and parsing ud 2.0 with udpipe
Milan Straka and Jana Straková. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
CoNLL 2017 shared task
Daniel Zeman, Martin Popel, Milan Straka, Jan Hajič, Joakim Nivre, Filip Ginter, Juhani Luotolahti, Sampo Pyysalo, Slav Petrov, Martin Potthast, et al. 2017 · 2017
Cited alongside, same era.
Towards better UD parsing: Deep contextualized word embeddings, ensemble, and treebank concatenation
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Later among the works it cites.
CoNLL 2018 shared task: multilingual parsing from raw text to universal dependencies
Daniel Zeman, Jan Hajič, Martin Popel, Martin Potthast, Milan Straka, Filip Ginter, Joakim Nivre, and Slav Petrov. 2018a · 2018
Later among the works it cites.
CoNLL 2018 shared task system outputs
Daniel Zeman, Martin Potthast, Elie Duthoo, Olivier Mesnard, Piotr Rybak, Alina Wróblewska, Wanxiang Che, Yijia Liu, Yuxuan Wang, Bo Zheng, et al. 2018b · 2018
Later among the works it cites.
Scibert: Pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Closest in time.
75 languages, 1 model: Parsing universal dependencies universally
Dan Kondratyuk and Milan Straka. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wanxiang Che, Yijia Liu, Yuxuan Wang, Bo Zheng, and Ting Liu. 2018 · 2018
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Improving named entity recognition by jointly learning to disambiguate morphological tags
Onur Güngör, Suzan Üsküdarlı, and Tunga Güngör. 2018 · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Turku Neural Parser Pipeline: An end-to-end system for the CoNLL 2018 shared task
Jenna Kanerva, Filip Ginter, Niko Miekka, Akseli Leino, and Tapio Salakoski. 2018 · 2018
Cited alongside, same era.
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Deep contextualized word embeddings in transition-based and graph-based dependency parsing–a tale of two parsers revisited
Artur Kulmizev, Miryam de Lhoneux, Johannes Gontrum, Elena Fano, and Joakim Nivre. 2019 · 2019
Closest in time.
Toward multilingual identification of online registers
Veronika Laippala, Roosa Kyllönen, Jesse Egbert, Douglas Biber, and Sampo Pyysalo. 2019 · 2019
Closest in time.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019 · 2019
Closest in time.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Closest in time.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, and Benoît Sagot. 2019 · 2019
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Closest in time.
Multilingual probing of deep pre-trained contextual encoders
Vinit Ravishankar, Memduh Gökırmak, Lilja Øvrelid, and Erik Velldal. 2019 · 2019
Closest in time.
Is multilingual BERT fluent in language generation?
Samuel Rönnqvist, Jenna Kanerva, Tapio Salakoski, and Filip Ginter. 2019 · 2019
Closest in time.
A Finnish news corpus for named entity recognition
Teemu Ruokolainen, Pekka Kauppinen, Miikka Silfverberg, and Krister Lindén. 2019 · 2019
Closest in time.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Closest in time.
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2019 · 2019
Closest in time.