Fetching the paper…
Reading the bibliography…
Large transformer-based language models, e.g.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Julian Salazar, Davis Liang, Toan Q Nguyen, and Katrin Kirchhoff. 2019 · 1910
Earlier work this paper cites.
Benjamin van der Burgh and Suzan Verberne. 2019 · 1910
Earlier work this paper cites.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric Villemonte de La Clergerie, Djamé Seddah, and Benoît Sagot. 2019 · 1911
Earlier work this paper cites.
Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. 2019 · 1912
Earlier work this paper cites.
What is concept drift and how to measure it?
Shenghui Wang, Stefan Schlobach, and Michel Klein. 2010 · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
BERT-NL a set of language models pre-trained on the Dutch SoNaR corpus
Alex Brandsen, Anne Dirkson, Suzan Verberne, Maya Sappelli, Dung Manh Chu, and Kimberly Stoutjesdijk. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Cited alongside, same era.
Automatically correcting Dutch pronouns “die” and “dat”
Liesbeth Allein, Artuur Leeuwenberg, and Marie-Francine Moens. 2020 · 2020
Cited alongside, same era.
RobBERT: a Dutch RoBERTa-based Language Model
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2020 · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
A primer in bertology: What we know about how bert works
Lifelong pretraining: Continually adapting language models to emerging corpora
Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, and Xiang Ren. 2021 · 2021
Later among the works it cites.
Temporal adaptation of BERT and performance on downstream document classification: Insights from social media
Paul Röttger and Janet Pierrehumbert. 2021 · 2021
Later among the works it cites.
Measuring shifts in attitudes towards COVID-19 measures in Belgium
Kristen Scott, Pieter Delobelle, and bettina Berendt. 2021 · 2021
Later among the works it cites.
MedRoBERTa.nl: A language model for Dutch electronic health records
Stella Verkijk and Piek Vossen. 2021 · 2021
Later among the works it cites.
Adapt-and-distill: Developing small, fast and effective pretrained language models for domains
Yunzhi Yao, Shaohan Huang, Wenhui Wang, Li Dong, and Furu Wei. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Cited alongside, same era.
Neural machine translation with byte-level subwords
Changhan Wang, Kyunghyun Cho, and Jiatao Gu. 2020 · 2020
Cited alongside, same era.
Forget me not: Reducing catastrophic forgetting for domain adaptation in reading comprehension
Ying Xu, Xu Zhong, Antonio Jose Jimeno Yepes, and Jey Han Lau. 2020 · 2020
Cited alongside, same era.
RobBERTje: A distilled Dutch BERT model
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2021 · 2021
Cited alongside, same era.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021 · 2021
Cited alongside, same era.
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, and Benoît Sagot. 2022 · 2022
Closest in time.
Legal-RobBERT: A Dutch BERT model specific to the legal domain
Thomas Boer, Jesse Tijsterman, and Lan Chu. 2022 · 2022
Closest in time.
Domain- and task-adaptation for VaccinChatNL, a Dutch COVID-19 FAQ answering corpus and classification model
Jeska Buhmann, Maxime De Bruyn, Ehsan Lotfi, and Walter Daelemans. 2022 · 2022
Closest in time.
A continual learning survey: Defying forgetting in classification tasks
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. 2022 · 2022
Closest in time.
CoNTACT: A Dutch COVID-19 adapted BERT for vaccine hesitancy and argumentation detection
Jens Lemmens, Jens Van Nooten, Tim Kreutz, and Walter Daelemans. 2022 · 2022
Closest in time.