Fetching the paper…
Reading the bibliography…
We introduce the Dutch Model Benchmark: DUMB.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
The merits of universal language model fine-tuning for small datasets–a case with dutch book reviews
Benjamin Van der Burgh and Suzan Verberne. 2019 · 1910
Earlier work this paper cites.
Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. 2019 · 1912
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
WordNet: A lexical database for English
George A. Miller. 1994 · 1994
Earlier work this paper cites.
Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition
Erik F. Tjong Kim Sang. 2002 · 2002
Earlier work this paper cites.
Part of speech tagging en lemmatisering van het Corpus Gesproken Nederlands
Frank van Eynde. 2004 · 2004
Earlier work this paper cites.
Integrating lexical units, synsets and ontology in the cornetto database
Piek Vossen, Isa Maks, Roxane Segers, and Hennie VanderVliet. 2008 · 2008
Earlier work this paper cites.
VerbNet overview, extensions, mappings and applications
Karin Kipper Schuler, Anna Korhonen, and Susan Brown. 2009 · 2009
Earlier work this paper cites.
SemEval-2010 task 1: Coreference resolution in multiple languages
Marta Recasens, Lluís Màrquez, Emili Sapena, M. Antònia Martí, Mariona Taulé, Véronique Hoste, Massimo Poesio, and Yannick Versley. 2010 · 2010
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2012 · 2012
Earlier work this paper cites.
The Winograd Schema Challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
DutchSemCor: Targeting the ideal sense-tagged corpus
Piek Vossen, Attila Görög, Rubén Izquierdo, and Antal van den Bosch. 2012 · 2012
Earlier work this paper cites.
The Construction of a 500-Million-Word Reference Corpus of Contemporary Written Dutch
Nelleke Oostdijk, Martin Reynaert, Véronique Hoste, and Ineke Schuurman. 2013 · 2013
Earlier work this paper cites.
Large Scale Syntactic Annotation of Written Dutch: Lassy
Gertjan van Noord, Gosse Bouma, Frank Van Eynde, Daniël de Kok, Jelmer van der Linde, Ineke Schuurman, Erik Tjong Kim Sang, and Vincent Vandeghinste. 2013 · 2013
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Open Dutch WordNet
Marten Postma, Emiel van Miltenburg, Roxane Segers, Anneleen Schoen, and Piek Vossen. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019a · 2019
Cited alongside, same era.
WiC: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019 · 2019
Cited alongside, same era.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Cited alongside, same era.
DeBERTa: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
Cross-lingual transferring of pre-trained contextualized language models
Zuchao Li, Kevin Parnow, Hai Zhao, Zhuosheng Zhang, Rui Wang, Masao Utiyama, and Eiichiro Sumita. 2021 · 2021
Later among the works it cites.
KLUE: Korean language understanding evaluation
Sungjoon Park, Jihyung Moon, Sungdong Kim, Won Ik Cho, Ji Yoon Han, Jangwon Park, Chisung Song, Junseong Kim, Youngsook Song, Taehwan Oh, Joohong Lee, Juhyun Oh, Sungwon Lyu, Younghoon Jeong, Inkwon Lee, Sangwoo Seo, Dongjun Lee, Hyunwoo Kim, Myeonghwa Lee, Seongbo Jang, Seungwon Do, Sunkyoung Kim, Kyungtae Lim, Jongwon Lee, Kyumin Park, Jamin Shin, Seonghyun Kim, Lucy Park, Alice Oh, Jung-Woo Ha, and Kyunghyun Cho. 2021 · 2021
Later among the works it cites.
XTREME-R: Towards more challenging and nuanced multilingual evaluation
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
RobBERT: a Dutch RoBERTa-based Language Model
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2020 · 2020
Cited alongside, same era.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
FlauBERT: Unsupervised language model pre-training for French
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabbé, Laurent Besacier, and Didier Schwab. 2020 · 2020
Cited alongside, same era.
XGLUE: A new benchmark datasetfor cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Cited alongside, same era.
Zirui Wang, Adams Wei Yu, Orhan Firat, and Yuan Cao. 2021 · 2021
Later among the works it cites.
SICK-NL: A dataset for Dutch natural language inference
Gijs Wijnholds and Michael Moortgat. 2021 · 2021
Later among the works it cites.
Towards a cleaner document-oriented multilingual crawled corpus
Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, and Benoît Sagot. 2022 · 2022
Later among the works it cites.
Language contamination helps explains the cross-lingual capabilities of English pretrained models
Terra Blevins and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Can BERT dig it? Named Entity Recognition for information retrieval in the archaeology domain
Alex Brandsen, Suzan Verberne, Karsten Lambers, and Milco Wansleeben. 2022 · 2022
Later among the works it cites.
Investigating cross-document event coreference for Dutch
Loic De Langhe, Orphee De Clercq, and Veronique Hoste. 2022 · 2022
Later among the works it cites.
Make the best of cross-lingual transfer: Evidence from POS tagging with over 100 languages
Wietse de Vries, Martijn Wieling, and Malvina Nissim. 2022 · 2022
Later among the works it cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
Pieter Delobelle, Ewoenam Tokpo, Toon Calders, and Bettina Berendt. 2022a · 2022
Later among the works it cites.
RobBERT-2022: Updating a dutch language model to account for evolving language use
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2022b · 2022
Later among the works it cites.
ORCA: A challenging benchmark for Arabic language understanding
AbdelRahim Elmadany, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2022 · 2022
Later among the works it cites.
Cross-lingual transfer of monolingual models
Evangelia Gogoulou, Ariel Ekgren, Tim Isbister, and Magnus Sahlgren. 2022 · 2022
Later among the works it cites.
JGLUE: Japanese general language understanding evaluation
Kentaro Kurihara, Daisuke Kawahara, and Tomohide Shibata. 2022 · 2022
Later among the works it cites.
“zo grof !”: A comprehensive corpus for offensive and abusive language in Dutch
Ward Ruitenbeek, Victor Zwart, Robin Van Der Noord, Zhenja Gnezdilov, and Tommaso Caselli. 2022 · 2022
Later among the works it cites.
Exploring language markers of mental health in psychiatric stories
Marco Spruit, Stephanie Verkleij, Kees de Schepper, and Floortje Scheepers. 2022 · 2022
Later among the works it cites.
BasqueGLUE: A natural language understanding benchmark for Basque
Gorka Urbizu, Iñaki San Vicente, Xabier Saralegi, Rodrigo Agerri, and Aitor Soroa. 2022 · 2022
Later among the works it cites.
Universal dependencies 2.11
Daniel Zeman, Joakim Nivre, et al. 2022 · 2022
Later among the works it cites.
DeBERTaV3: Improving deBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2023 · 2023
Closest in time.