Fetching the paper…
Reading the bibliography…
This paper explores the performance of encoder and decoder language models on multilingual Natural Language Understanding (NLU) tasks, with a broad focus on Germanic languages.
Language models are few-shot learners
Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and others. 2020 · 1901
Earlier work this paper cites.
Liii. on lines and planes of closest fit to systems of points in space
Karl Pearson. 1901 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
The alpino dependency treebank
L van der Beek, G Bouma, R Malouf, and G van Noord. 2002 · 2002
Earlier work this paper cites.
Stochastic neighbor embedding
Geoffrey E Hinton and Sam Roweis. 2002 · 2002
Earlier work this paper cites.
Introduction to the conll-2002 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang. 2002 · 2002
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2003
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Manual of the Stockholm Umeå corpus version 2.0
Sofia Gustafson-Capková and Britt Hartmann. 2006 · 2006
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2006
Earlier work this paper cites.
Universal dependency annotation for multilingual parsing
Ryan McDonald, Joakim Nivre, Yvonne Quirmbach-Brundage, Yoav Goldberg, Dipanjan Das, Kuzman Ganchev, Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, et al. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Nosta-d named entity annotation for german: Guidelines and dataset
Darina Benikova, Chris Biemann, and Marc Reznicek. 2014 · 2014
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
A twitter corpus and benchmark resources for german sentiment analysis
Mark Cieliebak, Jan Milan Deriu, Dominic Egger, and Fatih Uzdilli. 2017 · 2017
Earlier work this paper cites.
The JSON data interchange syntax
ISO/IEC 21778:2017. 2017 · 2017
Earlier work this paper cites.
Sentiment Analysis With Convolutional Neural Networks: Classifying sentiment in Swedish reviews
Kristoffer Svensson. 2017 · 2017
Earlier work this paper cites.
The gum corpus: creating multilayer resources in the classroom
Amir Zeldes. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. 2018 · 2018
Cited alongside, same era.
NoReC: The Norwegian review corpus
Erik Velldal, Lilja Øvrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, and Fredrik Jørgensen. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Massively multilingual transfer for ner
Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019 · 2019
Cited alongside, same era.
Dane: A named entity resource for danish
Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted, Lasse Malm Lidegaard, and Anders Søgaard. 2020 · 2020
Cited alongside, same era.
Chatgpt: Jack of all trades, master of none
Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, et al. 2023 · 2023
Later among the works it cites.
Outlines
Rémi Louf. 2023 · 2023
Later among the works it cites.
ScandEval: A Benchmark for Scandinavian Natural Language Processing
Dan Saattrup Nielsen. 2023 · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in transformers
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Max Vladymyrov, Razvan Pascanu, et al. 2023 · 2023
Later among the works it cites.
Norbench–a benchmark for norwegian language models
David Samuel, Andrey Kutuzov, Samia Touileb, Erik Velldal, Lilja Øvrelid, Egil Rønningstad, Elina Sigdel, and Anna Palatkina. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Svanhvít L Ingólfsdóttir, Ásmundur A Gudjónsson, and Hrafn Loftsson. 2020 · 2020
Cited alongside, same era.
NorNE: Annotating named entities for Norwegian
Fredrik Jørgensen, Tobias Aasmoe, Anne-Stine Ruud Husevåg, Lilja Øvrelid, and Erik Velldal. 2020 · 2020
Cited alongside, same era.
Germanquad and germandpr: Improving non-english question answering and passage retrieval
Timo Möller, Julian Risch, and Malte Pietsch. 2021 · 2021
Cited alongside, same era.
Danlp: An open-source toolkit for danish natural language processing
Amalie Brogaard Pauli, Maria Barrett, Ophélie Lacroix, and Rasmus Hvingelby. 2021 · 2021
Cited alongside, same era.
dutchsocial · Datasets at Hugging Face — huggingface.co
Aakash Gupta. 2022 · 2022
Cited alongside, same era.
Natural questions in icelandic
Vésteinn Snæbjarnarson and Hafsteinn Einarsson. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Transfer to a low-resource language via close relatives: The case study on faroese
Vésteinn Snæbjarnarson, Annika Simonsen, Goran Glavaš, and Ivan Vulić. 2023 · 2023
Later among the works it cites.
Dumb: A benchmark for smart evaluation of dutch models
Wietse de Vries, Martijn Wieling, and Malvina Nissim. 2023 · 2023
Later among the works it cites.
Is chatgpt a good sentiment analyzer? a preliminary study
Zengzhi Wang, Qiming Xie, Yi Feng, Zixiang Ding, Zinong Yang, and Rui Xia. 2023 · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Later among the works it cites.
Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023 · 2023
Later among the works it cites.
Open LLM Leaderboard - a Hugging Face Space by open-llm-leaderboard — huggingface.co
2024
Closest in time.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. 2024 · 2024
Closest in time.
Are gllms danoliterate? benchmarking generative nlp in danish
Søren Vejlgaard Holm. 2024 · 2024
Closest in time.
Government of Iceland: How Iceland is using GPT-4 to preserve its language
OpenAI. 2023a · 2024
Closest in time.
New models and developer products announced at DevDay
OpenAI. 2023b · 2024
Closest in time.
Chatgpt and finetuned bert: A comparative study for developing intelligent design support systems
Yunjian Qiu and Yan Jin. 2024 · 2024
Closest in time.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al. 2024 · 2024
Closest in time.