Fetching the paper…
Reading the bibliography…
Latest instruction-tuned large language models (LLM) show great results on various tasks, however, they often face performance degradation for non-English input.
S. Lakew, et al, ”Transfer Learning in Multilingual Neural Machine Translation with Dynamic Vocabulary,” in Proceedings of the 15th International Conference on Spoken Language Translation, 2018, pp. 54–61
2018
Earlier work this paper cites.
W. Yoon, et al, ”Pre-trained language model for biomedical question answering,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 2019, pp. 727–740
2019
Earlier work this paper cites.
J. Devlin, et al, ”BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 4171–4186
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown, et al, ”Language Models are Few-Shot Learners” //arXiv preprint arXiv:2005.14165. – 2020
2020
Earlier work this paper cites.
J. Zhang, et al, ”PEGASUS: Pre-Training with Extracted Gap-Sentences for Abstractive Summarization,” in Proceedings of the 37th International Conference on Machine Learning, 2020
2020
Earlier work this paper cites.
M. Tikhomirov, et al, ”Using bert and augmentation in named entity recognition for cybersecurity domain,” in Natural Language Processing and Information Systems: 25th International Conference on Applications of Natural Language to Information Systems, NLDB 2020, Saarbrucken, Germany, June 24–26, 2020, Proceedings 25, 2020, pp. 16–24
2020
Cited alongside, same era.
K. Bostrom, G. Durrett, ”Byte Pair Encoding is Suboptimal for Language Model Pretraining,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 4617–4624
2020
Cited alongside, same era.
A. Nayak, et al, ”Domain adaptation challenges of BERT in tokenization and sub-word representations of Out-of-Vocabulary words,” in Proceedings of the First Workshop on Insights from Negative Results in NLP, 2020, pp. 1–5
2020
Cited alongside, same era.
W. Vries, M. Nissim, ”As Good as New. How to Successfully Recycle English GPT-2 to Make Models for Other Languages,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2021, pp. 836–846
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
P. Rust, et al, ”How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, pp. 3118–3135
2021
Cited alongside, same era.
V. Hofmann, J. Pierrehumbert, H. Schutze, ”Superbizarre Is Not Superb: Derivational Morphology Improves BERT’s Interpretation of Complex Words,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, pp. 3594–3608
2021
Cited alongside, same era.
2023
Closest in time.