Fetching the paper…
Reading the bibliography…
Transformer-based pretrained language models (T-PTLMs) have achieved great success in almost every NLP task.
P. Singh, T. Lin, E. T. Mueller, G. Lim, T. Perkins, and W. L. Zhu, “Open mind common sense: Knowledge acquisition from the general public,” in OTM Confederated International Conferences” On the Move to Meaningful Internet Systems” . Springer, 2002, pp. 1223–1237
2002
Earlier work this paper cites.
O. Bodenreider, “The unified medical language system (umls): integrating biomedical terminology,” Nucleic acids research , vol. 32, no. suppl_1, pp. D267–D270, 2004
2004
Earlier work this paper cites.
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil, “Model compression,” in Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining , 2006, pp. 535–541
2006
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering , vol. 22, no. 10, pp. 1345–1359, 2009
2009
Earlier work this paper cites.
D. Erhan, A. Courville, Y. Bengio, and P. Vincent, “Why does unsupervised pre-training help deep learning?” in Proceedings of the thirteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 2010, pp. 201–208
2010
Earlier work this paper cites.
C. Fellbaum, “Wordnet,” in Theory and applications of ontology: computer applications . Springer, 2010, pp. 231–243
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems , vol. 25, pp. 1097–1105, 2012
2012
Earlier work this paper cites.
R. Speer, C. Havasi et al. , “Representing general relational knowledge in conceptnet 5.” in LREC , vol. 2012, 2012, pp. 3679–86
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
C. Xu, D. Tao, and C. Xu, “A survey on multi-view learning,” arXiv preprint arXiv:1304.5634 , 2013
2013
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
N. Kalchbrenner, E. Grefenstette, and P. Blunsom, “A convolutional neural network for modelling sentences,” in Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2014, pp. 655–665
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
L. J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Proceedings of the 27th International Conference on Neural Information Processing Systems-Volume 2 , 2014, pp. 2654–2662
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in neural information processing systems , vol. 28, pp. 91–99, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
P. Liu, X. Qiu, and X. Huang, “Recurrent neural network for text classification with multi-task learning,” in Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence , 2016, pp. 2873–2879
2016
Earlier work this paper cites.
P. Zhou, Z. Qi, S. Zheng, J. Xu, H. Bao, and B. Xu, “Text classification improved by integrating bidirectional lstm with two-dimensional max pooling,” in Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers , 2016, pp. 3485–3495
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2016, pp. 1715–1725
2016
Earlier work this paper cites.
A. E. Johnson, T. J. Pollard, L. Shen, H. L. Li-Wei, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark, “Mimic-iii, a freely accessible critical care database,” Scientific data , vol. 3, no. 1, pp. 1–9, 2016
2016
Earlier work this paper cites.
M. Ziemski, M. Junczys-Dowmunt, and B. Pouliquen, “The united nations parallel corpus v1. 0,” in Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16) , 2016, pp. 3530–3534
2016
Earlier work this paper cites.
R. He and J. McAuley, “Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering,” in proceedings of the 25th international conference on world wide web , 2016, pp. 507–517
2016
Earlier work this paper cites.
I. Yamada, H. Shindo, H. Takeda, and Y. Takefuji, “Joint learning of the embedding of words and entities for named entity disambiguation,” in Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning , 2016, pp. 250–259
2016
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100,000+ questions for machine comprehension of text,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , 2016, pp. 2383–2392
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , 2018, pp. 66–71
2018
Earlier work this paper cites.
J. A. Wagner Filho, R. Wilkens, M. Idiart, and A. Villavicencio, “The brwac corpus: A new open resource for brazilian portuguese,” in Proceedings of the eleventh international conference on language resources and evaluation (LREC 2018) , 2018
2018
Earlier work this paper cites.
A. Kunchukuttan, P. Mehta, and P. Bhattacharyya, “The iit bombay english-hindi parallel corpus,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) , 2018
2018
Earlier work this paper cites.
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , 2018, pp. 2227–2237
2018
Earlier work this paper cites.
A. Akbik, D. Blythe, and R. Vollgraf, “Contextual string embeddings for sequence labeling,” in Proceedings of the 27th international conference on computational linguistics , 2018, pp. 1638–1649
2018
Earlier work this paper cites.
T. Kudo, “Subword regularization: Improving neural network translation models with multiple subword candidates,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 66–75
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP , 2018, pp. 353–355
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” Advances in Neural Information Processing Systems , vol. 32, pp. 5753–5763, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning, “Electra: Pre-training text encoders as discriminators rather than generators,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “Albert: A lite bert for self-supervised learning of language representations,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International Conference on Machine Learning . PMLR, 2019, pp. 6105–6114
2019
Earlier work this paper cites.
T. Kaur and T. K. Gandhi, “Automated brain image classification based on vgg-16 and transfer learning,” in 2019 International Conference on Information Technology (ICIT) . IEEE, 2019, pp. 94–98
2019
Earlier work this paper cites.
I. Beltagy, K. Lo, and A. Cohan, “Scibert: A pretrained language model for scientific text,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 3606–3611
2019
Earlier work this paper cites.
E. Alsentzer, J. Murphy, W. Boag, W.-H. Weng, D. Jindi, T. Naumann, and M. McDermott, “Publicly available clinical bert embeddings,” in Proceedings of the 2nd Clinical Natural Language Processing Workshop , 2019, pp. 72–78
2019
Earlier work this paper cites.
Y. Peng, S. Yan, and Z. Lu, “Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets,” in Proceedings of the 18th BioNLP Workshop and Shared Task , 2019, pp. 58–65
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Arkhipov, M. Trofimova, Y. Kuratov, and A. Sorokin, “Tuning multilingual transformers for language-specific named entity recognition,” in Proceedings of the 7th Workshop on Balto-Slavic Natural Language Processing , 2019, pp. 89–93
2019
Earlier work this paper cites.
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh, “Large batch optimization for deep learning: Training bert in 76 minutes,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
J. Ni, J. Li, and J. McAuley, “Justifying recommendations using distantly-labeled reviews and fine-grained aspects,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 188–197
2019
Earlier work this paper cites.
R. Zellers, A. Holtzman, H. Rashkin, Y. Bisk, A. Farhadi, F. Roesner, and Y. Choi, “Defending against neural fake news,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems , 2019, pp. 9054–9065
2019
Earlier work this paper cites.
P. J. O. Suárez, B. Sagot, and L. Romary, “Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures,” in 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7) . Leibniz-Institut für Deutsche Sprache, 2019
2019
Earlier work this paper cites.
H. Huang, Y. Liang, N. Duan, M. Gong, L. Shou, D. Jiang, and M. Zhou, “Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 2485–2494
2019
Earlier work this paper cites.
S. Gururangan, T. Dang, D. Card, and N. A. Smith, “Variational pretraining for semi-supervised text classification,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 5880–5894
2019
Earlier work this paper cites.
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu, “Mass: Masked sequence to sequence pre-training for language generation,” in International Conference on Machine Learning . PMLR, 2019, pp. 5926–5936
2019
Earlier work this paper cites.
L. Dong, N. Yang, W. Wang, F. Wei, X. Liu, Y. Wang, J. Gao, M. Zhou, and H.-W. Hon, “Unified language model pre-training for natural language understanding and generation,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems , 2019, pp. 13 063–13 075
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Polignano, P. Basile, M. De Gemmis, G. Semeraro, and V. Basile, “Alberto: Italian bert language understanding model for nlp challenging tasks based on tweets,” in 6th Italian Conference on Computational Linguistics, CLiC-it 2019 , vol. 2481. CEUR, 2019, pp. 1–6
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov, “Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 5797–5808
2019
Earlier work this paper cites.
A. Fan, E. Grave, and A. Joulin, “Reducing transformer depth on demand with structured dropout,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Sun, Y. Cheng, Z. Gan, and J. Liu, “Patient knowledge distillation for bert model compression,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 4323–4332
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning in nlp,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 3645–3650
2019
Earlier work this paper cites.
Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, “Ernie: Enhanced language representation with informative entities,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 1441–1451
2019
Earlier work this paper cites.
M. E. Peters, M. Neumann, R. Logan, R. Schwartz, V. Joshi, S. Singh, and N. A. Smith, “Knowledge enhanced contextual word representations,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 43–54
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 3982–3992
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
X. Liu, P. He, W. Chen, and J. Gao, “Multi-task deep neural networks for natural language understanding,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 4487–4496
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Cengiz, U. Sert, and D. Yuret, “Ku_ai at mediqa 2019: Domain-specific pre-training and transfer learning for medical nli,” in Proceedings of the 18th BioNLP Workshop and Shared Task , 2019, pp. 427–436
2019
Earlier work this paper cites.
W. Yoon, J. Lee, D. Kim, M. Jeong, and J. Kang, “Pre-trained language model for biomedical question answering,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2019, pp. 727–740
2019
Earlier work this paper cites.
A. Wang, J. Hula, P. Xia, R. Pappagari, R. T. McCoy, R. Patel, N. Kim, I. Tenney, Y. Huang, K. Yu et al. , “Can you tell me how to get past sesame street? sentence-level pretraining beyond language modeling,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 4465–4476
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
A. C. Stickland and I. Murray, “Bert and pals: Projected attention layers for efficient adaptation in multi-task learning,” in International Conference on Machine Learning . PMLR, 2019, pp. 5986–5995
2019
Earlier work this paper cites.
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky, “Revealing the dark secrets of bert,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 4365–4374
2019
Earlier work this paper cites.
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller, “Language models as knowledge bases?” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 2463–2473
2019
Earlier work this paper cites.
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh, “Universal adversarial triggers for attacking and analyzing nlp,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 2153–2162
2019
Earlier work this paper cites.
H. Elsahar, P. Vougiouklis, A. Remaci, C. Gravier, J. Hare, E. Simperl, and F. Laforest, “T-rex: A large scale alignment of natural language with knowledge base triples,” 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “Fairseq: A fast, extensible toolkit for sequence modeling,” NAACL HLT 2019 , p. 48, 2019
2019
Earlier work this paper cites.
J. Vig, “Bertviz: A tool for visualizing multihead self-attention in the bert model,” in ICLR Workshop: Debugging Machine Learning Models , 2019
2019
Earlier work this paper cites.
D. Pruthi, B. Dhingra, and Z. C. Lipton, “Combating adversarial misspellings with robust word recognition,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 5582–5591
2019
Earlier work this paper cites.
V. Misra, “Black box attacks on transformer language models,” in ICLR 2019 Debugging Machine Learning Models Workshop , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7871–7880
2020
Earlier work this paper cites.
J. Zhang, Y. Zhao, M. Saleh, and P. Liu, “Pegasus: Pre-training with extracted gap-sentences for abstractive summarization,” in International Conference on Machine Learning . PMLR, 2020, pp. 11 328–11 339
2020
Earlier work this paper cites.
X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, and X. Huang, “Pre-trained models for natural language processing: A survey,” Science China Technological Sciences , pp. 1–26, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig, “How can we know what language models know?” Transactions of the Association for Computational Linguistics , vol. 8, pp. 423–438, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh, “Eliciting knowledge from language models using automatically generated prompts,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 4222–4235
2020
Later among the works it cites.
N. Kassner and H. Schütze, “Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7811–7818
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, “Don’t stop pretraining: Adapt language models to domains and tasks,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 8342–8360
2020
Cited alongside, same era.
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang et al. , “Codebert: A pre-trained model for programming and natural languages,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings , 2020, pp. 1536–1547
2020
Cited alongside, same era.
2020
Cited alongside, same era.
C.-S. Wu, S. C. Hoi, R. Socher, and C. Xiong, “Tod-bert: Pre-trained natural language understanding for task-oriented dialogue,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 917–929
2020
Cited alongside, same era.
A. Louis, “Netbert: A pre-trained language representation model for computer networking,” Ph.D. dissertation, Cisco Systems, 2020
2020
Cited alongside, same era.
J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, and J. Kang, “Biobert: a pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics , vol. 36, no. 4, pp. 1234–1240, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Schick and H. Schütze, “Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8766–8774
2020
Later among the works it cites.
Z. Jiang, A. Anastasopoulos, J. Araki, H. Ding, and G. Neubig, “X-factr: Multilingual factual knowledge retrieval from pretrained language models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 5943–5959
2020
Later among the works it cites.
J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff, “Masked language model scoring,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 2699–2712
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Liang, N. Duan, Y. Gong, N. Wu, F. Guo, W. Qi, M. Gong, L. Shou, D. Jiang, G. Cao et al. , “Xglue: A new benchmark datasetfor cross-lingual pre-training, understanding and generation,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 6008–6018
2020
Later among the works it cites.
G. Aguilar, S. Kar, and T. Solorio, “Lince: A centralized benchmark for linguistic code-switching evaluation,” in Proceedings of The 12th Language Resources and Evaluation Conference , 2020, pp. 1803–1813
2020
Later among the works it cites.
S. Khanuja, S. Dandapat, A. Srinivasan, S. Sitaram, and M. Choudhury, “Gluecos: An evaluation benchmark for code-switched nlp,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 3575–3585
2020
Later among the works it cites.
J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, “Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation,” in International Conference on Machine Learning . PMLR, 2020, pp. 4411–4421
2020
Later among the works it cites.
T. Shavrina, A. Fenogenova, E. Anton, D. Shevelev, E. Artemova, V. Malykh, V. Mikhailov, M. Tikhonova, A. Chertok, and A. Evlampiev, “Russiansuperglue: A russian language understanding evaluation benchmark,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 4717–4726
2020
Later among the works it cites.
L. Xu, H. Hu, X. Zhang, L. Li, C. Cao, Y. Li, Y. Xu, K. Sun, D. Yu, C. Yu et al. , “Clue: A chinese language understanding evaluation benchmark,” in Proceedings of the 28th International Conference on Computational Linguistics , 2020, pp. 4762–4772
2020
Later among the works it cites.
2020
Later among the works it cites.
T. Wolf, J. Chaumond, L. Debut, V. Sanh, C. Delangue, A. Moi, P. Cistac, M. Funtowicz, J. Davison, S. Shleifer et al. , “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , 2020, pp. 38–45
2020
Later among the works it cites.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020, pp. 3505–3506
2020
Later among the works it cites.
B. Hoover, H. Strobelt, and S. Gehrmann, “exbert: A visual analysis tool to explore learned representations in transformer models,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , 2020, pp. 187–196
2020
Later among the works it cites.
Z. Yang, Y. Cui, Z. Chen, W. Che, T. Liu, S. Wang, and G. Hu, “Textbrewer: An open-source knowledge distillation toolkit for natural language processing,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , 2020, pp. 9–16
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Hisamoto, M. Post, and K. Duh, “Membership inference attacks on sequence-to-sequence models: Is my data in your machine translation system?” Transactions of the Association for Computational Linguistics , vol. 8, pp. 49–63, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Mosbach, M. Andriushchenko, and D. Klakow, “On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines,” in International Conference on Learning Representations , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
W. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang, “Unified pre-training for program understanding and generation,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 2655–2668
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
S. Panda, A. Agrawal, J. Ha, and B. Bloch, “Shuffled-token detection for refining pre-trained roberta,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop , 2021, pp. 88–93
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
W.-R. Chen, M. Abdul-Mageed, H. Cavusoglu et al. , “Indt5: A text-to-text transformer for 10 indigenous languages,” in Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas , 2021, pp. 265–271
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
S. Yuan, H. Zhao, Z. Du, M. Ding, X. Liu, Y. Cen, X. Zou, Z. Yang, and J. Tang, “Wudaocorpora: A super large-scale chinese corpora for pre-training language models,” AI Open , 2021
2021
Closest in time.
2021
Closest in time.
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 483–498
2021
Closest in time.
H. Schwenk, V. Chaudhary, S. Sun, H. Gong, and F. Guzmán, “Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , 2021, pp. 1351–1361
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Y. Meng, W. F. Speier, M. K. Ong, and C. Arnold, “Bidirectional representation learning from transformers using multimodal electronic health record data to predict depression,” IEEE Journal of Biomedical and Health Informatics , 2021
2021
Closest in time.
L. Rasmy, Y. Xiang, Z. Xie, C. Tao, and D. Zhi, “Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction,” NPJ digital medicine , vol. 4, no. 1, pp. 1–13, 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
——, “Aragpt2: Pre-trained transformer for arabic language generation,” in Proceedings of the Sixth Arabic Natural Language Processing Workshop , 2021, pp. 196–207
2021
Closest in time.
——, “Araelectra: Pre-training text discriminators for arabic language understanding,” in Proceedings of the Sixth Arabic Natural Language Processing Workshop , 2021, pp. 191–195
2021
Closest in time.
2021
Closest in time.
H. Lee, J. Yoon, B. Hwang, S. Joe, S. Min, and Y. Gwon, “Korealbert: Pretraining a lite bert model for korean language understanding,” in 2020 25th International Conference on Pattern Recognition (ICPR) . IEEE, 2021, pp. 5551–5557
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
X. Cheng, “Dual-view distilled bert for sentence embedding,” arXiv preprint arXiv:2104.08675 , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
F. Liu, E. Shareghi, Z. Meng, M. Basaldella, and N. Collier, “Self-alignment pretraining for biomedical entity representations,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 4228–4238
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
N. Kassner, P. Dufter, and H. Schütze, “Multilingual lama: Investigating knowledge in multilingual pretrained language models,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , 2021, pp. 3250–3258
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
T. Schick and H. Schütze, “It’s not just size that matters: Small language models are also few-shot learners,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 2339–2352
2021
Closest in time.
2021
Closest in time.
X. Wang, Y. Xiong, Y. Wei, M. Wang, and L. Li, “Lightseq: A high performance inference library for transformers,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers , 2021, pp. 113–120
2021
Closest in time.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient gpu serving system for transformer models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 389–402
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
J. Wang, G. Zhang, W. Wang, K. Zhang, and Y. Sheng, “Cloud-based intelligent self-diagnosis and department recommendation service using chinese medical bert,” Journal of Cloud Computing , vol. 10, no. 1, pp. 1–12, 2021
2021
Closest in time.
P. P. Liang, C. Wu, L.-P. Morency, and R. Salakhutdinov, “Towards understanding and mitigating social biases in language models,” in International Conference on Machine Learning . PMLR, 2021, pp. 6565–6576
2021
Closest in time.