Fetching the paper…
Reading the bibliography…
Fine-tuning transformer models after unsupervised pre-training reaches a very high performance on many different natural language processing tasks.
Constrained deep networks: Lagrangian optimization via log-barrier extensions
Kervadec, H., Dolz, J., Yuan, J., Desrosiers, C., Granger, E., Ayed, I.B., 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019 · 1907
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., Wolf, T., 2019 · 1910
Earlier work this paper cites.
Poor man’s bert: Smaller and faster transformer models
Sajjad, H., Dalvi, F., Durrani, N., Nakov, P., 2020 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases, in: Proceedings of the Third International Workshop on Paraphrasing (IWP2005)
Dolan, W.B., Brockett, C., 2005 · 2005
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge., in: TAC
Bentivogli, L., Clark, P., Dagan, I., Giampiccolo, D., 2009 · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank, in: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Seattle, Washington, USA. pp. 1631–1642
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C.D., Ng, A., Potts, C., 2013 · 2013
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text, in: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392
Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P., 2016 · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation, in: Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pp. 1–14
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., Specia, L., 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization, in: International Conference on Learning Representations
Loshchilov, I., Hutter, F., 2018 · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding, in: Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 353–355
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S., 2018 · 2018
Cited alongside, same era.
Neural network acceptability judgments
Warstadt, A., Singh, A., Bowman, S.R., 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference, in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), Association for Computational Linguistics. pp. 1112–1122
Williams, A., Nangia, N., Bowman, S., 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics, Minneapolis, Minnesota. pp. 4171–4186
Tinybert: Distilling bert for natural language understanding, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, pp. 4163–4174
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., Liu, Q., 2020 · 2020
Later among the works it cites.
Contrastive distillation on intermediate representations for language model compression, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 498–508
Sun, S., Gan, Z., Fang, Y., Cheng, Y., Wang, S., Liu, J., 2020a · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics. pp. 2158–2170
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., Zhou, D., 2020b · 2020
Later among the works it cites.
Multi-task learning for natural language processing in the 2020s: Where are we going?
Worsham, J., Kalita, J., 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2019 · 2019
Cited alongside, same era.
Diagnosis of alzheimer’s disease with sobolev gradient-based optimization and 3d convolutional neural network
Goceri, E., 2019 · 2019
Cited alongside, same era.
Are sixteen heads really better than one?, in: Wallach, H.M., Larochelle, H., Beygelzimer, A., d’Alché Buc, F., Fox, E.B., Garnett, R. (Eds.), NeurIPS, pp. 14014–14024
Michel, P., Levy, O., Neubig, G., 2019 · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression, in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4323–4332
Sun, S., Cheng, Y., Gan, Z., Liu, J., 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned., in: ACL (1), Association for Computational Linguistics. pp. 5797–5808
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., Titov, I., 2019 · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout, in: International Conference on Learning Representations
Fan, A., Grave, E., Joulin, A., 2020 · 2020
Cited alongside, same era.
Capsnet topology to classify tumours from brain images and comparative evaluation
Goceri, E., 2020 · 2020
Cited alongside, same era.
Conflicting bundles: Adapting architectures towards the improved training of deep neural networks, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 256–265
Peer, D., Stabinger, S., Rodriguez-Sanchez, A., 2021a
Cited in the paper.
Limitation of capsule networks
Peer, D., Stabinger, S., Rodríguez-Sánchez, A., 2021b
Cited in the paper.
Compressing large-scale transformer-based models: A case study on bert
Ganesh, P., Chen, Y., Lou, X., Khan, M.A., Yang, Y., Sajjad, H., Nakov, P., Chen, D., Winslett, M., 2021 · 2021
Closest in time.
How to train bert with an academic budget , 10644–10652
Izsak, P., Berchansky, M., Levy, O., 2021 · 2021
Closest in time.
Transformer-based approach for joint handwriting and named entity recognition in historical document
Rouhou, A.C., Dhiaf, M., Kessentini, Y., Salem, S.B., 2021 · 2021
Closest in time.
Document-level relation extraction via graph transformer networks and temporal convolutional networks
Shi, Y., Xiao, Y., Quan, P., Lei, M., Niu, L., 2021 · 2021
Closest in time.
Know what you don’t need: Single-shot meta-pruning for attention heads
Zhang, Z., Qi, F., Liu, Z., Liu, Q., Sun, M., 2021 · 2021
Closest in time.