Fetching the paper…
Reading the bibliography…
Pre-trained language models (PLM) have demonstrated their effectiveness for a broad range of information retrieval and natural language processing tasks.
Ensemble of exemplar-SVMs for object detection and beyond. In IEEE International Conference on Computer Vision, ICCV . IEEE Computer Society, 89–96
Tomasz Malisiewicz, Abhinav Gupta, and Alexei A. Efros. 2011 · 2011
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Principal component analysis: A review and recent developments
Ian T Jolliffe and Jorge Cadima. 2016 · 2016
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Multi-Head Attention with Disagreement Regularization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP . 2897–2903
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Earlier work this paper cites.
Scaling Neural Machine Translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, WMT . 1–9
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Earlier work this paper cites.
An analysis of encoder representations in transformer-based machine translation. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@EMNLP
Alessandro Raganato, Jörg Tiedemann, et al · 2018
Earlier work this paper cites.
Lessons from natural language inference in the clinical domain. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP . 1586–1596
Alexey Romanov and Chaitanya Shivade. 2018 · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT . 1112–1122
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
Unsupervised Feature Learning via Non-Parametric Instance Discrimination. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 . 3733–3742
Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin. 2018 · 2018
Earlier work this paper cites.
Publicly Available Clinical BERT Embeddings
Emily Alsentzer, John R. Murphy, Willie Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew B. A. McDermott. 2019 · 2019
Earlier work this paper cites.
SciBERT: A Pretrained Language Model for Scientific Text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP . 3613–3618
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Earlier work this paper cites.
What Does BERT Look at? An Analysis of BERT’s Attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@ACL . 276–286
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Joint Source-Target Self Attention with Locality Constraints
José A. R. Fonollosa, Noe Casas, and Marta R. Costa-jussà. 2019 · 2019
Earlier work this paper cites.
Revealing the Dark Secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP . 4364–4373
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Are Sixteen Heads Really Better than One?. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems . 14014–14024
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
Transfer Learning in Biomedical Natural Language Processing: An Evaluation of BERT and ELMo on Ten Benchmarking Datasets. In Proceedings of the 18th BioNLP Workshop and Shared Task, BioNLP@ACL . 58–65
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 2019
Cited alongside, same era.
Understanding the Behaviors of BERT in Ranking
Yifan Qiao, Chenyan Xiong, Zhenghao Liu, and Zhiyuan Liu. 2019 · 2019
Cited alongside, same era.
MHSAN: Multi-Head Self-Attention Network for Visual Semantic Embedding. In IEEE Winter Conference on Applications of Computer Vision, WACV . IEEE, 1507–1515
Geondo Park, Chihye Han, Daeshik Kim, and Wonjun Yoon. 2020 · 2020
Later among the works it cites.
Multiple Structural Priors Guided Self Attention Network for Language Understanding
Le Qi, Yu Zhang, Qingyu Yin, and Ting Liu. 2020 · 2020
Later among the works it cites.
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation. In Findings of the Association for Computational Linguistics, EMNLP . 556–568
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2020 · 2020
Later among the works it cites.
On the localness modeling for the self-attention based end-to-end speech synthesis
Shan Yang, Heng Lu, Shiyin Kang, Liumeng Xue, Jinba Xiao, Dan Su, Lei Xie, and Dong Yu. 2020 · 2020
Later among the works it cites.
A direct method to Frobenius norm-based matrix regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP . 3980–3990
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
How to Fine-Tune BERT for Text Classification?. In Chinese Computational Linguistics - 18th China National Conference, CCL , Vol. 11856. 194–206
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. 2019 · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@ACL . 63–76
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Cited alongside, same era.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL . 5797–5808
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Pay Less Attention with Lightweight and Dynamic Convolutions. In 7th International Conference on Learning Representations, ICLR
Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Leveraging Local and Global Patterns for Self-Attention Networks. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL . 3069–3075
Mingzhou Xu, Derek F. Wong, Baosong Yang, Yue Zhang, and Lidia S. Chao. 2019 · 2019
Cited alongside, same era.
Convolutional Self-Attention Networks. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,NAACL-HLT . 4040–4045
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu. 2019 · 2019
Cited alongside, same era.
Guiding Attention for Self-Supervised Learning with Transformers. In Findings of the Association for Computational Linguistics, EMNLP . 4676–4686
Ameet Deshpande and Karthik Narasimhan. 2020 · 2020
Cited alongside, same era.
Shi-Fang Yuan, Yi-Bin Yu, Ming-Zhao Li, and Hua Jiang. 2020 · 2020
Later among the works it cites.
Querying across genres for medical claims in news. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP . 1783–1789
Chaoyuan Zuo, Narayan Acharya, and Ritwik Banerjee. 2020 · 2020
Later among the works it cites.
How Linguistically Fair Are Multilingual Pre-Trained Language Models?. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI . 12710–12718
Monojit Choudhury and Amit Deshpande. 2021 · 2021
Later among the works it cites.
Thank you BART! Rewarding Pre-Trained Models Improves Formality Style Transfer. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP . 484–494
Huiyuan Lai, Antonio Toral, and Malvina Nissim. 2021 · 2021
Later among the works it cites.
Clustering-friendly Representation Learning via Instance Discrimination and Feature Decorrelation. In 9th International Conference on Learning Representations, ICLR 2021
Yaling Tao, Kentaro Takagi, and Kouta Nakata. 2021 · 2021
Later among the works it cites.
BERT-DRE: BERT with Deep Recursive Encoder for Natural Language Sentence Matching
Ehsan Tavan, Ali Rahmati, Maryam Najafi, Saeed Bibak, and Zahed Rahmati. 2021 · 2021
Later among the works it cites.
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers. In Findings of the Association for Computational Linguistics, ACL/IJCNLP (Findings of ACL, Vol. ACL/IJCNLP 2021) . 2140–2151
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021a · 2021
Later among the works it cites.
Dodrio: Exploring Transformer Models with Interactive Visualization
Zijie J. Wang, Robert Turko, and Duen Horng Chau. 2021b · 2021
Later among the works it cites.
Using Prior Knowledge to Guide BERT’s Attention in Semantic Textual Matching Tasks. In WWW ’21: The Web Conference 2021 . 2466–2475
Tingyu Xia, Yue Wang, Yuan Tian, and Yi Chang. 2021 · 2021
Later among the works it cites.
Pretrained Transformers for Text Ranking: BERT and Beyond. In SIGIR ’21: The 44th International ACM Conference on Research and Development in Information Retrieval, SIGIR . 2666–2668
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Later among the works it cites.
Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning, ICML , Vol. 139. 12310–12320
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. 2021 · 2021
Later among the works it cites.