Fetching the paper…
Reading the bibliography…
Fine-tuning pre-trained language models such as BERT has become a common practice dominating leaderboards across various NLP tasks.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Fine-tune bert for extractive summarization
Yang Liu. 2019 · 1903
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, F. Hill, Omer Levy, and Samuel R. Bowman. 2019 · 1905
Earlier work this paper cites.
Adversarial training generalizes data-dependent spectral norm regularization
K. Roth, Yannic Kilcher, and Thomas Hofmann. 2019 · 1906
Earlier work this paper cites.
Stable rank normalization for improved generalization in neural networks and gans
Amartya Sanyal, P. Torr, and Puneet K. Dokania. 2020 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 1909
Earlier work this paper cites.
Mixout: Effective regularization to finetune large-scale pretrained language models
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang. 2020 · 1909
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2019 · 1910
Earlier work this paper cites.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Brian W Matthews. 1975 · 1975
Earlier work this paper cites.
Solutions of ill-posed problems (an tikhonov and vy arsenin)
Ralph A Willoughby. 1979 · 1979
Earlier work this paper cites.
Improving the generalization properties of radial basis function neural networks
Chris Bishop. 1991 · 1991
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat. 1992 · 1992
Earlier work this paper cites.
Training with noise is equivalent to tikhonov regularization
Chris M Bishop. 1995 · 1995
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah A. Smith. 2020 · 2002
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
The pascal recognising textual entailment challenge
I. Dagan, Oren Glickman, and B. Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
W. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Roy Bar-Haim, I. Dagan, B. Dolan, Lisa Ferro, Danilo Giampiccolo, and B. Magnini. 2006 · 2006
Earlier work this paper cites.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Marius Mosbach, Maksym Andriushchenko, and D. Klakow. 2020 · 2006
Cited alongside, same era.
Revisiting few-sample bert fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2006
Cited alongside, same era.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, B. Magnini, I. Dagan, and W. Dolan. 2007 · 2007
Cited alongside, same era.
Better fine-tuning by reducing representational collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta. 2020 · 2008
Cited alongside, same era.
The difficulty of training deep architectures and the effect of unsupervised pre-training
D. Erhan, Pierre-Antoine Manzagol, Yoshua Bengio, S. Bengio, and P. Vincent. 2009 · 2009
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, L. Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Spectral norm regularization for improving the generalizability of deep learning
Y. Yoshida and Takeru Miyato. 2017 · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, R. Ge, Behnam Neyshabur, and Yi Zhang. 2018 · 2018
Later among the works it cites.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, A. Sabharwal, and Peter Clark. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
D. Erhan, Aaron C. Courville, Yoshua Bengio, and Pascal Vincent. 2010 · 2010
Cited alongside, same era.
Supervised contrastive learning for pre-trained language model fine-tuning
Beliz Gunel, Jingfei Du, Alexis Conneau, and Ves Stoyanov. 2020 · 2011
Cited alongside, same era.
Adding noise to the input of a model trained with a regularized objective
Salah Rifai, Xavier Glorot, Yoshua Bengio, and Pascal Vincent. 2011 · 2011
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, G. S. Corrado, and J. Dean. 2013 · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
Li Wan, Matthew D. Zeiler, Sixin Zhang, Y. LeCun, and R. Fergus. 2013 · 2013
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, R. Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Yoshua Bengio, and Hod Lipson. 2014 · 2014
Cited alongside, same era.
Later among the works it cites.
Explicit inductive bias for transfer learning with convolutional networks
Xuhong Li, Y. Grandvalet, and F. Davoine. 2018 · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Y. Bahri, D. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
A. Radford. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, F. Hill, Omer Levy, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, Jeffrey Wu, R. Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Entity, relation, and event extraction with contextualized span representations
David Wadden, Ulme Wennberg, Yi Luan, and Hannaneh Hajishirzi. 2019 · 2019
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Zihang Dai, Yiming Yang, J. Carbonell, R. Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2020 · 2020
Later among the works it cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Later among the works it cites.
Incorporating bert into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, W. Zhou, H. Li, and T. Liu. 2020b · 2048
Closest in time.