Fetching the paper…
Reading the bibliography…
Fine-tuning a pre-trained model (such as BERT, ALBERT, RoBERTa, T5, GPT, etc.) has proven to be one of the most promising paradigms in recent NLP research.
Sub-sampled cubic regularization for non-convex optimization
Jonas Moritz Kohler and Aurelien Lucchi. 2017 · 1904
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff. 2002 · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe. 2004 · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Christopher M Bishop and Nasser M Nasrabadi. 2006 · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B Dolan. 2007 · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan. 2010 · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David. 2014 · 2014
Earlier work this paper cites.
{ \{ TensorFlow } \} : A system for { \{ Large-Scale } \} machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016 · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer. 2016 · 2016
Earlier work this paper cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang. 2018 · 2018
Earlier work this paper cites.
Jax: composable transformations of python+ numpy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne. 2018 · 2018
Earlier work this paper cites.
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos. 2018 · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph Lampert. 2018 · 2018
Earlier work this paper cites.
Towards robust neural networks via random self-ensemble
Xuanqing Liu, Minhao Cheng, Huan Zhang, and Cho-Jui Hsieh. 2018 · 2018
Earlier work this paper cites.
Lectures on convex optimization , volume 137
Yurii Nesterov et al. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. 2018 · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. 2019 · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
The commitmentbank: Investigating projection in naturally occurring discourse
Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Using pre-training can improve model robustness and uncertainty
Dan Hendrycks, Kimin Lee, and Mantas Mazeika. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. 2021 · 2021
Later among the works it cites.
Implicit gradient regularization
David Barrett and Benoit Dherin. 2021 · 2021
Later among the works it cites.
On the usability of transformers-based models for a french question-answering task
Oralie Cattan, Christophe Servan, and Sophie Rosset. 2021 · 2021
Later among the works it cites.
An empirical study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He. 2021 · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ziwei Ji and Matus Telgarsky. 2019 · 2019
Cited alongside, same era.
Mixout: Effective regularization to finetune large-scale pretrained language models
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang. 2019 · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019 · 2019
Cited alongside, same era.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E Peters, Sebastian Ruder, and Noah A Smith. 2019 · 2019
Cited alongside, same era.
Wic: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019 · 2019
Cited alongside, same era.
Stable rank normalization for improved generalization in neural networks and gans
Amartya Sanyal, Philip H Torr, and Puneet K Dokania. 2019 · 2019
Cited alongside, same era.
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021 · 2021
Later among the works it cites.
Robust transfer learning with pretrained language models through adapters
Wenjuan Han, Bo Pang, and Ying Nian Wu. 2021 · 2021
Later among the works it cites.
On the effectiveness of adapter-based tuning for pretrained language model adaptation
Ruidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jiawei Low, Lidong Bing, and Luo Si. 2021 · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. 2021 · 2021
Later among the works it cites.
Noise stability regularization for improving bert fine-tuning
Hang Hua, Xingjian Li, Dejing Dou, Chengzhong Xu, and Jiebo Luo. 2021 · 2021
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2021 · 2021
Later among the works it cites.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021 · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De. 2021 · 2021
Later among the works it cites.
Training neural networks with fixed sparse masks
Yi-Lin Sung, Varun Nair, and Colin A Raffel. 2021 · 2021
Later among the works it cites.
Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning
Colin Wei, Sang Michael Xie, and Tengyu Ma. 2021 · 2021
Later among the works it cites.
Raise a child in large language model: Towards effective and generalizable fine-tuning
Runxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan, Baobao Chang, Songfang Huang, and Fei Huang. 2021 · 2021
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Later among the works it cites.
On the usability of transformers-based models for a french question-answering task
Oralie Cattan, Christophe Servan, and Sophie Rosset. 2022 · 2022
Later among the works it cites.
On the effectiveness of parameter-efficient fine-tuning
Zihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam, Lidong Bing, and Nigel Collier. 2022 · 2022
Later among the works it cites.
Stability vs implicit bias of gradient methods on separable data and beyond
Matan Schliserman and Tomer Koren. 2022 · 2022
Later among the works it cites.
Gradient matching for domain generalization
Yuge Shi, Jeffrey Seely, Philip Torr, N Siddharth, Awni Hannun, Nicolas Usunier, and Gabriel Synnaeve. 2022 · 2022
Later among the works it cites.
How does sharpness-aware minimization minimizes sharpness?
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li. 2022 · 2022
Later among the works it cites.
Chenghao Yang and Xuezhe Ma. 2022 · 2022
Later among the works it cites.
Moebert: from bert to mixture-of-experts via importance-guided adaptation
Simiao Zuo, Qingru Zhang, Chen Liang, Pengcheng He, Tuo Zhao, and Weizhu Chen. 2022 · 2022
Later among the works it cites.