Fetching the paper…
Reading the bibliography…
While fine-tuning of pre-trained language models generally helps to overcome the lack of labelled training samples, it also displays model performance instability.
ChatGPT to replace crowdsourcing of paraphrases for intent classification: Higher diversity and comparable model robustness
Jan Cegin, Jakub Simko, and Peter Brusilovsky. 2023 · 1905
Earlier work this paper cites.
Augmenting data with mixup for sentence classification: An empirical study
Hongyu Guo, Yongyi Mao, and Richong Zhang. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. 2019 · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. 2022 · 1965
Earlier work this paper cites.
Building a question answering test collection
Ellen M. Voorhees and Dawn M. Tice. 2000 · 2000
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
DBpedia – a large-scale, multilingual knowledge base extracted from wikipedia
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick van Kleef, Sören Auer, and Christian Bizer. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Snapshot ensembles: Train 1, get m for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger. 2017 · 2017
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al. 2018 · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P Izmailov, AG Wilson, D Podoprikhin, D Vetrov, and T Garipov. 2018 · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018 · 2018
Earlier work this paper cites.
Metainit: Initializing learning by learning to initialize
Yann N Dauphin and Samuel Schoenholz. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Nonlinear mixup: Out-of-manifold data augmentation for text classification
Hongyu Guo. 2020 · 2020
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Cited alongside, same era.
Mixout: Effective regularization to finetune large-scale pretrained language models
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang. 2020 · 2020
Cited alongside, same era.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance
CAMERO: Consistency regularized ensemble of perturbed language models with weight sharing
Chen Liang, Pengcheng He, Yelong Shen, Weizhu Chen, and Tuo Zhao. 2022 · 2022
Later among the works it cites.
Improving generalization of pre-trained language models via stochastic weight averaging
Peng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh, Ahmad Rashid, Ali Ghodsi, and Phillippe Langlais. 2022 · 2022
Later among the works it cites.
UniPELT: A unified framework for parameter-efficient language model tuning
Yuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Jiawei Han, Scott Yih, and Madian Khabsa. 2022 · 2022
Later among the works it cites.
The MultiBERTs: BERT Reproductions for Robustness Analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, and Ellie Pavlick. 2022 · 2022
Later among the works it cites.
NoisyTune: A little noise can help you finetune pretrained language models better
Chuhan Wu, Fangzhao Wu, Tao Qi, and Yongfeng Huang. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Thomas McCoy, Junghyun Min, and Tal Linzen. 2020 · 2020
Cited alongside, same era.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Cited alongside, same era.
A survey of data augmentation approaches for NLP
Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. 2021 · 2021
Cited alongside, same era.
Noise stability regularization for improving BERT fine-tuning
Hang Hua, Xingjian Li, Dejing Dou, Chengzhong Xu, and Jiebo Luo. 2021 · 2021
Cited alongside, same era.
How emotionally stable is ALBERT? testing robustness with stochastic weight averaging on a sentiment analysis task
Urja Khurana, Eric Nalisnick, and Antske Fokkens. 2021 · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Cited alongside, same era.
Fine-tuning pre-trained language models effectively by optimizing subnetworks adaptively
Haojie Zhang, Ge Li, Jia Li, Zhongjin Zhang, Yuqi Zhu, and Zhi Jin. 2022 · 2022
Later among the works it cites.
Multi-CLS BERT: An efficient alternative to traditional ensembling
Haw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, and Andrew McCallum. 2023 · 2023
Later among the works it cites.
PTP: Boosting stability and performance of prompt tuning with perturbation-based regularizer
Lichang Chen, Jiuhai Chen, Heng Huang, and Minhao Cheng. 2023 · 2023
Later among the works it cites.
Knowledge is a region in weight space for fine-tuned language models
Almog Gueta, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen. 2023 · 2023
Later among the works it cites.
Improving pretrained language model fine-tuning with noise stability regularization
Hang Hua, Xingjian Li, Dejing Dou, Cheng-Zhong Xu, and Jiebo Luo. 2023 · 2023
Later among the works it cites.
A rank stabilization scaling factor for fine-tuning with lora
Damjan Kalajdzievski. 2023 · 2023
Later among the works it cites.
MEAL: Stable and active learning for few-shot prompting
Abdullatif Köksal, Timo Schick, and Hinrich Schuetze. 2023 · 2023
Later among the works it cites.
Gpt understands, too
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2023 · 2023
Later among the works it cites.
Effectiveness of data augmentation for parameter efficient tuning with limited data
Stephen Obadinma, Hongyu Guo, and Xiaodan Zhu. 2023 · 2023
Later among the works it cites.
Self-supervised meta-prompt learning with meta-gradient regularization for few-shot generalization
Kaihang Pan, Juncheng Li, Hongye Song, Jun Lin, Xiaozhong Liu, and Siliang Tang. 2023 · 2023
Later among the works it cites.
Residual prompt tuning: improving prompt tuning with residual reparameterization
Anastasiia Razdaibiedina, Yuning Mao, Madian Khabsa, Mike Lewis, Rui Hou, Jimmy Ba, and Amjad Almahairi. 2023 · 2023
Later among the works it cites.
Two-stage fine-tuning for improved bias and variance for large pretrained language models
Lijing Wang, Yingya Li, Timothy Miller, Steven Bethard, and Guergana Savova. 2023 · 2023
Later among the works it cites.
Jan Cegin, Branislav Pecher, Jakub Simko, Ivan Srba, Maria Bielikova, and Peter Brusilovsky. 2024 · 2024
Closest in time.
A survey on stability of learning with limited labelled data and its sensitivity to the effects of randomness
Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024 · 2024
Closest in time.