Fetching the paper…
Reading the bibliography…
Fine-tuning large language models (LLMs) is computationally intensive because it requires updating all parameters.
Parameter-efficient transfer learning for nlp, 2019
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 1902
Earlier work this paper cites.
Intrinsic dimension of data representations in deep neural networks, 2019
Alessio Ansuini, Alessandro Laio, Jakob H. Macke, and Davide Zoccolan · 1905
Earlier work this paper cites.
A geometric modeling of occam’s razor in deep learning, 2024
Ke Sun and Frank Nielsen · 1905
Earlier work this paper cites.
Optuna: A next-generation hyperparameter optimization framework, 2019
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama · 1907
Earlier work this paper cites.
On the mathematical foundations of theoretical statistics
R A Fisher · 1922
Earlier work this paper cites.
Statistical decision rules and optimal inference , volume 53 of Translations of Mathematical Monographs
N. N. Čencov · 1982
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia · 2001
Earlier work this paper cites.
Analyzing redundancy in pretrained transformer models, 2020
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov · 2004
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Adapterfusion: Non-destructive task composition for transfer learning, 2021
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych · 2005
Earlier work this paper cites.
The rank of a random matrix
Xinlong Feng and Zhinan Zhang · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, and Danilo Giampiccolo · 2006
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention, 2021
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2006
Earlier work this paper cites.
Hierarchical nucleation in deep neural networks, 2020
Diego Doimo, Aldo Glielmo, Alessio Ansuini, and Alessandro Laio · 2007
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan · 2007
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
Sumio Watanabe · 2009
Earlier work this paper cites.
Differential Topology
V. Guillemin and A. Pollack · 2010
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text, 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms, 2017
Sebastian Ruder · 2017
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes, 2018
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Cited alongside, same era.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models, 2022
Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg · 2022
Later among the works it cites.
The generalized ratios intrinsic dimension estimator
Francesco Denti, Diego Doimo, Alessandro Laio, and Antonietta Mira · 2022
Later among the works it cites.
The geometry of hidden representations of large transformer models, 2023
Lucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio, Alessio Ansuini, and Alberto Cazzaniga · 2023
Later among the works it cites.
Bridging information-theoretic and geometric compression in language models
Emily Cheng, Corentin Kervadec, and Marco Baroni · 2023
Later among the works it cites.
Henry Kvinge, Davis Brown, and Charles Godfrey · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Three factors influencing minima in sgd, 2018
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Cited alongside, same era.
Computing the free energy without collective variables
Alex Rodriguez, Maria d’Errico, Elena Facco, and Alessandro Laio · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning, 2020
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta · 2020
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms, 2023
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Later among the works it cites.
Adalora: Adaptive budget allocation for parameter-efficient fine-tuning, 2023
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Nikos Karampatziakis, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao · 2023
Later among the works it cites.
Structure-aware low-rank adaptation for parameter-efficient fine-tuning
Yahao Hu, Yifei Xie, Tianfeng Wang, Man Chen, and Zhisong Pan · 2023
Later among the works it cites.
Sparse low-rank adaptation of pre-trained language models, 2023
Ning Ding, Xingtai Lv, Qiaosen Wang, Yulin Chen, Bowen Zhou, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee · 2023
Later among the works it cites.
Parameter-efficient fine-tuning for large models: A comprehensive survey, 2024
Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang · 2024
Closest in time.
Investigating adversarial vulnerability and implicit bias through frequency analysis, 2024
Lorenzo Basile, Nikos Karantzas, Alberto D’Onofrio, Luca Bortolussi, Alex Rodriguez, and Fabio Anselmi · 2024
Closest in time.
When do prompting and prefix-tuning work? a theory of capabilities and limitations, 2024
Aleksandar Petrov, Philip H. S. Torr, and Adel Bibi · 2024
Closest in time.
Lora+: Efficient low rank adaptation of large models, 2024
Soufiane Hayou, Nikhil Ghosh, and Bin Yu · 2024
Closest in time.
Alora: Allocating low-rank adaptation for fine-tuning large language models, 2024
Zequan Liu, Jiawen Lyn, Wei Zhu, Xing Tian, and Yvette Graham · 2024
Closest in time.
airoboros-gpt4-1.4.1-mpt
Jon Durbin · 2024
Closest in time.
Estimating the intrinsic dimension of datasets by a minimal neighborhood information
Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio · 2045
Closest in time.