Fetching the paper…
Reading the bibliography…
Standard fine-tuning of language models typically performs well on in-distribution data, but suffers with generalization to distribution shifts.
Theory of statistical estimation
Rory A. Fisher. 1925 · 1925
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen. 1989 · 1989
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari. 1998 · 1998
Earlier work this paper cites.
Methods of information geometry
Amari Shun-ichi and Nagaoka Hiroshi. 2000 · 2000
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin A. Bejan, and Andrew S. Gordon. 2011 · 2011
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Critical learning periods in deep networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto. 2019 · 2019
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Ankur Bapna and Orhan Firat. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Aditya Golatkar, Alessandro Achille, and Stefano Soatto. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Limitations of the empirical fisher approximation for natural gradient descent
Frederik Kunstner, Philipp Hennig, and Lukas Balles. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
BERT and PALs: Projected attention layers for efficient adaptation in multi-task learning
Asa Cooper Stickland and Iain Murray. 2019 · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
MAD-G: Multilingual adapter generation for efficient cross-lingual transfer
Alan Ansell, Edoardo Maria Ponti, Jonas Pfeiffer, Sebastian Ruder, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2021 · 2021
Later among the works it cites.
DeBERTa: decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Stanislaw Jastrzebski, Devansh Arpit, Oliver Astrand, Giancarlo B Kerg, Huan Wang, Caiming Xiong, Richard Socher, Kyunghyun Cho, and Krzysztof J Geras. 2021 · 2021
Later among the works it cites.
UNKs everywhere: Adapting multilingual language models to new scripts
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021 · 2021
Later among the works it cites.
Training neural networks with fixed sparse masks
Yi-Lin Sung, Varun Nair, and Colin Raffel. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
MLQA: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens. 2020 · 2020
Cited alongside, same era.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Cited alongside, same era.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
Neural unsupervised domain adaptation in NLP—A survey
Alan Ramponi and Barbara Plank. 2020 · 2020
Cited alongside, same era.
Raise a child in large language model: Towards effective and generalizable fine-tuning
Runxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan, Baobao Chang, Songfang Huang, and Fei Huang. 2021 · 2021
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2022 · 2022
Later among the works it cites.
BAD-X: Bilingual adapters improve zero-shot cross-lingual transfer
Marinela Parović, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2022 · 2022
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2022 · 2022
Later among the works it cites.
Parameter-efficient tuning makes a good classification head
Zhuoyi Yang, Ming Ding, Yanhui Guo, Qingsong Lv, and Jie Tang. 2022 · 2022
Later among the works it cites.
Critical learning periods for multisensory integration in deep networks
Michael Kleinman, Alessandro Achille, and Stefano Soatto. 2023 · 2023
Closest in time.
Surgical fine-tuning improves adaptation to distribution shifts
Yoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar, Huaxiu Yao, Percy Liang, and Chelsea Finn. 2023 · 2023
Closest in time.
On surgical fine-tuning for language encoders
Abhilasha Lodha, Gayatri Belapurkar, Saloni Chalkapurkar, Yuanming Tao, Reshmi Ghosh, Samyadeep Basu, Dmitrii Petrov, and Soundararajan Srinivasan. 2023 · 2023
Closest in time.