Fetching the paper…
Reading the bibliography…
Adapting large-scale pretrained language models to downstream tasks via fine-tuning is the standard method for achieving state-of-the-art performance on NLP benchmarks.
Wordnet: a lexical database for english
George A Miller · 1995
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Verbnet: A broad-coverage, comprehensive verb lexicon
Karin Kipper Schuler · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, and Danilo Giampiccolo · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini · 2009
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Fastfood-approximating kernel expansions in loglinear time
Quoc Le, Tamás Sarlós, and Alex Smola · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Earlier work this paper cites.
Contextual parameter generation for universal neural machine translation
Emmanouil Antonios Platanios, Mrinmaya Sachan, Graham Neubig, and Tom Mitchell · 2018
Earlier work this paper cites.
Deep quaternion networks
Chase J Gaudet and Anthony S Maida · 2018
Cited alongside, same era.
Quaternion convolutional neural networks
Xuanyu Zhu, Yi Xu, Hongteng Xu, and Changjian Chen · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Cited alongside, same era.
Record: Bridging the gap between human and machine commonsense reading comprehension
Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme · 2018
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
BatchEnsemble: An Alternative Approach to Efficient Ensemble and Lifelong Learning
Yeming Wen, Dustin Tran, and Jimmy Ba · 2020
Later among the works it cites.
Tinytl: Reduce memory, not parameters for efficient on-device learning
Han Cai, Chuang Gan, Ligeng Zhu, and Song Han · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Udapter: Language adaptation for truly universal dependency parsing
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord · 2020
Later among the works it cites.
The lottery ticket hypothesis for pre-trained bert networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E Peters, Sebastian Ruder, and Noah A Smith · 2019
Cited alongside, same era.
Lightweight and efficient neural natural language processing with quaternion networks
Yi Tay, Aston Zhang, Anh Tuan Luu, Jinfeng Rao, Shuai Zhang, Shuohang Wang, Jie Fu, and Siu Cheung Hui · 2019
Cited alongside, same era.
Evaluating lottery tickets under distributional shifts
Shrey Desai, Hongyuan Zhan, and Ahmed Aly · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman · 2019
Cited alongside, same era.
Sai Prasanna, Anna Rogers, and Anna Rumshisky · 2020
Later among the works it cites.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag · 2020
Later among the works it cites.
True Few-Shot Learning with Language Models
Ethan Perez, Douwe Kiela, and Kyunghyun Cho · 2021
Closest in time.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Closest in time.
Warp: Word-level adversarial reprogramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May · 2021
Closest in time.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Closest in time.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta · 2021
Closest in time.
Rethinking Embedding Coupling in Pre-trained Language Models
Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson, and Sebastian Ruder · 2021
Closest in time.
AdapterDrop: On the Efficiency of Adapters in Transformers
Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers, and Iryna Gurevych · 2021
Closest in time.
Parameterized hypercomplex graph neural networks for graph classification
Tuan Le, Marco Bertolini, Frank Noé, and Djork-Arné Clevert · 2021
Closest in time.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson · 2021
Closest in time.
AdapterFusion: Non-destructive task composition for transfer learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rückĺe, Cho Kyunghyun, and Iryna Gurevych · 2021
Closest in time.
High-performance large-scale image recognition without normalization
Andrew Brock, Soham De, Samuel L Smith, and Karen Simonyan · 2021
Closest in time.
Conditionally adaptive multi-task learning: Improving transfer learning in NLP using fewer parameters & less data
Jonathan Pilault, Amine El hattami, and Christopher Pal · 2021
Closest in time.