Linguistic knowledge and transferability of contextual representations
Original
Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew E Peters, and Noah A Smith. 2019a · 1903
Earlier work this paper cites.
To tune or not to tune? adapting pretrained representations to diverse tasks
Original
Matthew E Peters, Sebastian Ruder, and Noah A Smith. 2019 · 1903
Earlier work this paper cites.
Ernie: Enhanced representation through knowledge integration
Original
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019 · 1904
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Original
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 1905
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Original
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019b · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Designing and interpreting probes with control tasks
Original
John Hewitt and Percy Liang. 2019 · 1909
Earlier work this paper cites.
Asymptotically efficient estimation of covariance matrices with linear structure
Theodore W Anderson. 1973 · 1973
Earlier work this paper cites.
Compressing bert: Studying the effects of weight pruning on transfer learning
Original
Mitchell A Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2002
Earlier work this paper cites.
On the regularization of canonical correlation analysis
Tijl De Bie and Bart De Moor. 2003 · 2003
Earlier work this paper cites.
Information-theoretic probing with minimum description length
Original
Elena Voita and Ivan Titov. 2020 · 2003
Earlier work this paper cites.
What happens to bert embeddings during fine-tuning?
Original
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020 · 2004
Earlier work this paper cites.
Information-theoretic probing for linguistic structure
Original
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2004
Earlier work this paper cites.
On the effect of dropping layers of pre-trained transformer models
Original
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020 · 2004
Earlier work this paper cites.
Investigating transferability in pretrained language models
Original
Alex Tamkin, Trisha Singh, Davide Giovanardi, and Noah Goodman. 2020 · 2004
Earlier work this paper cites.
Deebert: Dynamic early exiting for accelerating bert inference
Original
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2004
Earlier work this paper cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Original
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2020 · 2005
Earlier work this paper cites.
DeBERTa: Decoding-enhanced bert with disentangled attention
Original
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2006
Earlier work this paper cites.
Support vector machines-kernels and the kernel trick
Martin Hofmann. 2006 · 2006
Earlier work this paper cites.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Original
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2020a · 2006
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman. 2009 · 2009
Earlier work this paper cites.
On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained transformers
Original
Marius Mosbach, Anna Khokhlova, Michael A Hedderich, and Dietrich Klakow. 2020b · 2010
Earlier work this paper cites.