Fetching the paper…
Reading the bibliography…
It is widely accepted that fine-tuning pre-trained language models usually brings about performance improvements in downstream tasks.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Probing linguistic systematicity
Emily Goodwin, Koustuv Sinha, and Timothy J. O’Donnell. 2020 · 1969
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
A latent variable model approach to PMI-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018 · 2018
Earlier work this paper cites.
All-but-the-top: Simple and effective postprocessing for word representations
Jiaqi Mu and Pramod Viswanath. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Earlier work this paper cites.
Representation degeneration problem in training natural language generation models
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tieyan Liu. 2019 · 2019
Earlier work this paper cites.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them
Hila Gonen and Yoav Goldberg. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a · 2019
Cited alongside, same era.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
How fine can fine-tuning be? learning efficient language models
Evani Radiya-Dixit and Xin Wang. 2020 · 2020
Later among the works it cites.
oLMpics-on what language model pre-training captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020 · 2020
Later among the works it cites.
Improving neural language generation with spectrum control
Lingxiao Wang, Jing Huang, Kevin Huang, Ziniu Hu, Guangtao Wang, and Quanquan Gu. 2020 · 2020
Later among the works it cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020 · 2020
Later among the works it cites.
Revisiting representation degeneration problem in language modeling
Zhong Zhang, Chongming Gao, Cong Xu, Rui Miao, Qinli Yang, and Junming Shao. 2020 · 2020
Later among the works it cites.
Isotropy in the contextual embedding space: Clusters and manifolds
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019b · 2019
Cited alongside, same era.
Investigating learning dynamics of BERT fine-tuning
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2020 · 2020
Cited alongside, same era.
On the sentence embeddings from pre-trained language models
Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020 · 2020
Cited alongside, same era.
What happens to BERT embeddings during fine-tuning?
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020 · 2020
Cited alongside, same era.
Asking without telling: Exploring latent ontologies in contextual representations
Julian Michael, Jan A. Botha, and Ian Tenney. 2020 · 2020
Cited alongside, same era.
On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained transformers
Marius Mosbach, Anna Khokhlova, Michael A. Hedderich, and Dietrich Klakow. 2020 · 2020
Cited alongside, same era.
Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church. 2021 · 2021
Closest in time.
Probing BERT in hyperbolic spaces
Boli Chen, Yao Fu, Guangwei Xu, Pengjun Xie, Chuanqi Tan, Mosha Chen, and Liping Jing. 2021 · 2021
Closest in time.
How transfer learning impacts linguistic knowledge in deep NLP models?
Nadir Durrani, Hassan Sajjad, and Fahim Dalvi. 2021 · 2021
Closest in time.
A cluster-based approach for improving isotropy in contextual embedding space
Sara Rajaee and Mohammad Taher Pilehvar. 2021 · 2021
Closest in time.
Infusing Finetuning with Semantic Dependencies
Zhaofeng Wu, Hao Peng, and Noah A. Smith. 2021 · 2021
Closest in time.
On the interplay between fine-tuning and composition in transformers
Lang Yu and Allyson Ettinger. 2021 · 2021
Closest in time.
DirectProbe: Studying representations without classifiers
Yichu Zhou and Vivek Srikumar. 2021 · 2021
Closest in time.