Fetching the paper…
Reading the bibliography…
Pre-trained models are widely used in fine-tuning downstream tasks with linear classifiers optimized by the cross-entropy loss, which might face robustness and stability problems.
Cross-entropy loss and low-rank features have responsibility for adversarial examples
Nar, K.; Ocal, O.; Sastry, S. S.; and Ramchandran, K. 2019 · 1901
Earlier work this paper cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Cao, K.; Wei, C.; Gaidon, A.; Arechiga, N.; and Ma, T. 2019 · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R.; and Le, Q. V. 2019 · 1906
Earlier work this paper cites.
Is BERT Really Robust? Natural Language Attack on Text Classification and Entailment
Jin, D.; Jin, Z.; Zhou, J. T.; and Szolovits, P. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
On mutual information maximization for representation learning
Tschannen, M.; Djolonga, J.; Rubenstein, P. K.; Gelly, S.; and Lucic, M. 2019 · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2019 · 1909
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U.; Levy, O.; Jurafsky, D.; Zettlemoyer, L.; and Lewis, M. 2019 · 1911
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E.; Hinton, G. E.; and Williams, R. J. 1986 · 1986
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J.; Ilharco, G.; Schwartz, R.; Farhadi, A.; Hajishirzi, H.; and Smith, N. 2020 · 2002
Earlier work this paper cites.
Supervised contrastive learning
Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020 · 2004
Earlier work this paper cites.
Bert-attack: Adversarial attack against bert using bert
Li, L.; Ma, R.; Guo, Q.; Xue, X.; and Qiu, X. 2020 · 2004
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Dagan, I.; Glickman, O.; and Magnini, B. 2005 · 2005
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
BERT-kNN: Adding a kNN search component to pretrained language models for better QA
Kassner, N.; and Schütze, H. 2020 · 2005
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A.; Zhou, H.; Mohamed, A.; and Auli, M. 2020 · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R.; Chopra, S.; and LeCun, Y. 2006 · 2006
Earlier work this paper cites.
Revisiting few-sample BERT fine-tuning
Zhang, T.; Wu, F.; Katiyar, A.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2006
Cited alongside, same era.
Distance metric learning for large margin nearest neighbor classification
Weinberger, K. Q.; and Saul, L. K. 2009 · 2009
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M.; and Hyvärinen, A. 2010 · 2010
Cited alongside, same era.
Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning
Gunel, B.; Du, J.; Conneau, A.; and Stoyanov, V. 2020 · 2011
Cited alongside, same era.
Learning word vectors for sentiment analysis
Maas, A.; Daly, R. E.; Pham, P. T.; Huang, D.; Ng, A. Y.; and Potts, C. 2011 · 2011
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y. 2018 · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018 · 2018
Later among the works it cites.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Papernot, N.; and McDaniel, P. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning word embeddings efficiently with noise-contrastive estimation
Mnih, A.; and Kavukcuoglu, K. 2013 · 2013
Cited alongside, same era.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A.; and Potts, C. 2013 · 2013
Cited alongside, same era.
Explaining and harnessing adversarial examples
Goodfellow, I. J.; Shlens, J.; and Szegedy, C. 2014 · 2014
Cited alongside, same era.
Facenet: A unified embedding for face recognition and clustering
Schroff, F.; Kalenichenko, D.; and Philbin, J. 2015 · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Zhang, X.; Zhao, J.; and LeCun, Y. 2015 · 2015
Cited alongside, same era.
Towards Evaluating the Robustness of Neural Networks
Carlini, N.; and Wagner, D. A. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P.; Lynch, C.; Chebotar, Y.; Hsu, J.; Jang, E.; Schaal, S.; Levine, S.; and Brain, G. 2018 · 2018
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. 2018 · 2018
Later among the works it cites.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z.; Xiong, Y.; Yu, S. X.; and Lin, D. 2018 · 2018
Later among the works it cites.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhang, Z.; and Sabuncu, M. R. 2018 · 2018
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Later among the works it cites.
Data-efficient image recognition with contrastive predictive coding
Henaff, O. 2020 · 2020
Later among the works it cites.
TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
Morris, J.; Lifland, E.; Yoo, J. Y.; Grigsby, J.; Jin, D.; and Qi, Y. 2020 · 2020
Later among the works it cites.
Training convolutional networks with noisy labels
Sukhbaatar, S.; Bruna, J.; Paluri, M.; Bourdev, L.; and Fergus, R. 2014 · 2080
Closest in time.