Fetching the paper…
Reading the bibliography…
Transfer learning has fundamentally changed the landscape of natural language processing (NLP) research.
Liu, X · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y · 1907
Earlier work this paper cites.
Structbert: Incorporating language structures into pre-training for deep language understanding
Wang, W · 1908
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z · 1909
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Nie, Y · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C · 1910
Earlier work this paper cites.
The influence curve and its role in robust estimation
Hampel, F. R · 1974
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
Rockafellar, R. T · 1976
Earlier work this paper cites.
On the convergence of the proximal point algorithm for convex minimization
Güler, O · 1991
Earlier work this paper cites.
New proximal point algorithms for convex minimization
Güler, O · 1992
Earlier work this paper cites.
Nonlinear proximal point algorithms using bregman functions, with applications to convex programming
Eckstein, J · 1993
Earlier work this paper cites.
Virtual adversarial training: a regularization method for supervised and semi-supervised learning
Miyato, T · 1993
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Convergence of proximal-like algorithms
Teboulle, M · 1997
Earlier work this paper cites.
Trust region methods
Conn, A. R · 2000
Earlier work this paper cites.
The microsoft toolkit of multi-task deep neural networks for natural language understanding
Liu, X · 2002
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Bar-Haim, R · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Giampiccolo, D · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L · 2009
Cited alongside, same era.
A survey on transfer learning
Pan, S. J · 2009
Cited alongside, same era.
Robust statistics
Huber, P. J · 2011
Cited alongside, same era.
The winograd schema challenge
Levesque, H · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D · 2014
Cited alongside, same era.
Proximal algorithms
A dirt-t approach to unsupervised domain adaptation
Shu, R · 2018
Later among the works it cites.
Fever: a large-scale dataset for fact extraction and verification
Thorne, J · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A · 2018
Later among the works it cites.
I know what you want: Semantic learning for text comprehension
Zhang, Z · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parikh, N · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Bowman, S · 2015
Cited alongside, same era.
Representation learning using multi-task deep neural networks for semantic classification and information retrieval
Liu, X · 2015
Cited alongside, same era.
The information geometry of mirror descent
Raskutti, G · 2015
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D · 2017
Cited alongside, same era.
Devlin, J · 2019
Closest in time.
Unified language model pre-training for natural language understanding and generation 13042–13054
Dong, L · 2019
Closest in time.
A hybrid neural network model for commonsense reasoning
He, P · 2019
Closest in time.
Parameter-efficient transfer learning for nlp
Houlsby, N · 2019
Closest in time.
A surprisingly robust trick for the winograd schema challenge
Kocijan, V · 2019
Closest in time.
To tune or not to tune? adapting pretrained representations to diverse tasks
Peters, M. E · 2019
Closest in time.
Language models are unsupervised multitask learners
Radford, A · 2019
Closest in time.
Bert and pals: Projected attention layers for efficient adaptation in multi-task learning
Stickland, A. C · 2019
Closest in time.
Neural network acceptability judgments
Warstadt, A · 2019
Closest in time.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T · 2019
Closest in time.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z · 2019
Closest in time.
Theoretically principled trade-off between robustness and accuracy
Zhang, H · 2019
Closest in time.
Spanbert: Improving pre-training by representing and predicting spans
Joshi, M · 2020
Closest in time.
On the variance of the adaptive learning rate and beyond
Liu, L · 2020
Closest in time.
Freelb: Enhanced adversarial training for natural language understanding. https://openreview.net/forum?id=BygzbyHFvB
Zhu, C · 2020
Closest in time.