Fetching the paper…
Reading the bibliography…
Fine-tuning pretrained contextual word embedding models to supervised downstream tasks has become commonplace in natural language processing.
Statistical methods for research workers
Fisher, R. A · 1935
Earlier work this paper cites.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Matthews, B. W · 1975
Earlier work this paper cites.
A sequential algorithm for training text classifiers
Lewis, D. D. and Gale, W. A · 1994
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, B. and Brockett, C · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J., Bardenet, R., Bengio, Y., and Kegl, B · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Non-stochastic best arm identification and hyperparameter optimization
Jamieson, K. and Talwalkar, A · 2016
Cited alongside, same era.
Determinantal point processes for mini-batch diversification
Zhang, C., Kjellström, H., and Mandt, S · 2017
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A · 2018
Cited alongside, same era.
On the state of the art of evaluation in neural language models
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffe, C., Shazeer, N., Roberts, A., Lee, K. L., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Inducing brain-relevant bias in natural language processing models
Schwartz, D., Toneva, M., and Wehbe, L · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Melis, G., Dyer, C., and Blunsom, P · 2018
Cited alongside, same era.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Phang, J., Févry, T., and Bowman, S. R · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Show your work: Improved reporting of experimental results
Dodge, J., Gururangan, S., Card, D., Schwartz, R., and Smith, N. A · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V · 2019
Later among the works it cites.
Freelb: Enhanced adversarial training for language understanding
Zhu, C., Cheng, Y., Gan, Z., Sun, S., Goldstein, T., and Liu, J · 2019
Later among the works it cites.