Fetching the paper…
Reading the bibliography…
Self-supervised pre-training of large-scale transformer models on text corpora followed by finetuning has achieved state-of-the-art on a number of natural language processing tasks.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2012
Earlier work this paper cites.
Scope: Structural classification of proteins–extended, integrating scop and astral data and classification of new structures
Fox, N. K., Brenner, S. E., and Chandonia, J.-M · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, u., and Polosukhin, I · 2017
Cited alongside, same era.
Deepsf: Deep convolutional neural network for mapping protein sequences to folds
Hou, J., Adhikari, B., and Cheng, J · 2018
Cited alongside, same era.
On the state of the art of evaluation in neural language models
Melis, G., Dyer, C., and Blunsom, P · 2018
Cited alongside, same era.
Differentiable plasticity: training plastic neural networks with backpropagation
Miconi, T., Stanley, K., and Clune, J · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Evaluating protein transfer learning with tape
Rao, R., Bhattacharya, N., Thomas, N., Duan, Y., Chen, P., Canny, J., Abbeel, P., and Song, Y · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
On empirical comparisons of optimizers for deep learning
Choi, D., Shallue, C. J., Nado, Z., Lee, J., Maddison, C. J., and Dahl, G. E · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nangia, N. and Bowman, S · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A · 2019
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A · 2020
Later among the works it cites.
Pretrained transformers as universal computation engines
Lu, K., A. Grover, A., Abbeel, P., and Mordatch, I · 2021
Closest in time.
KILT: a benchmark for knowledge intensive language tasks
Petroni, F., Piktus, A., Fan, A., Lewis, P., Yazdani, M., De Cao, N., Thorne, J., Jernite, Y., Karpukhin, V., Maillard, J., Plachouras, V., Rocktäschel, T., and Riedel, S · 2021
Closest in time.
Long range arena : A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Closest in time.