Fetching the paper…
Reading the bibliography…
Recent developments in unsupervised representation learning have successfully established the concept of transfer learning in NLP.
Dual co-matching network for multi-choice reading comprehension
Zhang, S., Zhao, H., Wu, Y., Zhang, Z., Zhou, X., and Zhou, X. (2019) · 1901
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2019) · 1905
Earlier work this paper cites.
Defending against neural fake news
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y. (2019) · 1905
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
Strubell, E., Ganesh, A., and McCallum, A. (2019) · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V. (2019) · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2019) · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2019) · 1910
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
The trec-8 question answering track evaluation
Voorhees, E. M. and Tice, D. M. (1999) · 1999
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C. (2005) · 2005
Earlier work this paper cites.
Clueweb09 data set
Callan, J., Hoy, M., Yoo, C., and Zhao, L. (2009) · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011) · 2011
Earlier work this paper cites.
English gigaword fifth edition, june
Parker, R., Graff, D., Kong, J., Chen, K., and Maeda, K. (2011) · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T. (2013) · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. (2013) · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., Le Cun, Y., and Fergus, R. (2013) · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Deep pyramid convolutional neural networks for text categorization
Johnson, R. and Zhang, T. (2017) · 2017
Later among the works it cites.
Race: Large-scale reading comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E. (2017) · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R. (2017) · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. (2014) · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X. (2015) · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A. (2015) · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y. (2015) · 2015
Cited alongside, same era.
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Cited alongside, same era.
Convolutional neural networks for text categorization: Shallow word-level vs. deep character-level
Johnson, R. and Zhang, T. (2016) · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2016) · 2016
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Later among the works it cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S. (2018) · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018) · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018) · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P. (2018) · 2018
Later among the works it cites.
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V. (2018) · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2018) · 2018
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R. (2018) · 2018
Later among the works it cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V. (2019) · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Later among the works it cites.
Transfer learning in natural language processing
Ruder, S., Peters, M. E., Swayamdipta, S., and Wolf, T. (2019) · 2019
Later among the works it cites.