Fetching the paper…
Reading the bibliography…
State-of-the-art performance on language understanding tasks is now achieved with increasingly large networks; the current record holder has billions of parameters.
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. B. and Brockett, C. (2005) · 2005
Earlier work this paper cites.
The Second PASCAL Recognising Textual Entailment Challenge
Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I. (2006) · 2006
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Dagan, I., Glickman, O., and Magnini, B. (2006) · 2006
Earlier work this paper cites.
The Third PASCAL Recognizing Textual Entailment Challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B. (2007) · 2007
Earlier work this paper cites.
The Sixth PASCAL Recognizing Textual Entailment Challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D. (2009) · 2009
Earlier work this paper cites.
The Winograd Schema Challenge
Levesque, H. J., Davis, E., and Morgenstern, L. (2012) · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J. Y., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. (2013) · 2013
Earlier work this paper cites.
BinaryConnect: Training Deep Neural Networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P. (2015) · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Earlier work this paper cites.
BinaryNet: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
Courbariaux, M. and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Hendrycks, D. and Gimpel, K. (2016) · 2016
Earlier work this paper cites.
SQuad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016) · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L. (2018) · 2017
Cited alongside, same era.
Deep Learning Scaling is Predictable, Empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. (2017) · 2017
Cited alongside, same era.
Learning Sparse Neural Networks through $L_0$ Regularization
Louizos, C., Welling, M., and Kingma, D. P. (2017) · 2017
Cited alongside, same era.
First Quora Dataset Release: Question Pairs
Shankar Iyer, Dandekar, N., and Csernai, K. (2017) · 2017
Cited alongside, same era.
Attention Is All You Need
Vaswani, A., Uszkoreit, J., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Neural Network Acceptability Judgments
Warstadt, A., Singh, A., and Bowman, S. R. (2018) · 2018
Later among the works it cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A., Nangia, N., and Bowman, S. (2018) · 2018
Later among the works it cites.
Superposition of many models into one
Cheung, B., Terekhov, A., Chen, Y., Agrawal, P., and Olshausen, B. (2019) · 2019
Later among the works it cites.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019) · 2019
Later among the works it cites.
Multi-Task Deep Neural Networks for Natural Language Understanding
Liu, X., He, P., Chen, W., and Gao, J. (2019) · 2019
Later among the works it cites.
Adding new tasks to a single network with weight transformations using binary masks
Mancini, M., Ricci, E., Caputo, B., and Bulò, S. R. (2019) · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhu, M. and Gupta, S. (2017) · 2017
Cited alongside, same era.
Reconciling modern machine learning and the bias-variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S. (2018) · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
The Lottery Ticket Hypothesis: Finding Small, Trainable Neural Networks
Frankle, J. and Carbin, M. (2018) · 2018
Cited alongside, same era.
Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
Mallya, A., Davis, D., and Lazebnik, S. (2018) · 2018
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Later among the works it cites.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. (2019) · 2019
Later among the works it cites.
Are All Layers Created Equal?
Zhang, C., Bengio, S., and Singer, Y. (2019) · 2019
Later among the works it cites.
Extreme Language Model Compression with Optimal Subwords and Shared Projections
Zhao, S., Gupta, R., Song, Y., and Zhou, D. (2019) · 2019
Later among the works it cites.
Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J. (2019) · 2019
Later among the works it cites.