Fetching the paper…
Reading the bibliography…
Pre-trained language models (e.g., BERT (Devlin et al., 2018) and its variants) have achieved remarkable success in varieties of NLP tasks.
Improved knowledge distillation via teacher assistant: Bridging the gap between student and teacher
Mirzadeh, S., Farajtabar, M., Li, A., and Ghasemzadeh, H · 1902
Earlier work this paper cites.
Distilling task-specific knowledge from BERT into simple neural networks
Tang, R., Lu, Y., Liu, L., Mou, L., Vechtomova, O., and Lin, J · 1903
Earlier work this paper cites.
What does BERT look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 1906
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Turc, I., Chang, M., Lee, K., and Toutanova, K · 1908
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Turc, I., Chang, M., Lee, K., and Toutanova, K · 1908
Earlier work this paper cites.
Tinybert: Distilling BERT for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 1909
Earlier work this paper cites.
Knowledge distillation from internal representations
Aguilar, G., Ling, Y., Zhang, Y., Yao, B., Fan, X., and Guo, E · 1910
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 1910
Earlier work this paper cites.
MLQA: evaluating cross-lingual extractive question answering
Lewis, P. S. H., Oguz, B., Rinott, R., Riedel, S., and Schwenk, H · 1910
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V · 1911
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., and Giampiccolo, D · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Dagan, I., Dang, H. T., Giampiccolo, D., and Magnini, B · 2009
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Cited alongside, same era.
Ba, L. J., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P · 2018
Later among the works it cites.
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V · 2018
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Cited alongside, same era.
Get to the point: Summarization with pointer-generator networks
See, A., Liu, P. J., and Manning, C. D · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Zagoruyko, S. and Komodakis, N · 2017
Cited alongside, same era.
Paragraph-level neural question generation with maxout pointer and gated self-attention networks
Zhao, Y., Ni, X., Ding, Y., and Ke, Q · 2018
Later among the works it cites.
Cloze-driven pretraining of self-attention networks
Baevski, A., Edunov, S., Liu, Y., Zettlemoyer, L., and Auli, M · 2019
Later among the works it cites.
Unified language model pre-training for natural language understanding and generation
Dong, L., Yang, N., Wang, W., Wei, F., Liu, X., Wang, Y., Gao, J., Zhou, M., and Hon, H.-W · 2019
Later among the works it cites.
What does BERT learn about the structure of language?
Jawahar, G., Sagot, B., and Seddah, D · 2019
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., and Levy, O · 2019
Later among the works it cites.
Cross-lingual language model pretraining
Lample, G. and Conneau, A · 2019
Later among the works it cites.
Text summarization with pretrained encoders
Liu, Y. and Lapata, M · 2019
Later among the works it cites.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Mass: Masked sequence to sequence pre-training for language generation
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y · 2019
Later among the works it cites.
Patient knowledge distillation for BERT model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Later among the works it cites.
Small and practical BERT models for sequence labeling
Tsai, H., Riesa, J., Johnson, M., Arivazhagan, N., Li, X., and Archer, A · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V · 2019
Later among the works it cites.
Addressing semantic drift in question generation for semi-supervised question answering
Zhang, S. and Bansal, M · 2019
Later among the works it cites.
Unilmv2: Pseudo-masked language models for unified language model pre-training
Bao, H., Dong, L., Wei, F., Wang, W., Yang, N., Liu, X., Wang, Y., Piao, S., Gao, J., Zhou, M., and Hon, H.-W · 2020
Closest in time.