Fetching the paper…
Reading the bibliography…
Recent developments in natural language representations have been accompanied by large and expensive models that leverage vast amounts of general-domain text through self-supervised pre-training.
The proof and measurement of association between two things
Spearman · 1904
Earlier work this paper cites.
Model compression with multi-task knowledge distillation for web-scale question answering system
Ze Yang, Linjun Shou, Ming Gong, Wutao Lin, and Daxin Jiang · 1904
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le · 1906
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 1908
Earlier work this paper cites.
Frequency Analysis of English vocabulary and grammar, based on the LOB corpus
Johansson, Stig, and Knut Hofland · 1989
Earlier work this paper cites.
Measures for corpus similarity and homogeneity
Adam Kilgarriff and Tony Rose · 1998
Earlier work this paper cites.
Model compression
C Bucilă, R Caruana, and A Niculescu-Mizil · 2006
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini · 2009
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Learning small-size dnn with output-distribution-based criteria
Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Semi-supervised convolutional neural networks for text categorization via region embedding
Rie Johnson and Tong Zhang · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2018
Later among the works it cites.
Attention-guided answer distillation for machine reading comprehension
Minghao Hu, Yuxing Peng, Furu Wei, Zhen Huang, Dongsheng Li, Nan Yang, and Ming Zhou · 2018
Later among the works it cites.
Few-shot learning of neural networks from scratch by pseudo example optimization
Akisato Kimura, Zoubin Ghahramani, Koh Takeuchi, Tomoharu Iwata, and Naonori Ueda · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Song Han, Huizi Mao, and William J Dally · 2016
Cited alongside, same era.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Ruining He and Julian McAuley · 2016
Cited alongside, same era.
Adversarial adaptation of synthetic or stale data
Young-Bum Kim, Karl Stratos, and Dongchan Kim · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim · 2017
Cited alongside, same era.
Quora question pairs
Z. Chen, H. Zhang, X. Zhang, and L. Zhao · 2018
Cited alongside, same era.
Transformer to cnn: Label-scarce distillation for efficient text classification
Yew Ken Chia, Sam Witteveen, and Martin Andrews · 2018
Cited alongside, same era.
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Later among the works it cites.
BAM! born-again multi-task networks for natural language understanding
Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D. Manning, and Quoc V. Le · 2019
Closest in time.
Variational pretraining for semi-supervised text classification
Suchin Gururangan, Tam Dang, Dallas Card, and Noah A. Smith · 2019
Closest in time.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.
Smaller, faster, cheaper, lighter: Introducing DistilBERT, a distilled version of BERT
Victor Sanh · 2019
Closest in time.
Distilling task-specific knowledge from bert into simple neural networks, 2019
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin · 2019
Closest in time.