Fetching the paper…
Reading the bibliography…
We present, to our knowledge, the first application of BERT to document classification.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
XLNet: generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla. 1990 · 1990
Earlier work this paper cites.
Automated learning of decision rules for text categorization
Chidanand Apté, Fred Damerau, and Sholom M. Weiss. 1994 · 1994
Earlier work this paper cites.
Scikit-learn: machine learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011 · 2011
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
GloVe: global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Event extraction via dynamic multi-pooling convolutional neural networks
Yubo Chen, Liheng Xu, Kang Liu, Daojian Zeng, and Jun Zhao. 2015 · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Cited alongside, same era.
Effective use of word order for text categorization with convolutional neural networks
Rie Johnson and Tong Zhang. 2015 · 2015
Cited alongside, same era.
Document modeling with gated recurrent neural network for sentiment classification
Duyu Tang, Bing Qin, and Ting Liu. 2015 · 2015
Cited alongside, same era.
Hierarchical attention networks for document classification
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016 · 2016
Learning sparse neural networks through l 0 l_{0} regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
SGM: sequence generation model for multi-label classification
Pengcheng Yang, Xu Sun, Wei Li, Shuming Ma, Wei Wu, and Houfeng Wang. 2018 · 2018
Later among the works it cites.
Rethinking complex neural network architectures for document classification
Ashutosh Adhikari, Achyudh Ram, Raphael Tang, and Jimmy Lin. 2019 · 2019
Closest in time.
BERT: pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning efficient convolutional networks through network slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017b · 2017
Cited alongside, same era.
Deep learning for extreme multi-label text classification
Jingzhou Liu, Wei-Cheng Chang, Yuexin Wu, and Yiming Yang. 2017a
Cited in the paper.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.