Fetching the paper…
Reading the bibliography…
In the natural language processing literature, neural networks are becoming increasingly deeper and complex.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla. 1990 · 1990
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky. 2009 · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Learning continuous phrase representations and syntactic parsing with recursive neural networks
Richard Socher, Christopher D. Manning, and Andrew Y. Ng. 2010 · 2010
Earlier work this paper cites.
Extensions of recurrent neural network language model
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2011 · 2011
Earlier work this paper cites.
Parsing natural scenes and natural language with recursive neural networks
Richard Socher, Cliff C. Lin, Chris Manning, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Matthew D. Zeiler. 2012 · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Very deep convolutional networks for text classification
Alexis Conneau, Holger Schwenk, Loïc Barrault, and Yann Lecun. 2016 · 2016
Cited alongside, same era.
Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Pairwise word interaction modeling with deep neural networks for semantic similarity measurement
Hua He and Jimmy Lin. 2016 · 2016
Efficient methods and hardware for deep learning
Song Han. 2017 · 2017
Later among the works it cites.
Learning efficient convolutional networks through network slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017 · 2017
Later among the works it cites.
First Quora dataset release: Question pairs
Nikhil Dandekar Shankar Iyer and Kornél Csernai. 2017 · 2017
Later among the works it cites.
Bilateral multi-perspective matching for natural language sentences
Zhiguo Wang, Wael Hamza, and Radu Florian. 2017 · 2017
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2017 · 2017
Later among the works it cites.
Constraint-aware deep neural network compression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
UMD-TTIC-UW at SemEval-2016 task 1: Attention-based multi-perspective convolutional neural networks for textual similarity measurement
Hua He, John Wieting, Kevin Gimpel, Jinfeng Rao, and Jimmy Lin. 2016 · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Cited alongside, same era.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2016 · 2016
Cited alongside, same era.
Generating factoid questions with recurrent neural networks: The 30m factoid question-answer corpus
Iulian Vlad Serban, Alberto García-Durán, Caglar Gulcehre, Sungjin Ahn, Sarath Chandar, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
A deep architecture for semantic matching with multiple positional sentence representations
Shengxian Wan, Yanyan Lan, Jiafeng Guo, Jun Xu, Liang Pang, and Xueqi Cheng. 2016 · 2016
Cited alongside, same era.
The galactic dependencies treebanks: Getting more data by synthesizing new languages
Dingquan Wang and Jason Eisner. 2016 · 2016
Cited alongside, same era.
Text classification improved by integrating bidirectional LSTM with two-dimensional max pooling
Peng Zhou, Zhenyu Qi, Suncong Zheng, Jiaming Xu, Hongyun Bao, and Bo Xu. 2016 · 2016
Cited alongside, same era.
Changan Chen, Frederick Tung, Naveen Vedula, and Greg Mori. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Time Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
FLOPs as a direct optimization objective for learning sparse neural networks
Raphael Tang, Ashutosh Adhikari, and Jimmy Lin. 2018 · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Training and inference with integers in deep neural networks
Shuang Wu, Guoqi Li, Feng Chen, and Luping Shi. 2018 · 2018
Later among the works it cites.
On-device neural language model based word prediction
Seunghak Yu, Nilesh Kulkarni, Haejun Lee, and Jihie Kim. 2018 · 2018
Later among the works it cites.