Fetching the paper…
Reading the bibliography…
We consider the problem of producing compact architectures for text classification, such that the full model fits in a limited amount of memory.
Indexing by latent semantic analysis
Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harshman · 1990
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla · 1990
Earlier work this paper cites.
A threshold of ln n for approximating set cover
Uriel Feige · 1998
Earlier work this paper cites.
Text categorization with support vector machines: Learning with many relevant features
Thorsten Joachims · 1998
Earlier work this paper cites.
A comparison of event models for naive bayes text classification
Andrew McCallum and Kamal Nigam · 1998
Earlier work this paper cites.
Entropy-based pruning of backoff language models
Andreas Stolcke · 2000
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Moses S. Charikar · 2002
Earlier work this paper cites.
Locality-sensitive hashing scheme based on p-stable distributions
M. Datar, N. Immorlica, P. Indyk, and V.S. Mirrokni · 2004
Earlier work this paper cites.
Hamming embedding and weak geometric consistency for large scale image search
Hervé Jégou, Matthijs Douze, and Cordelia Schmid · 2008
Earlier work this paper cites.
The group lasso for logistic regression
Lukas Meier, Sara Van De Geer, and Peter Bühlmann · 2008
Earlier work this paper cites.
Opinion mining and sentiment analysis
Bo Pang and Lillian Lee · 2008
Earlier work this paper cites.
Randomized language models via perfect hash functions
David Talbot and Thorsten Brants · 2008
Earlier work this paper cites.
Feature hashing for large scale multitask learning
Kilian Q Weinberger, Anirban Dasgupta, John Langford, Alex Smola, and Josh Attenberg · 2009
Earlier work this paper cites.
Spectral hashing
Yair Weiss, Antonio Torralba, and Rob Fergus · 2009
Earlier work this paper cites.
Submodular secretary problem and extensions
Mohammad Hossein Bateni, Mohammad Taghi Hajiaghayi, and Morteza Zadimoghaddam · 2010
Earlier work this paper cites.
Max-cover in map-reduce
Flavio Chierichetti, Ravi Kumar, and Andrew Tomkins · 2010
Cited alongside, same era.
Iterative quantization: A procrustean approach to learning binary codes
Yunchao Gong and Svetlana Lazebnik · 2011
Cited alongside, same era.
Product quantization for nearest neighbor search
Hervé Jegou, Matthijs Douze, and Cordelia Schmid · 2011
Cited alongside, same era.
High-dimensional signature compression for large-scale image classification
Jorge Sánchez and Florent Perronnin · 2011
Cited alongside, same era.
Optimization with sparsity-inducing penalties
Francis Bach, Rodolphe Jenatton, Julien Mairal, and Guillaume Obozinski · 2012
Cited alongside, same era.
Statistical language models based on neural networks
Tomas Mikolov · 2012
Cited alongside, same era.
Asymmetric LSH for sublinear time maximum inner product search
Anshumali Shrivastava and Ping Li · 2014
Later among the works it cites.
Hashing for similarity search: A survey
Jingdong Wang, Heng Tao Shen, Jingkuan Song, and Jianqiu Ji · 2014
Later among the works it cites.
Strategies for training large vocabulary neural language models
Welin Chen, David Grangier, and Michael Auli · 2015
Later among the works it cites.
Neural networks with few multiplications
Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio · 2015
Later among the works it cites.
On symmetric and asymmetric lshs for inner product search
Behnam Neyshabur and Nathan Srebro · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Subword language modeling with neural networks
Tomas Mikolov, Ilya Sutskever, Anoop Deoras, Hai-Son Le, Stefan Kombrink, and J Cernocky · 2012
Cited alongside, same era.
Baselines and bigrams: Simple, good sentiment and topic classification
Sida Wang and Christopher D Manning · 2012
Cited alongside, same era.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Marc-Aurelio Ranzato, and Nando et all de Freitas · 2013
Cited alongside, same era.
Optimized product quantization for approximate nearest neighbor search
Tiezheng Ge, Kaiming He, Qifa Ke, and Jian Sun · 2013
Cited alongside, same era.
Cartesian k-means
Mohammad Norouzi and David Fleet · 2013
Cited alongside, same era.
A reliable effective terascale linear learning system
Alekh Agarwal, Olivier Chapelle, Miroslav Dudík, and John Langford · 2014
Cited alongside, same era.
Learning to hash for indexing big data - A survey
Jun Wang, Wei Liu, Sanjiv Kumar, and Shih-Fu Chang · 2015
Later among the works it cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Later among the works it cites.
Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Closest in time.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou · 2016
Closest in time.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Closest in time.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov · 2016
Closest in time.
How should we evaluate supervised hashing?
Alexandre Sablayrolles, Matthijs Douze, Hervé Jégou, and Nicolas Usunier · 2016
Closest in time.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Closest in time.
Efficient character-level document classification by combining convolution and recurrent layers
Yijun Xiao and Kyunghyun Cho · 2016
Closest in time.