Fetching the paper…
Reading the bibliography…
Learning continuous representations of discrete objects such as text, users, movies, and URLs lies at the heart of many applications including language and user modeling.
Compression of recurrent neural networks for efficient language modeling
Artem M. Grachev, Dmitry I. Ignatov, and Andrey V. Savchenko · 1902
Earlier work this paper cites.
Neural input search for large scale recommendation models
Manas R. Joglekar, Cong Li, Jay K. Adams, Pranav Khaitan, and Quoc V. Le · 1907
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 1936
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L.M. Bregman · 1967
Earlier work this paper cites.
Texture features for image classification
R. Haralick, K. Shanmugam, and I. Dinstein · 1973
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Leon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Em algorithms for pca and spca
Sam Roweis · 1998
Earlier work this paper cites.
Using tf-idf to determine word relevance in document queries, 1999
Juan Ramos · 1999
Earlier work this paper cites.
On coresets for k-means and k-median clustering
Sariel Har-Peled and Soham Mazumdar · 2004
Earlier work this paper cites.
Conceptnet — a practical commonsense reasoning tool-kit
H. Liu and P. Singh · 2004
Earlier work this paper cites.
Clustering with bregman divergences
Arindam Banerjee, Srujana Merugu, Inderjit S. Dhillon, and Joydeep Ghosh · 2005
Earlier work this paper cites.
Infinite latent feature models and the indian buffet process
Thomas L. Griffiths and Zoubin Ghahramani · 2005
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2005
Earlier work this paper cites.
K-svd: An algorithm for designing overcomplete dictionaries for sparse representation
M. Aharon, M. Elad, and A. Bruckstein · 2006
Earlier work this paper cites.
K-means++: the advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2007
Earlier work this paper cites.
P.: Bayesian nonparametric latent feature models
Z. Ghahramani, P. Sollich, and T. L. Griffiths · 2007
Earlier work this paper cites.
Collapsed variational inference for hdp
Yee Whye Teh, Kenichi Kurihara, and Max Welling · 2007
Earlier work this paper cites.
Adaptive importance sampling to accelerate training of a neural probabilistic language model
Y. Bengio and J. S. Senecal · 2008
Earlier work this paper cites.
Matrix factorization techniques for recommender systems
Yehuda Koren, Robert Bell, and Chris Volinsky · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvarinen · 2010
Earlier work this paper cites.
A stick-breaking construction of the beta process
John Paisley, Aimee Zaas, Christopher W. Woods, Geoffrey S. Ginsburg, and Lawrence Carin · 2010
Earlier work this paper cites.
Hierarchical Bayesian nonparametric models with applications , pp. 158–207
Yee Whye Teh and Michael I. Jordan · 2010
Earlier work this paper cites.
The indian buffet process: An introduction and review
Thomas L. Griffiths and Zoubin Ghahramani · 2011
Earlier work this paper cites.
Nonparametric bayesian sparse factor models with application to gene expression modeling
David Knowles and Zoubin Ghahramani · 2011
Earlier work this paper cites.
Low Rank Approximation: Algorithms, Implementation, Applications
Ivan Markovsky · 2011
Cited alongside, same era.
Truly nonparametric online variational inference for hierarchical dirichlet processes
Michael Bryant and Erik Sudderth · 2012
Cited alongside, same era.
Small-variance asymptotics for exponential family dirichlet process mixture models
Ke Jiang, Brian Kulis, and Michael I. Jordan · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
Andriy Mnih and Yee Whye Teh · 2012
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Cited alongside, same era.
Small-variance asymptotics for hidden markov models
Anirban Roychowdhury, Ke Jiang, and Brian Kulis · 2013
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Later among the works it cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Xin Dong, Shangyu Chen, and Sinno Pan · 2017
Later among the works it cites.
Learning to hash with optimized anchor embedding for scalable retrieval
Y. Guo, G. Ding, L. Liu, J. Han, and L. Shao · 2017
Later among the works it cites.
Learning efficient convolutional networks through network slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang · 2017
Later among the works it cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 2017
Later among the works it cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Cited alongside, same era.
SNAP Datasets: Stanford large network dataset collection
Jure Leskovec and Andrej Krevl · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
Coresets for nonparametric estimation - the case of dp-means
Olivier Bachem, Mario Lucic, and Andreas Krause · 2015
Cited alongside, same era.
Compressing neural networks with the hashing trick
Wenlin Chen, James T. Wilson, Stephen Tyree, Kilian Q. Weinberger, and Yixin Chen · 2015
Cited alongside, same era.
The movielens datasets: History and context
F. Maxwell Harper and Joseph A. Konstan · 2015
Cited alongside, same era.
Later among the works it cites.
A mixture model for learning multi-sense word embeddings
Dai Quoc Nguyen, Dat Quoc Nguyen, Ashutosh Modi, Stefan Thater, and Manfred Pinkal · 2017
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2017
Later among the works it cites.
An efficient deep learning hashing neural network for mobile visual search
Heng Qi, Wu Liu, and Liang Liu · 2017
Later among the works it cites.
Hash embeddings for efficient word representations
Dan Svenstrup, Jonas Meinertz Hansen, and Ole Winther · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Probabilistic FastText for multi-sense word embeddings
Ben Athiwaratkun, Andrew Wilson, and Anima Anandkumar · 2018
Later among the works it cites.
Towards learning sparsely used dictionaries with arbitrary supports
P. Awasthi and A. Vijayaraghavan · 2018
Later among the works it cites.
Learning k-way d-dimensional discrete codes for compact embedding representations
Ting Chen, Martin Renqiang Min, and Yizhou Sun · 2018
Later among the works it cites.
Bayesian compression for natural language processing
Nadezhda Chirkova, Ekaterina Lobacheva, and Dmitry Vetrov · 2018
Later among the works it cites.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Later among the works it cites.
Word embedding based on low-rank doubly stochastic matrix decomposition
Denis Sedov and Zhirong Yang · 2018
Later among the works it cites.
Compressing word embeddings via deep compositional code learning
Raphael Shu and Hideki Nakayama · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Online embedding compression for text classification using low rank matrix factorization
Anish Acharya, Rahul Goel, Angeliki Metallinou, and Inderjit Dhillon · 2019
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Later among the works it cites.
How large a vocabulary does text classification need? a variational approach to vocabulary selection
Wenhu Chen, Yu Su, Yilin Shen, Zhiyu Chen, Xifeng Yan, and William Yang Wang · 2019
Later among the works it cites.
Mixed dimension embeddings with application to memory-efficient recommendation systems
Antonio Ginart, Maxim Naumov, Dheevatsa Mudigere, Jiyan Yang, and James Zou · 2019
Later among the works it cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Jianmo Ni, Jiacheng Li, and Julian McAuley · 2019
Later among the works it cites.
Addressing marketing bias in product recommendations
Mengting Wan, Jianmo Ni, Rishabh Misra, and Julian McAuley · 2020
Closest in time.