Fetching the paper…
Reading the bibliography…
Embedding layers are commonly used to map discrete symbols into continuous embedding vectors that reflect their semantic meanings.
Matrix factorization techniques for recommender systems
Koren, Y., Bell, R., and Volinsky, C · 2009
Earlier work this paper cites.
Product quantization for nearest neighbor search
Jegou, H., Douze, M., and Schmid, C · 2010
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Bordes, A., Usunier, N., Garcia-Duran, A., Weston, J., and Yakhnenko, O · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Cartesian k-means
Norouzi, M. and Fleet, D. J · 2013
Earlier work this paper cites.
Reasoning with neural tensor networks for knowledge base completion
Socher, R., Chen, D., Manning, C. D., and Ng, A · 2013
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Compressing word embeddings via deep compositional code learning
Shu, R. and Nakayama, H · 2017
Later among the works it cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., et al · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
Fast decoding in sequence models using discrete latent variables
Kaiser, Ł., Roy, A., Vaswani, A., Parmar, N., Bengio, S., Uszkoreit, J., and Shazeer, N · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Cited alongside, same era.
Bag of tricks for efficient text classification
Joulin, A., Grave, E., Bojanowski, P., and Mikolov, T · 2017
Cited alongside, same era.
Neural machine translation (seq2seq) tutorial
Luong, M., Brevdo, E., and Zhao, R · 2017
Cited alongside, same era.
Adaptive mixture of low-rank factorizations for compact neural modeling
Chen, T., Lin, J., Lin, T., Han, S., Wang, C., and Zhou, D
Cited in the paper.
Learning k-way d-dimensional discrete codes for compact embedding representations
Chen, T., Min, M. R., and Sun, Y
Cited in the paper.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2018
Later among the works it cites.
Memory-Efficient Adaptive Optimization for Large-Scale Learning
Anil, R., Gupta, V., Koren, T., and Singer, Y · 2019
Closest in time.
On the downstream performance of compressed word embeddings
May, A., Zhang, J., Dao, T., and Ré, C · 2019
Closest in time.