Fetching the paper…
Reading the bibliography…
The embedding layers transforming input words into real vectors are the key components of deep neural networks used in natural language processing.
Analysis of individual differences in multidimensional scaling via an n-way generalization of Eckart-Young decomposition
Carroll, J. D. and Chang, J.-J · 1970
Earlier work this paper cites.
Foundations of the PARAFAC procedure: Models and conditions for an” explanatory” multimodal factor analysis
Harshman, R. A · 1970
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Approximation of 2 d × 2 d 2^{d}\times 2^{d} matrices using tensor decomposition
Oseledets, I. V · 2010
Earlier work this paper cites.
Matrix product operator representations
Pirvu, B., Murg, V., Cirac, J. I., and Verstraete, F · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Tensor-train decomposition
Oseledets, I. V · 2011
Earlier work this paper cites.
Japanese and korean voice search
Schuster, M. and Nakajima, K · 2012
Earlier work this paper cites.
A literature survey of low-rank tensor approximation techniques
Grasedyck, L., Kressner, D., and Tobler, C · 2013
Earlier work this paper cites.
Algebraic geometry , volume 52
Hartshorne, R · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Kaggle Display Advertising Challenge, 2014
Criteo Labs · 2014
Earlier work this paper cites.
Practical lessons from predicting clicks on ads at facebook
He, X., Pan, J., Jin, O., Xu, T., Liu, B., Xu, T., Shi, Y., Atallah, A., Herbrich, R., Bowers, S., et al · 2014
Earlier work this paper cites.
The hackbusch conjecture on tensor formats
Buczyńska, W., Buczyński, J., and Michałek, M · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Cited alongside, same era.
Speeding-up convolutional neural networks using fine-tuned CP-decomposition
Lebedev, V., Ganin, Y., Rakhuba, M., Oseledets, I., and Lempitsky, V · 2015
Cited alongside, same era.
Tensorizing neural networks
Novikov, A., Podoprikhin, D., Osokin, A., and Vetrov, D. P · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Cited alongside, same era.
Ultimate tensorization: compressing convolutional and FC layers alike
Compressing recurrent neural network with tensor train
Tjandra, A., Sakti, S., and Nakamura, S · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Long-term forecasting using tensor-train RNNs
Yu, R., Zheng, S., Anandkumar, A., and Yue, Y · 2017
Later among the works it cites.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2018
Later among the works it cites.
Word2Bits-quantized word vectors
Lam, M · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Garipov, T., Podoprikhin, D., Novikov, A., and Vetrov, D · 2016
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, H., Khosravi, K., and Socher, R · 2016
Cited alongside, same era.
Fasttext. zip: Compressing text classification models
Joulin, A., Grave, E., Bojanowski, P., Douze, M., Jégou, H., and Mikolov, T · 2016
Cited alongside, same era.
Field-aware factorization machines for CTR prediction
Juan, Y., Zhuang, Y., Chin, W.-S., and Lin, C.-J · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2016
Cited alongside, same era.
Compression of neural machine translation models via pruning
See, A., Luong, M.-T., and Manning, C. D · 2016
Cited alongside, same era.
A call for clarity in reporting bleu scores
Post, M · 2018
Later among the works it cites.
WEST: Word Encoded Sequence Transducers
Variani, E., Suresh, A. T., and Weintraub, M · 2018
Later among the works it cites.
Wide compression: Tensor ring nets
Wang, W., Sun, Y., Eriksson, B., Wang, W., and Aggarwal, V · 2018
Later among the works it cites.
Deep neural network compression with single and multiple level quantization
Xu, Y., Wang, Y., Zhou, A., Lin, W., and Xiong, H · 2018
Later among the works it cites.
transformers. zip: Compressing transformers with pruning and quantization
Cheong, R. and Daniel, R · 2019
Closest in time.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Closest in time.
Generalized tensor models for recurrent neural networks
Khrulkov, V., Hrinchuk, O., and Oseledets, I · 2019
Closest in time.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Closest in time.
A tensorized transformer for language modeling
Ma, X., Zhang, P., Zhang, S., Duan, N., Hou, Y., Song, D., and Zhou, M · 2019
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Closest in time.
Pay less attention with lightweight and dynamic convolutions
Wu, F., Fan, A., Baevski, A., Dauphin, Y. N., and Auli, M · 2019
Closest in time.