Fetching the paper…
Reading the bibliography…
In natural language processing (NLP), the "Transformer" architecture was proposed as the first transduction model replying entirely on self-attention mechanisms without using sequence-aligned recurrent neural networks (RNNs) or convolution, and it achieved significant improvements for sequence to sequence tasks.
Cooley-tukey FFT on the connection machine
S Lennart Johnsson and Robert L Krawitz. 1992 · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Mathematics of the discrete Fourier transform (DFT): with audio applications
Julius Orion Smith. 2007 · 2007
Earlier work this paper cites.
Learning word vectors for sentiment analysis. In ACL . Association for Computational Linguistics, 142–150
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Structured matrices and polynomials: unified superfast algorithms
Victor Pan. 2012 · 2012
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Towards Neural Machine Translation with Latent Tree Attention
James Bradbury and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Circnn: accelerating and compressing deep neural networks using block-circulant weight matrices. In Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture . 395–408
Caiwen Ding, Siyu Liao, Yanzhi Wang, Zhe Li, Ning Liu, Youwei Zhuo, Chao Wang, Xuehai Qian, Yu Bai, Geng Yuan, et al · 2017
Cited alongside, same era.
Yoon Kim, Carl Denton, Luong Hoang, and Alexander M Rush. 2017 · 2017
Cited alongside, same era.
Pointer Sentinel Mixture Models. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. [n. d.] · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Supervised Domain Enablement Attention for Personalized Domain Classification. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . 894–899
Joo-Kyung Kim and Young-Bum Kim. 2018 · 2018
Later among the works it cites.
Simple Attention-Based Representation Learning for Ranking Short Social Media Posts
Peng Shi, Jinfeng Rao, and Jimmy Lin. 2018 · 2018
Later among the works it cites.
C-LSTM: Enabling Efficient LSTM Using Structured Compression Techniques on FPGAs. In FPGA’18
Shuo Wang, Zhe Li, Caiwen Ding, Bo Yuan, Qinru Qiu, Yanzhi Wang, and Yun Liang. 2018 · 2018
Later among the works it cites.
Supercharge Your AI and Database Applications with Xilinx’s HBM-Enabled UltraScale+ Devices Featuring Samsung HBM2
2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank. In International Conference on Machine Learning . 4082–4090
Liang Zhao, Siyu Liao, Yanzhi Wang, Zhe Li, Jian Tang, and Bo Yuan. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Later among the works it cites.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling. In Proceedings of NAACL-HLT 2019: Demonstrations
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.