Fetching the paper…
Reading the bibliography…
Transformers have been proven a successful model for a variety of tasks in sequence modeling.
Switchboard: Telephone speech corpus for research and development
Godfrey, J. J., Holliman, E. C., and McDaniel, J · 1992
Earlier work this paper cites.
The design for the wall street journal-based csr corpus
Paul, D. B. and Baker, J. M · 1992
Earlier work this paper cites.
Advances in automatic text summarization
Maybury, M · 1999
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Hochreiter, S., Bengio, Y., Frasconi, P., and Schmidhuber, J · 2001
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
The kaldi speech recognition toolkit
Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N., Hannemann, M., Motlicek, P., Qian, Y., Schwarz, P., Silovsky, J., Stemmer, G., and Vesely, K · 2011
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Shrivastava, A. and Li, P · 2014
Earlier work this paper cites.
Parallel training of dnns with natural gradient and parameter averaging
Povey, D., Zhang, X., and Khudanpur, S · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Arjovsky, M., Shah, A., and Bengio, Y · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W., Jaitly, N., Le, Q., and Vinyals, O · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Purely sequence-trained neural networks for asr based on lattice-free mmi
Povey, D., Peddinti, V., Galvez, D., Ghahremani, P., Manohar, V., Na, X., Wang, Y., and Khudanpur, S · 2016
Cited alongside, same era.
Efficient attention using a fixed-size memory representation
Britz, D., Guan, M. Y., and Luong, M.-T · 2017
Cited alongside, same era.
Chiu, C.-C. and Raffel, C · 2017
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Later among the works it cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Lee, J., Lee, Y., Kim, J., Kosiorek, A., Choi, S., and Teh, Y. W · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the properties of the softmax function with application in game theory and reinforcement learning, 2017
Gao, B. and Pavel, L · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Dong, L., Xu, S., and Xu, B · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Sukhbaatar, S., Grave, E., Bojanowski, P., and Joulin, A · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Closest in time.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2020
Closest in time.
Monotonic multihead attention
Ma, X., Pino, J. M., Cross, J., Puzon, L., and Gu, J · 2020
Closest in time.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2020
Closest in time.