Fetching the paper…
Reading the bibliography…
The remarkable success of transformers in the field of natural language processing has sparked the interest of the speech-processing community, leading to an exploration of their potential for modeling long-range dependencies within speech sequences.
D. Griffin and J. Lim, “Signal estimation from modified short-time Fourier transform,” IEEE Transactions on acoustics, speech, and signal processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proceedings of the IEEE , vol. 78, no. 10, pp. 1550–1560, 1990
1990
Earlier work this paper cites.
J. Schmidhuber and S. Hochreiter, “Long short-term memory,” Neural Comput , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
T. Chen and R. Rao, “Audio-visual integration in multimodal communication,” Proceedings of the IEEE , vol. 86, no. 5, pp. 837–852, 1998
1998
Earlier work this paper cites.
H. Ney, “Speech translation: Coupling of recognition and translation,” in 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No. 99CH36258) , vol. 1. IEEE, 1999, pp. 517–520
1999
Earlier work this paper cites.
S. Schötz, “Paralinguistic phonetics in nlp models & methods,” NLP term paper , 2002
2002
Earlier work this paper cites.
E. Matusov, S. Kanthak, and H. Ney, “On the integration of speech recognition and statistical machine translation,” in Ninth European Conference on Speech Communication and Technology , 2005
2005
Earlier work this paper cites.
P. Taylor, Text-to-speech synthesis . Cambridge university press, 2009
2009
Earlier work this paper cites.
2012
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE/ACM Transactions on audio, speech, and language processing , vol. 22, no. 10, 2014
2014
Earlier work this paper cites.
Q. Le and T. Mikolov, “Distributed representations of sentences and documents,” in Proceedings of The 31st International Conference on Machine Learning , 2014, p. 1188–1196
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
L. Deng, “Deep learning: from speech recognition to language and multimodal processing,” APSIPA Transactions on Signal and Information Processing , vol. 5, p. e1, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Bérard, O. Pietquin, L. Besacier, and C. Servan, “Listen and translate: A proof of concept for end-to-end speech-to-text translation,” in NIPS Workshop on end-to-end learning for speech and audio processing , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 6000–6010
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” Proc. Interspeech 2017 , pp. 4006–4010, 2017
2017
Earlier work this paper cites.
S. Ö. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman et al. , “Deep voice: Real-time neural text-to-speech,” in International Conference on Machine Learning . PMLR, 2017, pp. 195–204
2017
Earlier work this paper cites.
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin, “Convolutional sequence to sequence learning,” in International Conference on Machine Learning . PMLR, 2017, pp. 1243–1252
2017
Earlier work this paper cites.
Q. T. Do, S. Sakti, and S. Nakamura, “Toward expressive speech translation: A unified sequence-to-sequence lstms approach for translating words and emphasis.” in INTERSPEECH , 2017, pp. 2640–2644
2017
Earlier work this paper cites.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly translate foreign speech,” Proc. Interspeech 2017 , pp. 2625–2629, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Earlier work this paper cites.
Z. Zhang, J. Geiger, J. Pohjalainen, A. E.-D. Mousa, W. Jin, and B. Schuller, “Deep learning for environmentally robust speech recognition: An overview of recent developments,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 9, no. 5, pp. 1–28, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
A. Vaswani, S. Bengio, E. Brevdo, F. Chollet, A. Gomez, S. Gouws, L. Jones, Ł. Kaiser, N. Kalchbrenner, N. Parmar et al. , “Tensor2tensor for neural machine translation,” in Proceedings of the 13th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track) , 2018, pp. 193–199
2018
Earlier work this paper cites.
S. Zhou, L. Dong, S. Xu, and B. Xu, “A comparison of modeling units in sequence-to-sequence speech recognition with the transformer on mandarin chinese,” in International Conference on Neural Information Processing . Springer, 2018, pp. 210–220
2018
Earlier work this paper cites.
S. Zhou, L. Dong, S. Xu, and B. Xu, “Syllable-based sequence-to-sequence speech recognition with the transformer in mandarin chinese,” Proc. Interspeech 2018 , pp. 791–795, 2018
2018
Earlier work this paper cites.
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: 2000-speaker neural text-to-speech,” Proc. ICLR , pp. 214–217, 2018
2018
Earlier work this paper cites.
W. Ping, K. Peng, and J. Chen, “Clarinet: Parallel wave generation in end-to-end text-to-speech,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “Fftnet: A real-time speaker-dependent neural vocoder,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 2251–2255
2018
Earlier work this paper cites.
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, “End-to-end automatic speech translation of audiobooks,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 6224–6228
2018
Earlier work this paper cites.
L.-C. Vila, C. Escolano, J. A. Fonollosa, and M.-R. Costa-Jussà, “End-to-end speech translation with the transformer,” Proc. IberSPEECH 2018 , pp. 60–63, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Povey, H. Hadian, P. Ghahremani, K. Li, and S. Khudanpur, “A time-restricted self-attention layer for asr,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5874–5878
2018
Earlier work this paper cites.
M. Sperber, J. Niehues, G. Neubig, S. Stüker, and A. Waibel, “Self-attentional acoustic models,” Proc. Interspeech 2018 , pp. 3723–3727, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Xian, C. H. Lampert, B. Schiele, and Z. Akata, “Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 9, pp. 2251–2265, 2018
2018
Earlier work this paper cites.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang et al. , “A comparative study on transformer vs RNN in speech applications,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 449–456
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. B. Nassif, I. Shahin, I. Attili, M. Azzeh, and K. Shaalan, “Speech recognition using deep neural networks: A systematic review,” IEEE access , vol. 7, pp. 19 143–19 165, 2019
2019
Earlier work this paper cites.
R. A. Khalil, E. Jones, M. I. Babar, T. Jan, M. H. Zafar, and T. Alhussain, “Speech emotion recognition using deep learning techniques: A review,” IEEE Access , vol. 7, pp. 117 327–117 345, 2019
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “XLNet: Generalized autoregressive pretraining for language understanding,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 6706–6713
2019
Earlier work this paper cites.
G. Wang, “Deep text-to-speech system with seq2seq model,” arXiv preprint arXiv:1903.07398 , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Wang and S. Zhang, “Bridging commonsense reasoning and probabilistic planning via a differentiable neural logic framework,” in Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI-19) , 2019, p. 1091–1097
2019
Earlier work this paper cites.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 2978–2988
2019
Earlier work this paper cites.
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, “A comparison of transformer and LSTM encoder decoder models for ASR,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 8–15
2019
Earlier work this paper cites.
J. Li, X. Wang, Y. Li et al. , “The speechtransformer for large-scale Mandarin Chinese speech recognition,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 7095–7099
2019
Earlier work this paper cites.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3617–3621
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
E. Tsunoo, Y. Kashiwagi, T. Kumakura, and S. Watanabe, “Transformer ASR with contextual block processing,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 427–433
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Zhang, N. Ge, B. Chen, and K. Fan, “Lattice transformer for speech translation,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 6475–6484
2019
Earlier work this paper cites.
M. A. Di Gangi, M. Negri, and M. Turchi, “Adapting transformer to end-to-end spoken language translation,” in INTERSPEECH 2019 . International Speech Communication Association (ISCA), 2019, pp. 1133–1137
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , J. Burstein, C. Doran, and T. Solorio, Eds. Association for Computational Linguistics, 2019
2019
Earlier work this paper cites.
A. Conneau and G. Lample, “Cross-lingual language model pretraining,” in NeurIPS , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Budzianowski and I. Vulic, “Hello, it’s GPT-2 - how can I help you? towards the use of pretrained language models for task-oriented dialogue systems,” in 3rd Workshop on Neural Generation and Translation@EMNLP-IJCNLP . Association for Computational Linguistics, 2019
2019
Earlier work this paper cites.
H. Takatsu, K. Yokoyama, Y. Matsuyama, H. Honda, S. Fujie, and T. Kobayashi, “Recognition of intentions of users’ short responses for conversational news delivery system,” in Interspeech . ISCA, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Liu, Z. Zhao, and L.-D. Gan, “Intention detection based on bert-bilstm in taskoriented dialogue system,” in International Computer Conference on Wavelet Active Media Technology and Information Processing , 2019
2019
Earlier work this paper cites.
M. Korpusik, Z. Liu, and J. R. Glass, “A comparison of deep learning methods for language understanding,” in Interspeech . ISCA, 2019
2019
Earlier work this paper cites.
L. Zhang and H. Wang, “Using bidirectional transformer-crf for spoken language understanding,” in International Conference on Natural Language Processing and Chinese Computing NLPCC , ser. Lecture Notes in Computer Science, vol. 11838. Springer, 2019
2019
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in International Conference on Learning Representations ICLR . OpenReview.net, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 2, pp. 423–443, 2019
2019
Earlier work this paper cites.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proceedings of the conference. Association for Computational Linguistics. Meeting , vol. 2019. NIH Public Access, 2019, p. 6558
2019
Earlier work this paper cites.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 7464–7473
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Salazar, K. Kirchhoff, and Z. Huang, “Self-attention networks for connectionist temporal classification in speech recognition,” in Icassp 2019-2019 ieee international conference on acoustics, speech and signal processing (icassp) . IEEE, 2019, pp. 7115–7119
2019
Earlier work this paper cites.
L. Dong, F. Wang, and B. Xu, “Self-attention aligner: A latency-control end-to-end model for ASR using self-attention network and chunk-hopping,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5656–5660
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Li, R. Dabre, X. Lu, P. Shen, T. Kawahara, and H. Kawai, “Improving transformer-based speech recognition systems with compressed structure and speech attributes augmentation.” in Interspeech , 2019, pp. 4400–4404
2019
Earlier work this paper cites.
X. Ma, P. Zhang, S. Zhang, N. Duan, Y. Hou, M. Zhou, and D. Song, “A tensorized transformer for language modeling,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Gao, Z. Jiang, H. You, Z. Lu, S. C. Hoi, and X. Wang, “Dynamic fusion with intra- and inter-modality attention flow for visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, p. 6639–6648
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, p. 5103–5114
2019
Earlier work this paper cites.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7829–7833
2020
Earlier work this paper cites.
——, “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations , 2020, pp. 38–45
2020
Earlier work this paper cites.
M. Alam, M. D. Samad, L. Vidyaratne, A. Glandon, and K. M. Iftekharuddin, “Survey on deep neural networks in speech and vision systems,” Neurocomputing , vol. 417, pp. 302–321, 2020
2020
Earlier work this paper cites.
A. M. Braşoveanu and R. Andonie, “Visualizing transformers for nlp: a brief survey,” in 2020 24th International Conference Information Visualisation (IV) . IEEE, 2020, pp. 270–279
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
N. R. Wu and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-networks,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, p. 3982–3992
2020
Earlier work this paper cites.
A. Baevski, M. A. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=r1gEjCNYwS
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Li, Y. Wu, Y. Gaur, C. Wang, R. Zhao, and S. Liu, “On the comparison of popular end-to-end models for large scale speech recognition,” Proc. Interspeech 2020 , pp. 1–5, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Wu, Y. Wang, Y. Shi, C.-F. Yeh, and F. Zhang, “Streaming transformer-based acoustic models using self-attention with augmented memory,” Proc. Interspeech 2020 , pp. 2132–2136, 2020
2020
Earlier work this paper cites.
O. Hrinchuk, M. Popova, and B. Ginsburg, “Correction of automatic speech recognition with transformer sequence-to-sequence model,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7074–7078
2020
Cited alongside, same era.
Z. Tian, J. Yi, Y. Bai, J. Tao, S. Zhang, and Z. Wen, “Synchronous transformers for end-to-end speech recognition,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7884–7888
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
T. Likhomanenko, Q. Xu, G. Synnaeve, R. Collobert, and A. Rogozhnikov, “CAPE: Encoding relative positions with continuous augmented positional embeddings,” Advances in Neural Information Processing Systems , vol. 34, pp. 16 079–16 092, 2021
2021
Later among the works it cites.
B. J. Woo, H. Y. Kim, J. Kim, and N. S. Kim, “Speech separation based on dptnet with sparse attention,” in 2021 7th IEEE International Conference on Network Intelligence and Digital Content (IC-NIDC) . IEEE, 2021, pp. 339–343
2021
Later among the works it cites.
M. Zaheer, G. Guruganesh, A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang et al. , “Constructing transformers for longer sequences with sparse attention methods,” Google AI Blog , 2021
2021
Later among the works it cites.
A. Roy, M. Saffar, A. Vaswani, and D. Grangier, “Efficient content-based sparse attention with routing transformers,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 53–68, 2021
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2020
Cited alongside, same era.
T. Moriya, T. Ochiai, S. Karita, H. Sato, T. Tanaka, T. Ashihara, R. Masumura, Y. Shinohara, and M. Delcroix, “Self-distillation for improving CTC-Transformer-Based ASR Systems,” in INTERSPEECH , 2020, pp. 546–550
2020
Cited alongside, same era.
A. Jain, A. Rouhe, S.-A. Grönroos, M. Kurimo et al. , “Finnish ASR with deep transformer models.” in Interspeech , 2020, pp. 3630–3634
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Li, Y. Liu, Y. Wu, S. Liu, S. Zhao, and M. Liu, “Robutrans: A robust transformer-based text-to-speech model,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8228–8235
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Zheng, X. Li, F. Xie, and L. Lu, “Improving end-to-end speech synthesis with local recurrent neural network enhanced transformer,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6734–6738
2020
Cited alongside, same era.
M. Chen, X. Tan, Y. Ren, J. Xu, H. Sun, S. Zhao, and T. Qin, “Multispeech: Multi-speaker text to speech with transformer,” Proc. Interspeech 2020 , pp. 4024–4028, 2020
2020
Cited alongside, same era.
Later among the works it cites.
2021
Later among the works it cites.
J. Li, R. Cotterell, and M. Sachan, “Differentiable subset pruning of transformer heads,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1442–1459, 2021
2021
Later among the works it cites.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5904–5908
2021
Later among the works it cites.
Y. Shi, Y. Wang, C. Wu, C.-F. Yeh, J. Chan, F. Zhang, D. Le, and M. Seltzer, “Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6783–6787
2021
Later among the works it cites.
J. Xu, S. Hu, J. Yu, X. Liu, and H. Meng, “Mixed precision quantization of transformer language models for speech recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7383–7387
2021
Later among the works it cites.
A. Miech, J.-B. Alayrac, I. Laptev, J. Sivic, and A. Zisserman, “Thinking fast and slow: Efficient text-to-visual retrieval with transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 9826–9836
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International conference on machine learning . PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Wen, J. Xia, Y. Huang, L. Li, J. Xu, and J. Shao, “Cookie: Contrastive cross-modal knowledge sharing pre-training for vision-language representation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 2208–2217
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
B. Xue, J. Yu, J. Xu, S. Liu, S. Hu, Z. Ye, M. Geng, X. Liu, and H. Meng, “Bayesian transformer language models for speech recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7378–7382
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Later among the works it cites.
C. Kervadec, C. Wolf, G. Antipov, M. Baccouche, and M. Nadri, “Supervising the transfer of reasoning patterns in vqa,” Advances in Neural Information Processing Systems , vol. 34, pp. 18 256–18 267, 2021
2021
Later among the works it cites.
X. Zhan, Y. Wu, X. Dong, Y. Wei, M. Lu, Y. Zhang, H. Xu, and X. Liang, “Product1M: Towards weakly supervised instance-level product retrieval via cross-modal pretraining,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 11 782–11 791
2021
Later among the works it cites.
Q. Xia, H. Huang, N. Duan, D. Zhang, L. Ji, Z. Sui, E. Cui, T. Bharti, and M. Zhou, “XGPT: Cross-modal generative pre-training for image captioning,” in Natural Language Processing and Chinese Computing: 10th CCF International Conference, NLPCC 2021, Qingdao, China, October 13–17, 2021, Proceedings, Part I 10 . Springer, 2021, pp. 786–797
2021
Later among the works it cites.
M. Zhou, L. Zhou, S. Wang, Y. Cheng, L. Li, Z. Yu, and J. Liu, “Uc2: Universal cross-lingual cross-modal vision-and-language pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 4155–4165
2021
Later among the works it cites.
M. Ni, H. Huang, L. Su, E. Cui, T. Bharti, L. Wang, D. Zhang, and N. Duan, “M3p: Learning universal representations via multitask multilingual multimodal pre-training,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 3977–3986
2021
Later among the works it cites.
Z. Li, Z. Liu, and X. Zhang, “Causal attention for vision-language tasks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, p. 10533–10542
2021
Later among the works it cites.
Y. Chen, Y. Wang, X. Liu, C. Qian, L. Lin, and C. C. Loy, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, p. 10442–10451
2021
Later among the works it cites.
C. Zhuge, H. Zhang, J. Liang, X. Zhang, and Z. Luo, “Kaleido-bert: Vision-language pre-training on fashion domain,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, p. 14529–14538
2021
Later among the works it cites.
2021
Later among the works it cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 4904–4916
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less is more: Clipbert for video-and-language learning via sparse sampling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7331–7341
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Narasimhan, A. Rohrbach, and T. Darrell, “Clip-it! language-guided video summarization,” Advances in Neural Information Processing Systems , vol. 34, pp. 13 988–14 000, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
——, “Gilbert: Generative vision-language pre-training for image-text retrieval,” in Proceedings of the Web Conference 2021 , 2021, p. 1143–1154
2021
Later among the works it cites.
T. Lin, Y. Wang, X. Liu, and X. Qiu, “A survey of transformers,” AI Open , 2022
2022
Later among the works it cites.
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient transformers: A survey,” ACM Computing Surveys , vol. 55, no. 6, pp. 1–28, 2022
2022
Later among the works it cites.
Q. Song, B. Sun, and S. Li, “Multimodal sparse transformer network for audio-visual speech recognition,” IEEE Transactions on Neural Networks and Learning Systems , 2022
2022
Later among the works it cites.
F. Dang, H. Chen, and P. Zhang, “Dpt-fsnet: Dual-path transformer based full-band and sub-band fusion network for speech enhancement,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6857–6861
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM computing surveys (CSUR) , vol. 54, no. 10s, pp. 1–41, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
K. B. Bhangale and M. Kothandaraman, “Survey of deep learning paradigms for speech processing,” Wireless Personal Communications , vol. 125, no. 2, pp. 1913–1949, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “Data2vec: A general framework for self-supervised learning in speech, vision and language,” in International Conference on Machine Learning . PMLR, 2022, pp. 1298–1312
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Zhang, D. S. Park, W. Han, J. Qin, A. Gulati, J. Shor, A. Jansen, Y. Xu, Y. Huang, S. Wang et al. , “BigSSL: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1519–1532, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, Z. Chen, Z. Chen, S. Liu, J. Wu, Y. Qian, F. Wei, J. Li et al. , “UniSpeech-SAT: Universal speech representation learning with speaker aware pre-training,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6152–6156
2022
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Later among the works it cites.
L.-W. Chen and A. Rudnicky, “Fine-grained style control in transformer-based text-to-speech synthesis,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7907–7911
2022
Later among the works it cites.
Huawei. (2022) Speaking your language: The transformer in machine translation. [Online]. Available: https://blog.huawei.com/2022/02/01/speaking-your-language-transformer-machine-translation/
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Shor and S. Venugopalan, “TRILLsson: Distilling universal paralinguistic speech representations,” 2022
2022
Later among the works it cites.
J. Shor, A. Jansen, W. Han, D. Park, and Y. Zhang, “Universal paralinguistic speech representations using self-supervised conformers,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 3169–3173
2022
Later among the works it cites.
2022
Later among the works it cites.
W. Yu, J. Zhou, H. Wang, and L. Tao, “SETransformer: speech enhancement transformer,” Cognitive Computation , pp. 1–7, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
A. G. C. P. Ramos, A. Mehrotra, N. D. Lane, and S. Bhattacharya, “Conditioning sequence-to-sequence networks with learned activations,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
T. Wu and B. Juang, “Induce spoken dialog intents via deep unsupervised context contrastive clustering,” in Interspeech . ISCA, 2022
2022
Later among the works it cites.
W. A. Abro, A. Aicher, N. Rach, S. Ultes, W. Minker, and G. Qi, “Natural language understanding for argumentative dialogue systems in the opinion building domain,” Knowl. Based Syst. , vol. 242, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Jang, J. Lee, and K. Kim, “GPT-Critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems,” in ICLR . OpenReview.net, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Sakuma, S. Fujie, and T. Kobayashi, “Response timing estimation for spoken dialog systems based on syntactic completeness prediction,” in SLT . IEEE, 2022
2022
Later among the works it cites.
T. Lin, Y. Wu, F. Huang, L. Si, J. Sun, and Y. Li, “Duplex conversation: Towards human-like interaction in spoken dialogue systems,” in KDD . ACM, 2022
2022
Later among the works it cites.
A. Waheed, “Combining neural networks with knowledge for spoken dialogue systems,” Ph.D. dissertation, Ulm University, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Dong, J. Fu, P. Zhou, H. Li, and X. Wang, “Improving spoken language understanding with cross-modal contrastive learning,” in Interspeech . ISCA, 2022
2022
Later among the works it cites.
J. Svec, A. Frémund, M. Bulín, and J. Lehecka, “Transfer learning of transformers for spoken language understanding,” in International Conference on Text, Speech, and Dialogue (TSD) , ser. Lecture Notes in Computer Science, vol. 13502. Springer, 2022
2022
Later among the works it cites.
W. Shen, X. He, C. Zhang, X. Zhang, and J. Xie, “A transformer-based user satisfaction prediction for proactive interaction mechanism in dueros,” in International Conference on Information & Knowledge Management (CIKM) , M. A. Hasan and L. Xiong, Eds. ACM, 2022
2022
Later among the works it cites.
J. Yang, P. Wang, Y. Zhu, M. Feng, M. Chen, and X. He, “Gated multimodal fusion with contrastive learning for turn-taking prediction in human-robot dialogue,” in ICASSP . IEEE, 2022
2022
Later among the works it cites.
V. Sunder, S. Thomas, H. J. Kuo, J. Ganhotra, B. Kingsbury, and E. Fosler-Lussier, “Towards end-to-end integration of dialog history for improved spoken language understanding,” in ICASSP . IEEE, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Yu, D. Park, and H. K. Kim, “Auxiliary loss of transformer with residual connection for end-to-end speaker diarization,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 8377–8381
2022
Later among the works it cites.
Z. Duan, G. Gao, J. Chen, S. Li, J. Ruan, G. Yang, and X. Yu, “Dual-residual transformer network for speech recognition,” Journal of the Audio Engineering Society , vol. 70, no. 10, pp. 871–881, 2022
2022
Later among the works it cites.
W. G. Tech, “Sparse transformers and longformers: A comprehensive summary of space and time optimizations on transformer architectures,” https://medium.com/walmartglobaltech/sparse-transformers-and-longformers-a-comprehensive-summary-of-space-and-time-optimizations-on-4caa5c388693 , 2021, accessed: 2022-01-28
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Pham Minh, N. Nguyen Xuan, and S. Tran Thai, “Tt-vit: Vision transformer compression using tensor-train decomposition,” in Computational Collective Intelligence: 14th International Conference, ICCCI 2022, Hammamet, Tunisia, September 28–30, 2022, Proceedings . Springer, 2022, pp. 755–767
2022
Later among the works it cites.
S. Li, P. Zhang, G. Gan, X. Lv, B. Wang, J. Wei, and X. Jiang, “Hypoformer: Hybrid decomposition transformer for edge-friendly neural machine translation,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 7056–7068
2022
Later among the works it cites.
D. Timonin, B. Y. Hsueh, and V. Nguyen, “Accelerated inference for large transformer models using nvidia triton inference server,” https://developer.nvidia.com/blog/accelerated-inference-for-large-transformer-models-using-nvidia-fastertransformer-and-nvidia-triton-inference-server/ , Aug 2022
2022
Later among the works it cites.
Z. Gan, Y.-C. Chen, L. Li, T. Chen, Y. Cheng, S. Wang, J. Liu, L. Wang, and Z. Liu, “Playing lottery tickets with vision and language,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 1, 2022, pp. 652–660
2022
Later among the works it cites.
S. Yan, X. Xiong, A. Arnab, Z. Lu, M. Zhang, C. Sun, and C. Schmid, “Multiview transformers for video recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3333–3343
2022
Later among the works it cites.
2022
Later among the works it cites.
D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi, “Align and prompt: Video-and-language pre-training with entity prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4953–4963
2022
Later among the works it cites.
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,” Neurocomputing , vol. 508, pp. 293–304, 2022
2022
Later among the works it cites.
S. Latif, H. Cuayáhuitl, F. Pervez, F. Shamshad, H. S. Ali, and E. Cambria, “A survey on deep reinforcement learning for audio-based applications,” Artificial Intelligence Review , vol. 56, no. 3, pp. 2193–2240, 2023
2023
Closest in time.
2023
Closest in time.
J. Bgn, “Timeline of transformers for speech,” https://jonathanbgn.com/2021/12/31/timeline-transformers-speech.html , 2021, accessed: 2023-03-07
2023
Closest in time.
2023
Closest in time.
W. Chen, X. Xing, X. Xu, J. Pang, and L. Du, “Speechformer++: A hierarchical efficient framework for paralinguistic speech processing,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Closest in time.
2023
Closest in time.
A. López-Zorrilla, M. I. Torres, and H. Cuayáhuitl, “Audio embedding-aware dialogue policy learning,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 31, 2023
2023
Closest in time.
M. Firdaus, A. Ekbal, and E. Cambria, “Multitask learning for multilingual intent detection and slot filling in dialogue systems,” Inf. Fusion , vol. 91, 2023
2023
Closest in time.
J. Mei, Y. Wang, X. Tu, M. Dong, and T. He, “Incorporating BERT with probability-aware gate for spoken language understanding,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 31, 2023
2023
Closest in time.
M. Burchi and R. Timofte, “Audio-visual efficient conformer for robust speech recognition,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 2258–2267
2023
Closest in time.
W. Jiang, C. Sun, F. Chen, Y. Leng, Q. Guo, J. Sun, and J. Peng, “Low complexity speech enhancement network based on frame-level Swin transformer,” Electronics , 2023. [Online]. Available: https://www.mdpi.com/2079-9292/12/6/1330
2079
Closest in time.