Fetching the paper…
Reading the bibliography…
While the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Anantapadmanabhan, A. Bellur, and H. A. Murthy, “Modal analysis and transcription of strokes of the mridangam using non-negative matrix factorization,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing , 2013
2013
Earlier work this paper cites.
M. Tian, A. Srinivasamurthy, M. Sandler, and X. Serra, “A study of instrument-wise onset detection in beijing opera percussion ensembles,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 2159–2163
2014
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE transactions on affective computing , vol. 5, no. 4, pp. 377–390, 2014
2014
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for Environmental Sound Classification,” in Proceedings of the 23rd Annual ACM Conference on Multimedia . ACM Press, 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, D. Eck, K. Simonyan, and M. Norouzi, “Neural audio synthesis of musical notes with wavenet autoencoders,” 2017
2017
Earlier work this paper cites.
F.-R. Stöter, S. Chakrabarty, B. Edler, and E. A. Habets, “Classification vs. regression in supervised learning for single channel speaker count estimation,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Earlier work this paper cites.
P. Warden, “Speech commands: A dataset for limited-vocabulary speech recognition,” 2018
2018
Earlier work this paper cites.
B. Kim, M. Ghei, B. Pardo, and Z. Duan, “Vocal imitation set: a dataset of vocally imitated sound events using the audioset ontology.” in DCASE , 2018, pp. 148–152
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019
2019
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 12 449–12 460
2020
Cited alongside, same era.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International conference on machine learning . PMLR, 2020, pp. 5156–5165
2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units.” IEEE Transactions on Audio, Speech, and Language Processing , pp. 1–1, 2021
2021
Cited alongside, same era.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “Mlp-mixer: An all-mlp architecture for vision,” Advances in neural information processing systems , vol. 34, pp. 24 261–24 272, 2021
A. Gu, K. Goel, and C. Re, “Efficiently modeling long sequences with structured state spaces,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation,” in HEAR: Holistic Evaluation of Audio Representations . PMLR, 2022, pp. 1–24
2022
Later among the works it cites.
J. Turian, J. Shier, H. R. Khan, B. Raj, B. W. Schuller, C. J. Steinmetz, C. Malloy, G. Tzanetakis, G. Velarde, K. McNally, M. Henry, N. Pinto, C. Noufi, C. Clough, D. Herremans, E. Fonseca, J. Engel, J. Salamon, P. Esling, P. Manocha, S. Watanabe, Z. Jin, and Y. Bisk, “HEAR: Holistic Evaluation of Audio Representations,” in Proceedings of the NeurIPS 2021 Competitions and Demonstrations Track . PMLR, Jul. 2022, pp. 125–145, iSSN: 2640-3498
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009
2022
Cited alongside, same era.
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “Ssast: Self-supervised audio spectrogram transformer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, 2022, pp. 10 699–10 709
2022
Cited alongside, same era.
P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked autoencoders that listen,” Advances in Neural Information Processing Systems , vol. 35, pp. 28 708–28 720, 2022
2022
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Cited alongside, same era.
2022
Later among the works it cites.
2023
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “Beats: audio pre-training with acoustic tokenizers,” in Proceedings of the 40th International Conference on Machine Learning , 2023, pp. 5178–5193
2023
Later among the works it cites.
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: efficient visual representation learning with bidirectional state space model,” in Proceedings of the 41st International Conference on Machine Learning . JMLR.org, 2024
2024
Closest in time.
S. Yadav and Z.-H. Tan, “Audio mamba: Selective state spaces for self-supervised audio representations,” in Proc. INTERSPEECH 2024 – 25 th Annual Conference of the International Speech Communication Association , Kos Island, Greece, Sep. 2024
2024
Closest in time.
M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. K. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter, “xLSTM: Extended long short-term memory,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024
2024
Closest in time.
S. Yadav, S. Theodoridis, L. K. Hansen, and Z.-H. Tan, “Masked autoencoders with multi-window local-global attention are better audio learners,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
B. Alkin, M. Beck, K. Pöppel, S. Hochreiter, and J. Brandstetter, “Vision-LSTM: xLSTM as generic vision backbone,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.