Fetching the paper…
Reading the bibliography…
Transformers have revolutionized deep learning across various tasks, including audio representation learning, due to their powerful modeling capabilities.
“A new approach to linear filtering and prediction problems,”
Rudolph Emil Kalman, · 1960
Earlier work this paper cites.
“IEMOCAP: interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan, · 2008
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Esc: Dataset for environmental sound classification,”
Karol J Piczak, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Earlier work this paper cites.
“Voxceleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Speech commands: A dataset for limited-vocabulary speech recognition,”
Pete Warden, · 2018
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning,” 2019
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby, · 2021
Cited alongside, same era.
“AST: Audio Spectrogram Transformer,”
Yuan Gong, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Mamba: Linear-time sequence modeling with selective state spaces,”
Albert Gu and Tri Dao, · 2023
Later among the works it cites.
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang, · 2024
Closest in time.
“Swin-umamba: Mamba-based unet with imagenet-based pretraining,”
Jiarun Liu, Hao Yang, Hong-Yu Zhou, Yan Xi, Lequan Yu, Yizhou Yu, Yong Liang, Guangming Shi, Shaoting Zhang, Hairong Zheng, et al., · 2024
Closest in time.
“U-mamba: Enhancing long-range dependency for biomedical image segmentation,” 2024
Jun Ma, Feifei Li, and Bo Wang, · 2024
Closest in time.
“Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation,” 2024
Zhaohu Xing, Tian Ye, Yijun Yang, Guang Liu, and Lei Zhu, · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré, · 2021
Cited alongside, same era.
“Efficiently modeling long sequences with structured state spaces,”
Albert Gu, Karan Goel, and Christopher Ré, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Self-Supervised Speech Representation Learning: A Review,”
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D. Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaloe, Tara N. Sainath, and Shinji Watanabe, · 2022
Cited alongside, same era.
“Ssast: Self-supervised audio spectrogram transformer,”
Yuan Gong, Cheng-I Lai, Yu-An Chung, and James Glass, · 2022
Cited alongside, same era.
“Videomamba: State space model for efficient video understanding,” 2024
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao, · 2024
Closest in time.
“Graph-mamba: Towards long-range graph sequence modeling with selective state spaces,” 2024
Chloe Wang, Oleksii Tsepa, Jun Ma, and Bo Wang, · 2024
Closest in time.
“Multichannel long-term streaming neural speech enhancement for static and moving speakers,”
Changsheng Quan and Xiaofei Li, · 2024
Closest in time.
“Tramba: A hybrid transformer and mamba architecture for practical audio and bone conduction speech super resolution and enhancement on mobile and wearable platforms,” 2024
Yueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang, and Stephen Xia, · 2024
Closest in time.
Xilin Jiang, Cong Han, and Nima Mesgarani, · 2024
Closest in time.