Fetching the paper…
Reading the bibliography…
In this paper, we present Reshape Dimensions Network (ReDimNet), a novel neural network architecture for extracting utterance-level speaker representations.
S.-C. Yin, R. Rose, and P. Kenny, “Adaptive score normalization for progressive model adaptation in text independent speaker verification,” in 2008 IEEE International Conference on Acoustics, Speech and Signal Processing , 2008, pp. 4857–4860
2008
Earlier work this paper cites.
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition.” in Interspeech , vol. 2015, 2015, p. 3586
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
M. McLaren, L. Ferrer, D. Castan, and A. Lawson, “The speakers in the wild (sitw) speaker recognition database.” in Interspeech , 2016, pp. 818–822
2016
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Interspeech 2018 . ISCA, Sep. 2018
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” in Interspeech 2018 . ISCA, Sep. 2018
2018
Earlier work this paper cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 4690–4699
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
I. Szöke, M. Skácel, L. Mošner, J. Paliesek, and J. Černockỳ, “Building and evaluation of a real room impulse response dataset,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 4, pp. 863–876, 2019
2019
Earlier work this paper cites.
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Science and Language , 2019
2019
Earlier work this paper cites.
M. K. Nandwana, J. van Hout, M. McLaren, C. Richey, A. Lawson, and M. A. Barrios, “The voices from a distance challenge 2019 evaluation plan,” 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Y.-Q. Yu and W.-J. Li, “Densely connected time delay neural network for speaker verification,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 921–925
2020
Cited alongside, same era.
D. Garcia-Romero, G. Sell, and A. Mccree, “Magneto: X-vector magnitude estimation network plus offset for improved speaker recognition.” in Odyssey , 2020, pp. 1–8
2020
Cited alongside, same era.
S. Zheng, L. Cheng, Y. Chen, H. Wang, and Q. Chen, “3d-speaker: A large-scale multi-device, multi-distance, and multi-dialect corpus for speech representation disentanglement,” 2023
2023
Later among the works it cites.
I. Yakovlev, A. Okhotnikov, N. Torgashov, R. Makarov, Y. Voevodin, and K. Simonchik, “VoxTube: a multilingual speaker recognition dataset,” in Proc. INTERSPEECH 2023 , 2023, pp. 2238–2242
2023
Later among the works it cites.
H.-J. Heo, U.-H. Shin, R. Lee, Y. Cheon, and H.-M. Park, “Next-tdnn: Modernizing multi-scale temporal convolution backbone for speaker verification,” 2023
2023
Later among the works it cites.
Z. Zhao, Z. Li, W. Wang, and P. Zhang, “Pcf: Ecapa-tdnn with progressive channel fusion for speaker verification,” 2023
2023
Later among the works it cites.
H. Wang, S. Zheng, Y. Chen, L. Cheng, and Q. Chen, “CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking,” in Proc. INTERSPEECH 2023 , 2023, pp. 5301–5305
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
M. Tan and Q. V. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
T. Liu, R. K. Das, K. A. Lee, and H. Li, “Mfa: Tdnn with multi-scale frequency-channel attention for text-independent speaker verification with short utterances,” 2022
2022
Cited alongside, same era.
Y. Zhang, Z. Lv, H. Wu, S. Zhang, P. Hu, Z. Wu, H. yi Lee, and H. Meng, “Mfa-conformer: Multi-scale feature aggregation conformer for automatic speaker verification,” 2022
2022
Cited alongside, same era.
B. Liu, Z. Chen, S. Wang, H. Wang, B. Han, and Y. Qian, “DF-ResNet: Boosting Speaker Verification Performance with Depth-First Design,” in Proc. Interspeech 2022 , 2022, pp. 296–300
2022
Cited alongside, same era.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” 2022
2022
Cited alongside, same era.
Y. Lin, X. Qin, G. Zhao, M. Cheng, N. Jiang, H. Wu, and M. Li, “Voxblink: A large scale speaker verification dataset on camera,” 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
Y. Chen, S. Zheng, H. Wang, L. Cheng, Q. Chen, and J. Qi, “An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification,” in Proc. INTERSPEECH 2023 , 2023, pp. 2228–2232
2023
Later among the works it cites.
T. Liu, K. A. Lee, Q. Wang, and H. Li, “Golden gemini is all you need: Finding the sweet spots for speaker verification,” 2023
2023
Later among the works it cites.
B. Han, Z. Chen, and Y. Qian, “Exploring binary classification loss for speaker verification,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Later among the works it cites.
H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y. Deng, and Y. Qian, “Wespeaker: A research and production oriented speaker embedding learning toolkit,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” 2023
2023
Later among the works it cites.
K. Nam, Y. Kim, J. Huh, H. S. Heo, J. weon Jung, and J. S. Chung, “Disentangled representation learning for multilingual speaker recognition,” 2023
2023
Later among the works it cites.
J. Thienpondt and K. Demuynck, “Ecapa2: A hybrid neural network architecture and training strategy for robust speaker embeddings,” 2024
2024
Closest in time.