Fetching the paper…
Reading the bibliography…
Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task.
“Representation of functions by superpositions of a step or sigmoid function and their applications to neural network theory,”
Yoshifusa Ito, · 1991
Earlier work this paper cites.
“Switchboard cellular part 1 audio,”
David Miller David Graff, Kevin Walker, · 2001
Earlier work this paper cites.
“A method of estimating the equal error rate for automatic speaker verification,”
Jyh-Min Cheng and Hsiao-Chuan Wang, · 2004
Earlier work this paper cites.
“Mfcc and its applications in speaker recognition,”
Vibha Tiwari, · 2010
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P Murphy, · 2012
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Earlier work this paper cites.
“Unsupervised feature learning based on deep models for environmental audio tagging,”
Yong Xu, Qiang Huang, Wenwu Wang, Peter Foster, Siddharth Sigtia, Philip JB Jackson, and Mark D Plumbley, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Voxceleb: A large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Cited alongside, same era.
“Attention mechanism in speaker recognition: What does it learn in deep speaker embedding?,”
Qiongqiong Wang, Koji Okabe, Kong Aik Lee, Hitoshi Yamamoto, and Takafumi Koshinaka, · 2018
Cited alongside, same era.
“Weakly supervised training of speaker identification models,”
Martin Karu and Tanel Alumäe, · 2018
Cited alongside, same era.
“A brief introduction to weakly supervised learning,”
“Self-attentive speaker embeddings for text-independent speaker verification.,”
Yingke Zhu, Tom Ko, David Snyder, Brian Mak, and Daniel Povey, · 2018
Later among the works it cites.
“Attentive statistics pooling for deep speaker embedding,”
Koji Okabe, Takafumi Koshinaka, and Koichi Shinoda, · 2018
Later among the works it cites.
“Attention-based models for text-dependent speaker verification,”
FA Rezaur rahman Chowdhury, Quan Wang, Ignacio Lopez Moreno, and Li Wan, · 2018
Later among the works it cites.
“Leveraging weakly supervised data to improve end-to-end speech-to-text translation,”
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu, · 2019
Later among the works it cites.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhi-Hua Zhou, · 2018
Cited alongside, same era.
“Large-scale weakly supervised audio classification using gated convolutional neural network,”
Yong Xu, Qiuqiang Kong, Wenwu Wang, and Mark D Plumbley, · 2018
Cited alongside, same era.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
Yanpei Shi, Qiang Huang, and Thomas Hain, · 2020
Closest in time.
“H-vectors: Utterance-level speaker embedding using a hierarchical attention model,”
Yanpei Shi, Qiang Huang, and Thomas Hain, · 2020
Closest in time.
Sandesh V Katta, S Umesh, et al., · 2020
Closest in time.