Fetching the paper…
Reading the bibliography…
Many recent works on deep speaker embeddings train their feature extraction networks on large classification tasks, distinguishing between all speakers in a training set.
“Catastrophic forgetting in connectionist networks,”
Robert M. French, · 1999
Earlier work this paper cites.
“Front end factor analysis for speaker verification,”
Najim Dehak, Patrick J. Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Earlier work this paper cites.
“Improving deep neural networks for LVCSR using rectified linear units and dropout,”
George E. Dahl, Tara N. Sainath, and Geoffrey E. Hinton, · 2013
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“The dropout learning algorithm,”
Pierre Baldi and Peter Sadowski, · 2014
Earlier work this paper cites.
“An empirical analysis of dropout in piecewise linear networks,”
David Warde-Farley, Ian J. Goodfellow, Aaron Courville, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“An empirical investigation of catastrophic forgetting in gradient-based neural networks,”
Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Deep metric learning using triplet network,”
Elad Hoffer and Nir Ailon, · 2015
Earlier work this paper cites.
“Speaker recognition by machines and humans: A tutorial review,”
John H.L. Hansen and Taufiq Hasan, · 2015
Earlier work this paper cites.
“Visual domain adaptation: A survey of recent advances,”
V. M. Patel, R. Gopalan, R. Li, and R. Chellappa, · 2015
Earlier work this paper cites.
“End-to-end text-dependent speaker verification,”
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer, · 2016
Earlier work this paper cites.
“The speakers in the wild (SITW) speaker recognition database,”
Mitchell McLaren, Luciana Ferrer, Diego Castan, and Aaron Lawson, · 2016
Cited alongside, same era.
“Deep Neural Network Embeddings for Text-Independent Speaker Verification,”
David Snyder, Daniel Garcia-Romero, Daniel Povey, and Sanjeev Khudanpur, · 2017
Cited alongside, same era.
“Generalized End-to-End Loss for Speaker Verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2017
Cited alongside, same era.
“SphereFace: Deep hypersphere embedding for face recognition,”
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song, · 2017
Cited alongside, same era.
“Tacotron: Towards end-To-end speech synthesis,”
Yuxuan Wang, R. J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, and Quoc Le, · 2017
Cited alongside, same era.
“Model-agnostic meta-learning for fast adaptation of deep networks,”
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Later among the works it cites.
“Cosface: Large margin cosine loss for deep face recognition,”
Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu, · 2018
Later among the works it cites.
“Attentive statistics pooling for deep speaker embedding,”
Koji Okabe, Takafumi Koshinaka, and Koichi Shinoda, · 2018
Later among the works it cites.
“Bayesian HMM based x-vector clustering for Speaker Diarization,”
Mireia Diez, Lukas Burget, Shuai Wang, Johan Rohdin, and Honza Cernocký, · 2019
Later among the works it cites.
“Utterance-level aggregation for speaker recognition in the wild,”
W. Xie, A. Nagrani, J. S. Chung, and A. Zisserman, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chelsea Finn, Pieter Abbeel, and Sergey Levine, · 2017
Cited alongside, same era.
“VoxCeleb: A large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Cited alongside, same era.
“X-vectors: robust DNN embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge,”
Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jesús Villalba, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Additive margin softmax for face verification,”
Feng Wang, Jian Cheng, Weiyang Liu, and Haijun Liu, · 2018
Cited alongside, same era.
“A systematic study of the class imbalance problem in convolutional neural networks,”
Mateusz Buda, Atsuto Maki, and Maciej A. Mazurowski, · 2018
Cited alongside, same era.
“Cost-sensitive learning of deep feature representations from imbalanced data,”
Salman H. Khan, Munawar Hayat, Mohammed Bennamoun, Ferdous A. Sohel, and Roberto Togneri, · 2018
Cited alongside, same era.
Yun Tang, Guo-Hong Ding, Jing Huang, Xiaodong He, and Bowen Zhou, · 2019
Later among the works it cites.
“Arcface: Additive angular margin loss for deep face recognition,”
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou, · 2019
Later among the works it cites.
“AdaCos: Adaptively scaling cosine logits for effectively learning deep face representations,”
Xiao Zhang, Rui Zhao, Yu Qiao, Xiaogang Wang, and Hongsheng Li, · 2019
Later among the works it cites.
“Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML,” 2019
Aniruddh Raghu, Maithra Raghu, Samy Bengio, and Oriol Vinyals, · 2019
Later among the works it cites.
“Deep imbalanced learning for face recognition and attribute prediction,”
Chen Huang, Yining Li, Change Loy Chen, and Xiaoou Tang, · 2019
Later among the works it cites.
“Learning from less data: A unified data subset selection and active learning framework for computer vision,”
V. Kaushal, R. Iyer, S. Kothawade, R. Mahadev, K. Doctor, and G. Ramakrishnan, · 2019
Later among the works it cites.
“State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and Speakers in the Wild evaluations,”
Jesús Villalba, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Leibny Paola García-Perera, Fred Richardson, Réda Dehak, Pedro A. Torres-Carrasquillo, and Najim Dehak, · 2020
Closest in time.