Fetching the paper…
Reading the bibliography…
In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level.
“Improved end-of-query detection for streaming speech recognition,”
Matt Shannon, Gabor Simko, Shuo-Yiin Chang, and Carolina Parada, · 1913
Earlier work this paper cites.
“Multi-style training for robust isolated-word speech recognition,”
Richard Lippmann, Edward Martin, and D Paul, · 1987
Earlier work this paper cites.
“Recall, precision and average precision,”
Mu Zhu, · 2004
Earlier work this paper cites.
“Voice activity detection using harmonic frequency components in likelihood ratio test,”
Lee Ngee Tan, Bengt J Borgstrom, and Abeer Alwan, · 2010
Earlier work this paper cites.
“All for one: Feature combination for highly channel-degraded speech activity detection,”
Martin Graciarena, Abeer Alwan, Dan Ellis, Horacio Franco, Luciana Ferrer, John HL Hansen, Adam Janin, Byung Suk Lee, Yun Lei, Vikramjit Mitra, et al., · 2013
Earlier work this paper cites.
“Unsupervised speech activity detection using voicing measures and perceptual spectral flux,”
Seyed Omid Sadjadi and John HL Hansen, · 2013
Earlier work this paper cites.
“Real-life voice activity detection with lstm recurrent neural networks and an application to hollywood movies,”
Florian Eyben, Felix Weninger, Stefano Squartini, and Björn Schuller, · 2013
Earlier work this paper cites.
“Small-footprint keyword spotting using deep neural networks,”
Guoguo Chen, Carolina Parada, and Georg Heigold, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Improvements to the IBM speech activity detection system for the DARPA RATS program,”
Samuel Thomas, George Saon, Maarten Van Segbroeck, and Shrikanth S Narayanan, · 2015
Earlier work this paper cites.
“Voice activity detection: Merging source and filter-based information,”
Thomas Drugman, Yannis Stylianou, Yusuke Kida, and Masami Akamine, · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Cited alongside, same era.
“TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015,
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al., · 2015
Cited alongside, same era.
“On the efficient representation and execution of deep acoustic models,”
Raziel Alvarez, Rohit Prabhavalkar, and Anton Bakhtin, · 2016
Cited alongside, same era.
“Endpoint detection using grid long short-term memory networks for streaming speech recognition,”
Shuo-Yiin Chang, Bo Li, Tara N Sainath, Gabor Simko, and Carolina Parada, · 2017
Cited alongside, same era.
“Direct modeling of raw audio with dnns for wake word detection,”
Kenichi Kumatani, Sankaran Panchapagesan, Minhua Wu, Minjae Kim, Nikko Strom, Gautam Tiwari, and Arindam Mandai, · 2017
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“Speaker diarization with LSTM,”
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno, · 2018
Later among the works it cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, Yonghui Wu, et al., · 2018
Later among the works it cites.
“Wavenet based low rate speech coding,”
W Bastiaan Kleijn, Felicia SC Lim, Alejandro Luebs, Jan Skoglund, Florian Stimberg, Quan Wang, and Thomas C Walters, · 2018
Later among the works it cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Deep speaker: an end-to-end neural speaker embedding system,”
Chao Li, Xiaokong Ma, Bing Jiang, Xiangang Li, Xuewei Zhang, Xiao Liu, Ying Cao, Ajay Kannan, and Zhenyao Zhu, · 2017
Cited alongside, same era.
“Voxceleb: a large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Cited alongside, same era.
“A study on data augmentation of reverberant speech for robust speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L Seltzer, and Sanjeev Khudanpur, · 2017
Cited alongside, same era.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
Chanwoo Kim, Ananya Misra, Kean Chin, Thad Hughes, Arun Narayanan, Tara Sainath, and Michiel Bacchiani, · 2017
Cited alongside, same era.
“Temporal modeling using dilated convolution and gating for voice-activity-detection,”
Shuo-Yiin Chang, Bo Li, Gabor Simko, Tara N Sainath, Anshuman Tripathi, Aäron van den Oord, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2018
Cited alongside, same era.
Aonan Zhang, Quan Wang, Zhenyao Zhu, John Paisley, and Chong Wang, · 2019
Closest in time.
“Asvspoof 2019: a large-scale public database of synthetic, converted and replayed speech,”
Xin Wang, Junichi Yamagishi, Massimiliano Todisco, Hector Delgado, Andreas Nautsch, Nicholas Evans, Md Sahidullah, Ville Vestman, Tomi Kinnunen, Kong Aik Lee, et al., · 2019
Closest in time.
“VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2019
Closest in time.
“Direct Speech-to-Speech Translation with a Sequence-to-Sequence Model,”
Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu, · 2019
Closest in time.
“Sample efficient adaptive text-to-speech,”
Yutian Chen, Yannis Assael, Brendan Shillingford, David Budden, Scott Reed, Heiga Zen, Quan Wang, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, and Nando de Freitas, · 2019
Closest in time.
“Tuplemax loss for language identification,”
Li Wan, Prashant Sridhar, Yang Yu, Quan Wang, and Ignacio Lopez Moreno, · 2019
Closest in time.