Fetching the paper…
Reading the bibliography…
The field of speech processing has undergone a transformative shift with the advent of deep learning.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
The cocktail party problem
Simon Haykin and Zhe Chen. 2005 · 1902
Earlier work this paper cites.
Structured Pruning of Large Language Models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019c · 1910
Earlier work this paper cites.
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen. 2017 · 1913
Earlier work this paper cites.
End-to-end generation of talking faces from noisy speech. In ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 1948–1952
Sefik Emre Eskimez, Ross K Maddox, Chenliang Xu, and Zhiyao Duan. 2020 · 1952
Earlier work this paper cites.
A comparative performance study of several pitch detection algorithms
Lawrence Rabiner, Md Cheng, A Rosenberg, and C McGonegal. 1976 · 1976
Earlier work this paper cites.
All-pole modeling of degraded speech
Jae Lim and Alan Oppenheim. 1978 · 1978
Earlier work this paper cites.
Suppression of acoustic noise in speech using spectral subtraction
Steven Boll. 1979 · 1979
Earlier work this paper cites.
A hybrid time-frequency domain articulatory speech synthesizer
Man Sondhi and J. Schroeter. 1987 · 1987
Earlier work this paper cites.
A tutorial on hidden Markov models and selected applications in speech recognition
Lawrence R Rabiner. 1989 · 1989
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990
Charles T Hemphill, John J Godfrey, and George R Doddington. 1990 · 1990
Earlier work this paper cites.
A Bayesian estimation approach for speech enhancement using hidden Markov models
Yariv Ephraim. 1992 · 1992
Earlier work this paper cites.
Timit acoustic phonetic continuous speech corpus
John S Garofolo. 1993 · 1993
Earlier work this paper cites.
Connectionist speech recognition: a hybrid approach . Vol. 247
Herve A Bourlard and Nelson Morgan. 1994 · 1994
Earlier work this paper cites.
Speech enhancement based on a priori signal to noise estimation. In 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings , Vol. 2. IEEE, 629–632
Pascal Scalart et al · 1996
Earlier work this paper cites.
Signal subspace methods for speech enhancement
Peter SK Hansen. 1997 · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal. 1997 · 1997
Earlier work this paper cites.
Practice of usage of spectral analysis for forensic speaker identification. In RLA2C 1998-Speaker Recognition and its Commercial and Forensic Applications . 136–140
Sergey Koval and Sergey Krynov. 2020 · 1998
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs. In 2001 IEEE international conference on acoustics, speech, and signal processing. Proceedings (Cat. No. 01CH37221) , Vol. 2. IEEE, 749–752
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra. 2001 · 2001
Earlier work this paper cites.
Speech recognition using SVMs
Nathan Smith and Mark Gales. 2001 · 2001
Earlier work this paper cites.
Evaluation of a noise-robust DSR front-end on Aurora databases. In Seventh International Conference on Spoken Language Processing
Duncan Macho, Laurent Mauuary, Bernhard Noé, Yan Ming Cheng, Doug Ealey, Denis Jouvet, Holly Kelleher, David Pearce, and Fabien Saadoun. 2002 · 2002
Earlier work this paper cites.
Phonetic speaker recognition with support vector machines
William Campbell, Joseph Campbell, Douglas Reynolds, Douglas Jones, and Timothy Leek. 2003 · 2003
Earlier work this paper cites.
Channel robust speaker verification via feature mapping. In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings.(ICASSP’03). , Vol. 2. IEEE, II–53
Douglas A Reynolds. 2003 · 2003
Earlier work this paper cites.
Pitch detection algorithm: autocorrelation method and AMDF. In Proceedings of the 3rd international symposium on communications and information technology , Vol. 2. 551–556
Li Tan and Montri Karnjanadecha. 2003 · 2003
Earlier work this paper cites.
The CMU Arctic speech databases. In Fifth ISCA workshop on speech synthesis
John Kominek and Alan W Black. 2004 · 2004
Earlier work this paper cites.
Channel compensation for SVM speaker recognition.. In Odyssey , Vol. 4. 219–226
Alex Solomonoff, Carl Quillen, and William M Campbell. 2004 · 2004
Earlier work this paper cites.
The AMI meeting corpus: A pre-announcement. In Machine Learning for Multimodal Interaction: Second International Workshop, MLMI 2005, Edinburgh, UK, July 11-13, 2005, Revised Selected Papers 2 . Springer, 28–39
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et al · 2005
Earlier work this paper cites.
Levinson-durbin algorithm
Paolo Castiglioni. 2005 · 2005
Earlier work this paper cites.
Real-time speaker identification and verification
Tomi Kinnunen, Evgeny Karpov, and Pasi Franti. 2005 · 2005
Earlier work this paper cites.
Advances in channel compensation for SVM speaker recognition. In Proceedings.(ICASSP’05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005. , Vol. 1. IEEE, I–629
Alex Solomonoff, William M Campbell, and Ian Boardman. 2005 · 2005
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao. 2006 · 2006
Earlier work this paper cites.
Within-class covariance normalization for SVM-based speaker recognition. In Ninth international conference on spoken language processing
Andrew O Hatch, Sachin Kajarekar, and Andreas Stolcke. 2006 · 2006
Earlier work this paper cites.
An overview of automatic speaker diarization systems
Sue E Tranter and Douglas A Reynolds. 2006 · 2006
Earlier work this paper cites.
Evaluation of objective quality measures for speech enhancement
Yi Hu and Philipos C Loizou. 2007 · 2007
Earlier work this paper cites.
The application of hidden Markov models in speech recognition
Mark Gales, Steve Young, et al · 2008
Earlier work this paper cites.
Supervised sequence labelling with recurrent neural networks
Kazuya Kawakami. 2008 · 2008
Earlier work this paper cites.
Speech enhancement using harmonic emphasis and adaptive comb filtering
Wen Jin, Xin Liu, Michael S Scordilis, and Lu Han. 2009 · 2009
Earlier work this paper cites.
The importance of phase in speech enhancement
Kuldip Paliwal, Kamil Wójcicki, and Benjamin Shannon. 2011 · 2011
Earlier work this paper cites.
HAWQV3: Dyadic Neural Network Quantization
Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W. Mahoney, and Kurt Keutzer. 2020 · 2011
Earlier work this paper cites.
Applying Convolutional Neural Networks concepts to hybrid NN-HMM model for speech recognition. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 4277–4280
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, and Gerald Penn. 2012 · 2012
Earlier work this paper cites.
Speaker diarization: A review of recent research
Xavier Anguera, Simon Bozonnet, Nicholas Evans, Corinne Fredouille, Gerald Friedland, and Oriol Vinyals. 2012 · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves. 2012 · 2012
Earlier work this paper cites.
Connectionist temporal classification
Alex Graves and Alex Graves. 2012 · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
The RSR2015: Database for text-dependent speaker verification using multiple pass-phrases. In Annual Conference of the International Speech Communication Association (Interspeech)
Anthony Larcher, Kong Aik Lee, Bin Ma, and Haizhou Li. 2012 · 2012
Earlier work this paper cites.
TED-LIUM: an Automatic Speech Recognition dedicated corpus.. In LREC . 125–129
Anthony Rousseau, Paul Deléglise, and Yannick Esteve. 2012 · 2012
Earlier work this paper cites.
Perceptual objective listening quality assessment (polqa), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment
John G Beerends, Christian Schmidmer, Jens Berger, Matthias Obermann, Raphael Ullmann, Joachim Pomy, and Michael Keyhl. 2013 · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing . Ieee, 6645–6649
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013 · 2013
Earlier work this paper cites.
Semi-supervised learning
Mohamed Farouk Abdel Hady and Friedhelm Schwenker. 2013 · 2013
Earlier work this paper cites.
Asgard: A portable architecture for multilingual dialogue systems. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 8386–8390
Jingjing Liu, Panupong Pasupat, Scott Cyphers, and Jim Glass. 2013 · 2013
Earlier work this paper cites.
Speech enhancement based on deep denoising autoencoder.. In Interspeech , Vol. 2013. 436–440
Xugang Lu, Yu Tsao, Shigeki Matsuda, and Chiori Hori. 2013 · 2013
Earlier work this paper cites.
Convolutional neural networks for speech recognition
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, Li Deng, Gerald Penn, and Dong Yu. 2014a · 2014
Earlier work this paper cites.
Convolutional Neural Networks for Speech Recognition
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, Li Deng, Gerald Penn, and Dong Yu. 2014b · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Robust speech recognition with speech enhanced deep neural networks. In Fifteenth annual conference of the international speech communication association
Jun Du, Qing Wang, Tian Gao, Yong Xu, Li-Rong Dai, and Chin-Hui Lee. 2014 · 2014
Earlier work this paper cites.
TTS synthesis with bidirectional LSTM based recurrent neural networks. In Fifteenth annual conference of the international speech communication association
Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong. 2014 · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks. In International conference on machine learning . PMLR, 1764–1772
Alex Graves and Navdeep Jaitly. 2014 · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014 · 2014
Earlier work this paper cites.
Nearest neighbor discriminant analysis for robust speaker recognition. In Fifteenth Annual Conference of the International Speech Communication Association
Seyed Omid Sadjadi, Jason Pelecanos, and Weizhong Zhu. 2014 · 2014
Earlier work this paper cites.
Deep neural networks for small footprint text-dependent speaker verification. In 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 4052–4056
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez. 2014a · 2014
Earlier work this paper cites.
Deep neural networks for small footprint text-dependent speaker verification. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 4052–4056
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez. 2014b · 2014
Earlier work this paper cites.
A regression approach to speech enhancement based on deep neural networks
Yong Xu, Jun Du, Li-Rong Dai, and Chin-Hui Lee. 2014 · 2014
Earlier work this paper cites.
Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks. In Sixteenth Annual Conference of the International Speech Communication Association
Zhuo Chen, Shinji Watanabe, Hakan Erdogan, and John R Hershey. 2015 · 2015
Earlier work this paper cites.
Describing multimedia content using attention-based encoder-decoder networks
Kyunghyun Cho, Aaron Courville, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
TCD-TIMIT: An audio-visual corpus of continuous speech
Naomi Harte and Eoin Gillen. 2015 · 2015
Earlier work this paper cites.
Convolutional neural networks for patient-specific ECG classification. In 2015 37th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) . IEEE, 2608–2611
Serkan Kiranyaz, Turker Ince, Ridha Hamila, and Moncef Gabbouj. 2015 · 2015
Earlier work this paper cites.
The RedDots data collection for speaker recognition. In Interspeech 2015
Kong Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David Van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma, et al · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5206–5210
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 815–823
Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning . PMLR, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Multi-Scale Context Aggregation by Dilated Convolutions
Fisher Yu and Vladlen Koltun. 2015 · 2015
Earlier work this paper cites.
A comparison of several computational auditory scene analysis (CASA) techniques for monaural speech segregation
Jihen Zeremdini, Mohamed Anouar Ben Messaoud, and Aicha Bouzid. 2015 · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin. In International conference on machine learning . PMLR, 173–182
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Lip reading in the wild. In Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13 . Springer, 87–103
Joon Son Chung and Andrew Zisserman. 2017 · 2016
Earlier work this paper cites.
SNR-Aware Convolutional Neural Network Modeling for Speech Enhancement.. In Interspeech . 3768–3772
Szu-Wei Fu, Yu Tsao, and Xugang Lu. 2016 · 2016
Earlier work this paper cites.
Deep reconstruction-classification networks for unsupervised domain adaptation. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14 . Springer, 597–613
Muhammad Ghifary, W Bastiaan Kleijn, Mengjie Zhang, David Balduzzi, and Wen Li. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep clustering: Discriminative embeddings for segmentation and separation. In 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 31–35
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe. 2016 · 2016
Earlier work this paper cites.
Neural machine translation in linear time
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
The speakers in the wild (SITW) speaker recognition database.. In Interspeech . 818–822
Mitchell McLaren, Luciana Ferrer, Diego Castan, and Aaron Lawson. 2016 · 2016
Earlier work this paper cites.
SampleRNN: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Purely sequence-trained neural networks for ASR based on lattice-free MMI.. In Interspeech . 2751–2755
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur. 2016 · 2016
Earlier work this paper cites.
Novel deep autoencoder features for non-intrusive speech quality assessment. In 2016 24th European Signal Processing Conference (EUSIPCO) . IEEE, 2315–2319
Meet H Soni and Hemant A Patil. 2016 · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Robust spoken language understanding for house service robots
Andrea Vanzo, Danilo Croce, Emanuele Bastianelli, Roberto Basili, and Daniele Nardi. 2016 · 2016
Earlier work this paper cites.
Survey on the attention based RNN model and its applications in computer vision
Feng Wang and David MJ Tax. 2016 · 2016
Earlier work this paper cites.
Automatic speech recognition . Vol. 1
Dong Yu and Li Deng. 2016 · 2016
Earlier work this paper cites.
Real-time vibration-based structural damage detection using one-dimensional convolutional neural networks
Osama Abdeljaber, Onur Avci, Serkan Kiranyaz, Moncef Gabbouj, and Daniel J Inman. 2017 · 2017
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech. In International conference on machine learning . PMLR, 195–204
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Earlier work this paper cites.
Chung-Cheng Chiu and Colin Raffel. 2017 · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks. In International conference on machine learning . PMLR, 933–941
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Earlier work this paper cites.
Improved speech reconstruction from silent video. In Proceedings of the IEEE International Conference on Computer Vision Workshops . 455–462
Ariel Ephrat, Tavi Halperin, and Shmuel Peleg. 2017 · 2017
Earlier work this paper cites.
Vid2speech: speech reconstruction from silent video. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5095–5099
Ariel Ephrat and Shmuel Peleg. 2017 · 2017
Earlier work this paper cites.
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg. 2017 · 2017
Earlier work this paper cites.
Deep voice 2: Multi-speaker neural text-to-speech
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. 2017 · 2017
Earlier work this paper cites.
Streaming small-footprint keyword spotting using sequence-to-sequence models. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 474–481
Yanzhang He, Rohit Prabhavalkar, Kanishka Rao, Wei Li, Anton Bakhtin, and Ian McGraw. 2017 · 2017
Earlier work this paper cites.
Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, and Hsin-Min Wang. 2017 · 2017
Earlier work this paper cites.
Squeeze-and-Excitation Networks
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu. 2017 · 2017
Earlier work this paper cites.
The LJ Speech Dataset
Keith Ito and Linda Johnson. 2017 · 2017
Earlier work this paper cites.
Acoustic Modeling for Google Home.. In Interspeech . 399–403
Bo Li, Tara N Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K Chin, et al · 2017
Earlier work this paper cites.
Deep voice 3: Scaling text-to-speech with convolutional sequence learning
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2017 · 2017
Earlier work this paper cites.
A Comparison of sequence-to-sequence models for speech recognition.. In Interspeech . 939–943
Rohit Prabhavalkar, Kanishka Rao, Tara N Sainath, Bo Li, Leif Johnson, and Navdeep Jaitly. 2017 · 2017
Earlier work this paper cites.
Online and linear-time attention by enforcing monotonic alignments. In International conference on machine learning . PMLR, 2837–2846
Colin Raffel, Minh-Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck. 2017 · 2017
Earlier work this paper cites.
Recent advances in recurrent neural networks
Hojjat Salehinejad, Sharan Sankar, Joseph Barfett, Errol Colak, and Shahrokh Valaee. 2017 · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel. 2017 · 2017
Earlier work this paper cites.
Deep neural network embeddings for text-independent speaker verification.. In Interspeech , Vol. 2017. 999–1003
David Snyder, Daniel Garcia-Romero, Daniel Povey, and Sanjeev Khudanpur. 2017 · 2017
Earlier work this paper cites.
Lip reading sentences in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6447–6456
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
End-to-end text-independent speaker verification with triplet loss on short utterances.. In Interspeech . 1487–1491
Chunlei Zhang and Kazuhito Koishida. 2017 · 2017
Earlier work this paper cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun. 2018 · 2018
Earlier work this paper cites.
The fifth’CHiME’speech separation and recognition challenge: dataset, task and baselines
Jon Barker, Shinji Watanabe, Emmanuel Vincent, and Jan Trmal. 2018 · 2018
Earlier work this paper cites.
VoxCeleb2: Deep Speaker Recognition. In INTERSPEECH
J. S. Chung, A. Nagrani, and A. Zisserman. 2018 · 2018
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Chris Donahue, Julian McAuley, and Miller Puckette. 2018 · 2018
Earlier work this paper cites.
Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 5884–5888
Linhao Dong, Shuang Xu, and Bo Xu. 2018 · 2018
Earlier work this paper cites.
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. 2018 · 2018
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Training Pruned Neural Networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Earlier work this paper cites.
Few-shot learning with attention-based sequence-to-sequence models
Bertrand Higy and Peter Bell. 2018 · 2018
Earlier work this paper cites.
Audio-visual speech enhancement using multimodal deep convolutional neural networks
Jen-Cheng Hou, Syu-Siang Wang, Ying-Hui Lai, Yu Tsao, Hsiu-Wen Chang, and Hsin-Min Wang. 2018 · 2018
Earlier work this paper cites.
Hierarchical Generative Modeling for Controllable Speech Synthesis. In International Conference on Learning Representations
Wei-Ning Hsu, Yu Zhang, Ron J Weiss, Heiga Zen, Yonghui Wu, Yuxuan Wang, Yuan Cao, Ye Jia, Zhifeng Chen, Jonathan Shen, et al · 2018
Earlier work this paper cites.
Reinforcement learning of speech recognition system based on policy gradient and hypothesis selection. In 2018 ieee international conference on acoustics, speech and signal processing (icassp) . IEEE, 5759–5763
Taku Kala and Takahiro Shinozaki. 2018 · 2018
Earlier work this paper cites.
Efficient neural audio synthesis. In International Conference on Machine Learning . PMLR, 2410–2419
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu. 2018 · 2018
Earlier work this paper cites.
FloWaveNet: A generative flow for raw audio
Sungwon Kim, Sang-gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon. 2018 · 2018
Earlier work this paper cites.
Emorl: continuous acoustic emotion classification using deep reinforcement learning. In 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 4445–4450
Egor Lakomkin, Mohammad Ali Zamani, Cornelius Weber, Sven Magg, and Stefan Wermter. 2018 · 2018
Earlier work this paper cites.
Time-Frequency Networks for Audio Super-Resolution. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 646–650
Teck Yian Lim, Raymond A. Yeh, Yijia Xu, Minh N. Do, and Mark Hasegawa-Johnson. 2018 · 2018
Earlier work this paper cites.
Towards achieving robust universal neural vocoding
Jaime Lorenzo-Trueba, Thomas Drugman, Javier Latorre, Thomas Merritt, Bartosz Putrycz, Roberto Barra-Chicote, Alexis Moinet, and Vatsal Aggarwal. 2018 · 2018
Earlier work this paper cites.
Real-time single-channel dereverberation and separation with time-domain audio separation network.. In Interspeech . 342–346
Yi Luo and Nima Mesgarani. 2018 · 2018
Earlier work this paper cites.
Y-Net: joint segmentation and classification for diagnosis of breast biopsy images. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Proceedings, Part II 11 . Springer, 893–901
Sachin Mehta, Ezgi Mercan, Jamen Bartlett, Donald Weaver, Joann G Elmore, and Linda Shapiro. 2018 · 2018
Earlier work this paper cites.
Unspeech: Unsupervised speech context embeddings
Benjamin Milde and Chris Biemann. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018a · 2018
Earlier work this paper cites.
Clarinet: Parallel wave generation in end-to-end text-to-speech
Wei Ping, Kainan Peng, and Jitong Chen. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Speaker recognition from raw waveform with sincnet. In 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 1021–1028
Mirco Ravanelli and Yoshua Bengio. 2018 · 2018
Earlier work this paper cites.
Voices obscured in complex environmental settings (voices) corpus
Colleen Richey, Maria A Barrios, Zeb Armstrong, Chris Bartels, Horacio Franco, Martin Graciarena, Aaron Lawson, Mahesh Kumar Nandwana, Allen Stauffer, Julien van Hout, et al · 2018
Earlier work this paper cites.
First DIHARD challenge evaluation plan
Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristia, Jun Du, Sriram Ganapathy, and Mark Liberman. 2018 · 2018
Earlier work this paper cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 4779–4783
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Earlier work this paper cites.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron. In international conference on machine learning . PMLR, 4693–4702
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous. 2018 · 2018
Earlier work this paper cites.
X-vectors: Robust dnn embeddings for speaker recognition. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5329–5333
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. 2018 · 2018
Earlier work this paper cites.
Talking face generation by conditional recurrent adversarial network
Yang Song, Jingwen Zhu, Dawei Li, Xiaolong Wang, and Hairong Qi. 2018 · 2018
Earlier work this paper cites.
Wave-u-net: A multi-scale neural network for end-to-end audio source separation
Daniel Stoller, Sebastian Ewert, and Simon Dixon. 2018 · 2018
Earlier work this paper cites.
Sequence-to-sequence ASR optimization via reinforcement learning. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5829–5833
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2018 · 2018
Earlier work this paper cites.
Petar Veličković, William Fedus, William L Hamilton, Pietro Liò, Yoshua Bengio, and R Devon Hjelm. 2018 · 2018
Earlier work this paper cites.
Audio source separation and speech enhancement
Emmanuel Vincent, Tuomas Virtanen, and Sharon Gannot. 2018 · 2018
Earlier work this paper cites.
Generalized end-to-end loss for speaker verification. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4879–4883
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno. 2018 · 2018
Earlier work this paper cites.
Speaker diarization with LSTM. In 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5239–5243
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno. 2018a · 2018
Earlier work this paper cites.
Unsupervised Domain Adaptation via Domain Adversarial Training for Speaker Recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 4889–4893
Qing Wang, Wei Rao, Sining Sun, Leib Xie, Eng Siong Chng, and Haizhou Li. 2018c · 2018
Earlier work this paper cites.
A bi-model based rnn semantic frame parsing model for intent detection and slot filling
Yu Wang, Yilin Shen, and Hongxia Jin. 2018d · 2018
Earlier work this paper cites.
Alternative objective functions for deep clustering. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 686–690
Zhong-Qiu Wang, Jonathan Le Roux, and John R Hershey. 2018b · 2018
Earlier work this paper cites.
Speech commands: A dataset for limited-vocabulary speech recognition
Pete Warden. 2018 · 2018
Earlier work this paper cites.
Improving Attention Based Sequence-to-Sequence Models for End-to-End English Conversational Speech Recognition.. In Interspeech . 761–765
Chao Weng, Jia Cui, Guangsen Wang, Jun Wang, Chengzhu Yu, Dan Su, and Dong Yu. 2018 · 2018
Earlier work this paper cites.
Joint slot filling and intent detection via capsule neural networks
Chenwei Zhang, Yaliang Li, Nan Du, Wei Fan, and Philip S Yu. 2018a · 2018
Earlier work this paper cites.
Forward attention in sequence-to-sequence acoustic modeling for speech synthesis. In 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP) . IEEE, 4789–4793
Jing-Xuan Zhang, Zhen-Hua Ling, and Li-Rong Dai. 2018b · 2018
Earlier work this paper cites.
L2-ARCTIC: A non-native English speech corpus.. In Interspeech . 2783–2787
Guanlong Zhao, Sinem Sonsaat, Alif Silpachai, Ivana Lucic, Evgeny Chukharev-Hudilainen, John Levis, and Ricardo Gutierrez-Osuna. 2018 · 2018
Earlier work this paper cites.
Arbitrary talking face generation via attentional audio-visual coherence learning
Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng, and Ran He. 2018 · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. 2019 · 2019
Earlier work this paper cites.
Forget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition.. In Interspeech . 2618–2622
Kartik Audhkhasi, George Saon, Zoltán Tüske, Brian Kingsbury, and Michael Picheny. 2019 · 2019
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed. 2019a · 2019
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2019b · 2019
Earlier work this paper cites.
A comparative study on end-to-end speech to text translation. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 792–799
Parnia Bahar, Tobias Bieschke, and Hermann Ney. 2019 · 2019
Earlier work this paper cites.
Semi-supervised sequence-to-sequence ASR using unpaired speech and text
Murali Karthick Baskar, Shinji Watanabe, Ramon Astudillo, Takaaki Hori, Lukáš Burget, and Jan Černockỳ. 2019 · 2019
Earlier work this paper cites.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan. 2019a · 2019
Earlier work this paper cites.
Temporal FiLM: Capturing Long-Range Sequence Dependencies with Feature-Wise Modulations
Sawyer Birnbaum, Volodymyr Kuleshov, Zayd Enam, Pang Wei W Koh, and Stefano Ermon. 2019 · 2019
Earlier work this paper cites.
On robustness of unsupervised domain adaptation for speaker recognition. In Interspeech
Pierre-Michel Bousquet and Mickael Rouvier. 2019 · 2019
Earlier work this paper cites.
Non-intrusive speech quality prediction using modulation energies and lstm-network
Benjamin Cauchi, Kai Siedenburg, Joao F Santos, Tiago H Falk, Simon Doclo, and Stefan Goetze. 2019 · 2019
Earlier work this paper cites.
Bert for joint intent classification and slot filling
Qian Chen, Zhu Zhuo, and Wen Wang. 2019b · 2019
Earlier work this paper cites.
Ju-chieh Chou, Cheng-chieh Yeh, and Hung-yi Lee. 2019 · 2019
Earlier work this paper cites.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 2019
Earlier work this paper cites.
Improving post-filtering of artificial speech using pre-trained LSTM neural networks
Marvin Coto-Jiménez. 2019 · 2019
Earlier work this paper cites.
Efficient keyword spotting using dilated convolutions and gating. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6351–6355
Alice Coucke, Mohammed Chlieh, Thibault Gisselbrecht, David Leroy, Mathieu Poumeyrol, and Thibaut Lavril. 2019 · 2019
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4690–4699
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. 2019 · 2019
Earlier work this paper cites.
Adapting transformer to end-to-end spoken language translation
Mattia A Di Gangi, Matteo Negri, and Marco Turchi. 2019 · 2019
Earlier work this paper cites.
Explicit alignment of text and speech encodings for attention-based end-to-end speech recognition. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 913–919
Jennifer Drexler and James Glass. 2019 · 2019
Earlier work this paper cites.
Application of different statistical tests for validation of synthesized speech parameterized by cepstral coefficients and lsp
Carlos Franco-Galván, Abel Herrera-Camacho, and Boris Escalante-Ramírez. 2019 · 2019
Earlier work this paper cites.
End-to-end neural speaker diarization with self-attention. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 296–303
Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Yawen Xue, Kenji Nagamatsu, and Shinji Watanabe. 2019 · 2019
Earlier work this paper cites.
Attention wave-u-net for speech enhancement. In 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 249–253
Ritwik Giri, Umut Isik, and Arvindh Krishnaswamy. 2019 · 2019
Earlier work this paper cites.
Streaming end-to-end speech recognition for mobile devices. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6381–6385
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al · 2019
Earlier work this paper cites.
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. In International Conference on Learning Representations
Dan Hendrycks and Thomas Dietterich. 2019 · 2019
Earlier work this paper cites.
Deep domain adaptation for anti-spoofing in speaker verification systems
Ivan Himawan, Fernando Villavicencio, Sridha Sridharan, and Clinton Fookes. 2019 · 2019
Earlier work this paper cites.
Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factorization. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5901–5905
Wei-Ning Hsu, Yu Zhang, Ron J Weiss, Yu-An Chung, Yuxuan Wang, Yonghui Wu, and James Glass. 2019 · 2019
Earlier work this paper cites.
Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, and Tomoki Toda. 2019 · 2019
Earlier work this paper cites.
Parkinson disease detection using deep neural networks. In 2019 twelfth international conference on contemporary computing (IC3) . IEEE, 1–4
Anubhav Johri, Ashish Tripathi, et al · 2019
Earlier work this paper cites.
Jee-weon Jung, Hee-Soo Heo, Ju-ho Kim, Hye-jin Shim, and Ha-Jin Yu. 2019 · 2019
Earlier work this paper cites.
ACVAE-VC: Non-parallel voice conversion with auxiliary classifier variational autoencoder
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, and Nobukatsu Hojo. 2019 · 2019
Earlier work this paper cites.
Cyclegan-vc2: Improved cyclegan-based non-parallel voice conversion. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6820–6824
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, and Nobukatsu Hojo. 2019 · 2019
Earlier work this paper cites.
An active learning paradigm for online audio-visual emotion recognition
Ioannis Kansizoglou, Loukas Bampis, and Antonios Gasteratos. 2019 · 2019
Earlier work this paper cites.
A comparative study on transformer vs rnn in speech applications. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 449–456
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al · 2019
Earlier work this paper cites.
CHiVE: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network. In International Conference on Machine Learning . PMLR, 3331–3340
Tom Kenter, Vincent Wan, Chun-An Chan, Rob Clark, and Jakub Vit. 2019 · 2019
Earlier work this paper cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville. 2019 · 2019
Earlier work this paper cites.
An Effective Style Token Weight Control Technique for End-to-End Emotional Speech Synthesis
Ohsung Kwon, Inseon Jang, ChungHyun Ahn, and Hong-Goo Kang. 2019 · 2019
Earlier work this paper cites.
The CORAL+ algorithm for unsupervised domain adaptation of PLDA. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5821–5825
Kong Aik Lee, Qiongqiong Wang, and Takafumi Koshinaka. 2019 · 2019
Earlier work this paper cites.
Temporal convolutional networks for speech and music detection in radio broadcast. In 20th International Society for Music Information Retrieval Conference, ISMIR 2019, 4-8 November 2019 . International Society for Music Information Retrieval
Quentin Lemaire and Andre Holzapfel. 2019 · 2019
Earlier work this paper cites.
Federated learning for keyword spotting. In ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 6341–6345
David Leroy, Alice Coucke, Thibaut Lavril, Thibault Gisselbrecht, and Joseph Dureau. 2019 · 2019
Earlier work this paper cites.
Single channel speech enhancement using temporal convolutional recurrent neural networks. In 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . IEEE, 896–900
Jingdong Li, Hui Zhang, Xueliang Zhang, and Changliang Li. 2019b · 2019
Earlier work this paper cites.
Speech enhancement using forked generative adversarial networks with spectral subtraction
Ju Lin, Sufeng Niu, Zice Wei, Xiang Lan, Adriaan J Wijngaarden, Melissa C Smith, and Kuang-Ching Wang. 2019 · 2019
Earlier work this paper cites.
Exploiting unlabeled data in cnns by self-supervised learning to rank
Xialei Liu, Joost Van De Weijer, and Andrew D Bagdanov. 2019 · 2019
Earlier work this paper cites.
Speech model pre-training for end-to-end spoken language understanding
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio. 2019 · 2019
Earlier work this paper cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani. 2019 · 2019
Earlier work this paper cites.
Neural TTS stylization with adversarial and collaborative games. In International Conference on Learning Representations
Shuang Ma, Daniel Mcduff, and Yale Song. 2019 · 2019
Earlier work this paper cites.
Transformers with convolutional context for asr
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Combining Speaker Recognition and Metric Learning for Speaker-Dependent Representation Learning.. In INTERSPEECH . 4015–4019
Joao Monteiro, Md Jahangir Alam, and Tiago H Falk. 2019 · 2019
Earlier work this paper cites.
Improving transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration. In Proc. Interspeech , Vol. 2019
Tomohiro Nakatani. 2019 · 2019
Earlier work this paper cites.
Speech recognition using deep neural networks: A systematic review
Ali Bou Nassif, Ismail Shahin, Imtinan Attili, Mohammad Azzeh, and Khaled Shaalan. 2019 · 2019
Earlier work this paper cites.
Cycle-gans for domain adaptation of acoustic features for speaker recognition. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6206–6210
Phani Sankar Nidadavolu, Jesús Villalba, and Najim Dehak. 2019 · 2019
Earlier work this paper cites.
A review of deep learning based speech synthesis
Yishuang Ning, Sheng He, Zhiyong Wu, Chunxiao Xing, and Liang-Jie Zhang. 2019 · 2019
Earlier work this paper cites.
A novel bi-directional interrelated model for joint intent detection and slot filling
Peiqing Niu, Zhongfu Chen, Meina Song, et al · 2019
Earlier work this paper cites.
Tacotron-Based Acoustic Model Using Phoneme Alignment for Practical Neural Text-to-Speech Systems. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . 214–221
Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, and Hisashi Kawai. 2019b · 2019
Earlier work this paper cites.
Improving deep models of speech quality prediction through voice activity detection and entropy-based measures. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 636–640
Jasper Ooster and Bernd T Meyer. 2019 · 2019
Cited alongside, same era.
TCNN: Temporal convolutional neural network for real-time speech enhancement in the time domain. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6875–6879
Ashutosh Pandey and DeLiang Wang. 2019 · 2019
Cited alongside, same era.
Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap
Tae Jin Park, Kyu J Han, Manoj Kumar, and Shrikanth Narayanan. 2019 · 2019
Cited alongside, same era.
Learning problem-agnostic speech representations from multiple self-supervised tasks
Santiago Pascual, Mirco Ravanelli, Joan Serra, Antonio Bonafonte, and Yoshua Bengio. 2019 · 2019
Cited alongside, same era.
Nu-wave: A diffusion probabilistic model for neural audio upsampling
Junhyeok Lee and Seungu Han. 2021 · 2021
Later among the works it cites.
A better and faster end-to-end model for streaming asr. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5634–5638
Bo Li, Anmol Gulati, Jiahui Yu, Tara N Sainath, Chung-Cheng Chiu, Arun Narayanan, Shuo-Yiin Chang, Ruoming Pang, Yanzhang He, James Qin, et al · 2021
Later among the works it cites.
Real-time monaural speech enhancement with short-time discrete cosine transform
Qinglong Li, Fei Gao, Haixin Guan, and Kaichi Ma. 2021a · 2021
Later among the works it cites.
Confidence estimation for attention-based sequence-to-sequence models for speech recognition. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6388–6392
Qiujia Li, David Qiu, Yu Zhang, Bo Li, Yanzhang He, Philip C Woodland, Liangliang Cao, and Trevor Strohman. 2021c · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A hybrid of deep CNN and bidirectional LSTM for automatic speech recognition
Vishal Passricha and Rajesh Kumar Aggarwal. 2019 · 2019
Cited alongside, same era.
TTS skins: Speaker conversion via ASR
Adam Polyak, Lior Wolf, and Yaniv Taigman. 2019 · 2019
Cited alongside, same era.
Waveglow: A flow-based generative network for speech synthesis. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3617–3621
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Dual supervised learning for non-native speech recognition
Kacper Radzikowski, Robert Nowak, Le Wang, and Osamu Yoshie. 2019 · 2019
Cited alongside, same era.
The pytorch-kaldi speech recognition toolkit. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6465–6469
Mirco Ravanelli, Titouan Parcollet, and Yoshua Bengio. 2019 · 2019
Cited alongside, same era.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
The second dihard diarization challenge: Dataset, task, and baselines
Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristia, Jun Du, Sriram Ganapathy, and Mark Liberman. 2019 · 2019
Cited alongside, same era.
Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Association for Computational Linguistics, Online, 4582–4597
Xiang Lisa Li and Percy Liang. 2021 · 2021
Later among the works it cites.
Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks
Ju Lin, Adriaan J. de Lind van Wijngaarden, Kuang-Ching Wang, and Melissa C. Smith. 2021b · 2021
Later among the works it cites.
S2vc: a framework for any-to-any voice conversion with self-supervised pretrained representations
Jheng-hao Lin, Yist Y Lin, Chung-Ming Chien, and Hung-yi Lee. 2021a · 2021
Later among the works it cites.
Tera: Self-supervised learning of transformer encoder representation for speech
Andy T Liu, Shang-Wen Li, and Hung-yi Lee. 2021b · 2021
Later among the works it cites.
Improving RNN transducer based ASR with auxiliary tasks. In 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 172–179
Chunxi Liu, Frank Zhang, Duc Le, Suyoun Kim, Yatharth Saraf, and Geoffrey Zweig. 2021f · 2021
Later among the works it cites.
Expressive TTS Training With Frame and Style Reconstruction Loss
Rui Liu, Berrak Sisman, Guanglai Gao, and Haizhou Li. 2021d · 2021
Later among the works it cites.
Graphspeech: Syntax-aware graph attention network for neural speech synthesis. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6059–6063
Rui Liu, Berrak Sisman, and Haizhou Li. 2021c · 2021
Later among the works it cites.
Any-to-many voice conversion with location-relative sequence-to-sequence modeling
Songxiang Liu, Yuewen Cao, Disong Wang, Xixin Wu, Xunying Liu, and Helen Meng. 2021a · 2021
Later among the works it cites.
Delightfultts: The microsoft speech synthesis system for blizzard challenge 2021
Yanqing Liu, Zhihang Xu, Gang Wang, Kuan Chen, Bohan Li, Xu Tan, Jinzhu Li, Lei He, and Sheng Zhao. 2021e · 2021
Later among the works it cites.
A Study on Speech Enhancement Based on Diffusion Probabilistic Model. In 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . 659–666
Yen-Ju Lu, Yu Tsao, and Shinji Watanabe. 2021 · 2021
Later among the works it cites.
FlowVocoder: A small Footprint Neural Vocoder based Normalizing flow for Speech Synthesis
Manh Luong and Viet Anh Tran. 2021 · 2021
Later among the works it cites.
Lira: Learning visual speech representations from audio through self-supervision
Pingchuan Ma, Rodrigo Mira, Stavros Petridis, Björn W Schuller, and Maja Pantic. 2021 · 2021
Later among the works it cites.
Somshubra Majumdar, Jagadeesh Balam, Oleksii Hrinchuk, Vitaly Lavrukhin, Vahid Noroozi, and Boris Ginsburg. 2021 · 2021
Later among the works it cites.
NORESQA: A framework for speech quality assessment using non-matching references
Pranay Manocha, Buye Xu, and Anurag Kumar. 2021 · 2021
Later among the works it cites.
S-vectors and TESA: Speaker embeddings and a speaker authenticator based on transformer encoder
Narla John Metilda Sagaya Mary, Srinivasan Umesh, and Sandesh Varadaraju Katta. 2021 · 2021
Later among the works it cites.
Efficienttts: An efficient and high-quality text-to-speech architecture. In International Conference on Machine Learning . PMLR, 7700–7709
Chenfeng Miao, Liang Shuang, Zhengchen Liu, Chen Minchuan, Jun Ma, Shaojun Wang, and Jing Xiao. 2021 · 2021
Later among the works it cites.
An overview of deep-learning-based audio-visual speech enhancement and separation
Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu, Meng Yu, Dong Yu, and Jesper Jensen. 2021 · 2021
Later among the works it cites.
A cappella: Audio-visual singing voice separation
Juan F Montesinos, Venkatesh S Kadandale, and Gloria Haro. 2021 · 2021
Later among the works it cites.
Stylemelgan: An efficient high-fidelity adversarial vocoder with temporal adaptive normalization. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6034–6038
Ahmed Mustafa, Nicola Pia, and Guillaume Fuchs. 2021 · 2021
Later among the works it cites.
Yoshihiko Nankaku, Kenta Sumiya, Takenori Yoshimura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, and Keiichi Tokuda. 2021 · 2021
Later among the works it cites.
Deep variational generative models for audio-visual speech separation. In 2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP) . IEEE, 1–6
Viet-Nhat Nguyen, Mostafa Sadeghi, Elisa Ricci, and Xavier Alameda-Pineda. 2021 · 2021
Later among the works it cites.
Speech Recognition: a review of the different deep learning approaches
Ilias Papastratis. 2021 · 2021
Later among the works it cites.
A Universal Multi-Speaker Multi-Style Text-to-Speech via Disentangled Representation Learning Based on Rényi Divergence Minimization.. In Interspeech . 3625–3629
Dipjyoti Paul, Sankar Mukherjee, Yannis Pantazis, and Yannis Stylianou. 2021 · 2021
Later among the works it cites.
Wave-GAN: a deep learning approach for the prediction of nonlinear regular wave loads and run-up on a fixed cylinder
Blanca Pena and Luofeng Huang. 2021 · 2021
Later among the works it cites.
Shrinking Bigfoot: Reducing wav2vec 2.0 footprint. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing . Association for Computational Linguistics, Virtual, 134–141
Zilun Peng, Akshay Budhkar, Ilana Tuil, Jason Levy, Parinaz Sobhani, Raphael Cohen, and Jumana Nassour. 2021 · 2021
Later among the works it cites.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Diffusion-based voice conversion with fast maximum likelihood sampling scheme
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Kudinov, and Jiansheng Wei. 2021b · 2021
Later among the works it cites.
Self-attention for audio super-resolution. In 2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP) . IEEE, 1–6
Nathanaël Carraz Rakotonirina. 2021 · 2021
Later among the works it cites.
Portaspeech: Portable and high-quality generative text-to-speech
Yi Ren, Jinglin Liu, and Zhou Zhao. 2021 · 2021
Later among the works it cites.
Wav2vec-c: A self-supervised model for speech representation learning
Samik Sadhu, Di He, Che-Wei Huang, Sri Harish Mallidi, Minhua Wu, Ariya Rastrow, Andreas Stolcke, Jasha Droppo, and Roland Maas. 2021 · 2021
Later among the works it cites.
Perceptual-similarity-aware deep speaker representation learning for multi-speaker generative modeling
Yuki Saito, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2021 · 2021
Later among the works it cites.
Wav2kws: Transfer learning from speech representations for keyword spotting
Deokjin Seo, Heung-Seon Oh, and Yuchul Jung. 2021 · 2021
Later among the works it cites.
SESQA: semi-supervised learning for speech quality assessment. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 381–385
Joan Serrà, Jordi Pons, and Santiago Pascual. 2021 · 2021
Later among the works it cites.
Representation transfer learning from deep end-to-end speech recognition networks for the classification of health states from speech
Benjamin Sertolli, Zhao Ren, Björn W Schuller, and Nicholas Cummins. 2021 · 2021
Later among the works it cites.
RAD-TTS: Parallel flow-based TTS with robust alignment learning and diverse synthesis. In ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models
Kevin J Shih, Rafael Valle, Rohan Badlani, Adrian Lancucki, Wei Ping, and Bryan Catanzaro. 2021 · 2021
Later among the works it cites.
Spoken language identification using deep learning
Gundeep Singh, Sahil Sharma, Vijay Kumar, Manjit Kaur, Mohammed Baz, and Mehedi Masud. 2021 · 2021
Later among the works it cites.
Self-Supervised Metric Learning With Graph Clustering For Speaker Diarization. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . 90–97
Prachi Singh and Sriram Ganapathy. 2021 · 2021
Later among the works it cites.
Attention is all you need in speech separation. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 21–25
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong. 2021 · 2021
Later among the works it cites.
Graphpb: Graphical representations of prosody boundary in speech synthesis. In 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 438–445
Aolan Sun, Jianzong Wang, Ning Cheng, Huayi Peng, Zhen Zeng, Lingwei Kong, and Jing Xiao. 2021 · 2021
Later among the works it cites.
EdiTTS: Score-based Editing for Controllable Text-to-Speech
Jaesung Tae, Hyeongju Kim, and Taesu Kim. 2021 · 2021
Later among the works it cites.
Joint time-frequency and time domain learning for speech enhancement. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence . 3816–3822
Chuanxin Tang, Chong Luo, Zhiyuan Zhao, Wenxuan Xie, and Wenjun Zeng. 2021 · 2021
Later among the works it cites.
Multi-channel speech enhancement using graph neural networks. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3415–3419
Panagiotis Tzirakis, Anurag Kumar, and Jacob Donley. 2021 · 2021
Later among the works it cites.
Thilo von Neumann, Keisuke Kinoshita, Christoph Boeddeker, Marc Delcroix, and Reinhold Haeb-Umbach. 2021 · 2021
Later among the works it cites.
A modulation-domain loss for neural-network-based real-time speech enhancement. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6643–6647
Tyler Vuong, Yangyang Xia, and Richard M Stern. 2021 · 2021
Later among the works it cites.
Learning efficient representations for keyword spotting with triplet loss. In Speech and Computer: 23rd International Conference, SPECOM 2021, St. Petersburg, Russia, September 27–30, 2021, Proceedings 23 . Springer, 773–785
Roman Vygon and Nikolay Mikhaylovskiy. 2021 · 2021
Later among the works it cites.
Auto-KWS 2021 Challenge: Task, datasets, and baselines
Jingsong Wang, Yuxuan He, Chunyu Zhao, Qijie Shao, Wei-Wei Tu, Tom Ko, Hung-yi Lee, and Lei Xie. 2021a · 2021
Later among the works it cites.
Transformer in Action: A Comparative Study of Transformer-Based Acoustic Models for Large Scale Speech Recognition Applications. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 6778–6782
Yongqiang Wang, Yangyang Shi, Frank Zhang, Chunyang Wu, Julian Chan, Ching-Feng Yeh, and Alex Xiao. 2021b · 2021
Later among the works it cites.
Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5679–5683
Ron J Weiss, RJ Skerry-Ryan, Eric Battenberg, Soroosh Mariooryad, and Diederik P Kingma. 2021 · 2021
Later among the works it cites.
ItoTTS and ItoWave: Linear Stochastic Differential Equation Is All You Need For Audio Generation
Shoule Wu and Ziqiang Shi. 2021 · 2021
Later among the works it cites.
Microsoft speaker diarization system for the voxceleb speaker recognition challenge 2020. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5824–5828
Xiong Xiao, Naoyuki Kanda, Zhuo Chen, Tianyan Zhou, Takuya Yoshioka, Sanyuan Chen, Yong Zhao, Gang Liu, Yu Wu, Jian Wu, et al · 2021
Later among the works it cites.
Chen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Qi Ju, Tong Xiao, Jingbo Zhu, et al · 2021
Later among the works it cites.
Self-training and pre-training are complementary for speech recognition. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3030–3034
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. 2021a · 2021
Later among the works it cites.
Adaspeech 2: Adaptive text to speech with untranscribed data. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6613–6617
Yuzi Yan, Xu Tan, Bohan Li, Tao Qin, Sheng Zhao, Yuan Shen, and Tie-Yan Liu. 2021 · 2021
Later among the works it cites.
Multi-band melgan: Faster waveform generation for high-quality text-to-speech. In 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 492–498
Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie. 2021b · 2021
Later among the works it cites.
SUPERB: Speech processing Universal PERformance Benchmark
Shu-Wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung-yi Lee. 2021a · 2021
Later among the works it cites.
A deep neural network model for speaker identification
Feng Ye and Jun Yang. 2021 · 2021
Later among the works it cites.
End-to-end speech translation via cross-modal progressive training
Rong Ye, Mingxuan Wang, and Lei Li. 2021 · 2021
Later among the works it cites.
Gan vocoder: Multi-resolution discriminator is all you need
Jaeseong You, Dalhyun Kim, Gyuhyeon Nam, Geumbyeol Hwang, and Gyeongsu Chae. 2021 · 2021
Later among the works it cites.
Wavesplit: End-to-end speech separation by speaker clustering
Neil Zeghidour and David Grangier. 2021 · 2021
Later among the works it cites.
Librispeech transducer model with internal language model prior correction
Albert Zeyer, André Merboldt, Wilfried Michel, Ralf Schlüter, and Hermann Ney. 2021 · 2021
Later among the works it cites.
Denoispeech: Denoising text to speech with frame-level noise modeling. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7063–7067
Chen Zhang, Yi Ren, Xu Tan, Jinglin Liu, Kejun Zhang, Tao Qin, Sheng Zhao, and Tie-Yan Liu. 2021b · 2021
Later among the works it cites.
Meta-learning for cross-channel speaker verification. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5839–5843
Hanyi Zhang, Longbiao Wang, Kong Aik Lee, Meng Liu, Jianwu Dang, and Hui Chen. 2021c · 2021
Later among the works it cites.
Non-parallel Sequence-to-Sequence Voice Conversion for Arbitrary Speakers. In 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) . 1–5
Ying Zhang, Hao Che, and Xiaorui Wang. 2021a · 2021
Later among the works it cites.
Monaural speech enhancement with complex convolutional block attention module and joint time frequency losses. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6648–6652
Shengkui Zhao, Trung Hieu Nguyen, and Bin Ma. 2021a · 2021
Later among the works it cites.
Towards natural and controllable cross-lingual voice conversion based on neural tts model and phonetic posteriorgram. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5969–5973
Shengkui Zhao, Hao Wang, Trung Hieu Nguyen, and Bin Ma. 2021b · 2021
Later among the works it cites.
Serialized multi-layer multi-head attention for neural speaker embedding
Hongning Zhu, Kong Aik Lee, and Haizhou Li. 2021 · 2021
Later among the works it cites.
Conformer-1
2022 · 2022
Later among the works it cites.
Speech Recognition With Conformer
2022 · 2022
Later among the works it cites.
Mel Frequency Cepstral Coefficient and its Applications: A Review
Zrar Kh. Abdul and Abdulbasit K. Al-Talabani. 2022 · 2022
Later among the works it cites.
CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement
Sherif Abdulatif, Ruizhe Cao, and Bin Yang. 2022 · 2022
Later among the works it cites.
One TTS alignment to rule them all. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6092–6096
Rohan Badlani, Adrian Łańcucki, Kevin J Shih, Rafael Valle, Wei Ping, and Bryan Catanzaro. 2022a · 2022
Later among the works it cites.
One TTS Alignment to Rule Them All. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 6092–6096
Rohan Badlani, Adrian Łańcucki, Kevin J. Shih, Rafael Valle, Wei Ping, and Bryan Catanzaro. 2022b · 2022
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language. In International Conference on Machine Learning . PMLR, 1298–1312
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022 · 2022
Later among the works it cites.
A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing. In International Conference on Machine Learning . PMLR, 1399–1411
He Bai, Renjie Zheng, Junkun Chen, Mingbo Ma, Xintong Li, and Liang Huang. 2022 · 2022
Later among the works it cites.
Avocodo: Generative adversarial network for artifact-free vocoder
Taejun Bak, Junmo Lee, Hanbin Bae, Jinhyeok Yang, Jae-Sung Bae, and Young-Sun Joo. 2022 · 2022
Later among the works it cites.
An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks
Kai-Wei Chang et al · 2022
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Later among the works it cites.
Unispeech-sat: Universal speech representation learning with speaker aware pre-training. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6152–6156
Sanyuan Chen, Yu Wu, Chengyi Wang, Zhengyang Chen, Zhuo Chen, Shujie Liu, Jian Wu, Yao Qian, Furu Wei, Jinyu Li, et al · 2022
Later among the works it cites.
Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8432–8436
Zehua Chen, Xu Tan, Ke Wang, Shifeng Pan, Danilo Mandic, Lei He, and Sheng Zhao. 2022a · 2022
Later among the works it cites.
Self-supervised learning with random-projection quantizer for speech recognition. In International Conference on Machine Learning . PMLR, 3915–3924
Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu, and Yonghui Wu. 2022 · 2022
Later among the works it cites.
NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis
Hyeong-Seok Choi, Jinhyeok Yang, Juheon Lee, and Hyeongju Kim. 2022 · 2022
Later among the works it cites.
Domain Adaptation for Speaker Recognition in Singing and Spoken Voice. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7192–7196
Anurag Chowdhury, Austin Cozzo, and Arun Ross. 2022 · 2022
Later among the works it cites.
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Later among the works it cites.
Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 8517–8521
Keqi Deng, Songjun Cao, Yike Zhang, Long Ma, Gaofeng Cheng, Ji Xu, and Pengyuan Zhang. 2022a · 2022
Later among the works it cites.
Improving CTC-based speech recognition via knowledge transferring from pre-trained language models. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8517–8521
Keqi Deng, Songjun Cao, Yike Zhang, Long Ma, Gaofeng Cheng, Ji Xu, and Pengyuan Zhang. 2022b · 2022
Later among the works it cites.
Adversarial Text-to-Speech for low-resource languages. In Proceedings of the The Seventh Arabic Natural Language Processing Workshop (WANLP) . 76–84
Ashraf Elneima and Mikołaj Bińkowski. 2022 · 2022
Later among the works it cites.
Self-supervised representation learning: Introduction, advances, and challenges
Linus Ericsson, Henry Gouk, Chen Change Loy, and Timothy M Hospedales. 2022 · 2022
Later among the works it cites.
Voice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7902–7906
Adam Gabryś, Goeric Huybrechts, Manuel Sam Ribeiro, Chung-Ming Chien, Julian Roth, Giulia Comini, Roberto Barra-Chicote, Bartek Perz, and Jaime Lorenzo-Trueba. 2022 · 2022
Later among the works it cites.
A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS
Haohan Guo, Fenglong Xie, Frank K Soong, Xixin Wu, and Helen Meng. 2022 · 2022
Later among the works it cites.
A motion matching-based framework for controllable gesture synthesis from speech. In ACM SIGGRAPH 2022 Conference Proceedings . 1–9
Ikhsanul Habibie, Mohamed Elgharib, Kripasindhu Sarkar, Ahsan Abdullah, Simbarashe Nyatsanga, Michael Neff, and Christian Theobalt. 2022 · 2022
Later among the works it cites.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al · 2022
Later among the works it cites.
NU-Wave 2: A general neural audio upsampling model for various sampling rates
Seungu Han and Junhyeok Lee. 2022 · 2022
Later among the works it cites.
Yosuke Higuchi, Brian Yan, Siddhant Arora, Tetsuji Ogawa, Tetsunori Kobayashi, and Shinji Watanabe. 2022 · 2022
Later among the works it cites.
Wei-Ning Hsu, Tal Remez, Bowen Shi, Jacob Donley, and Yossi Adi. 2022b · 2022
Later among the works it cites.
Language model compression with weighted low-rank factorization
Yen-Chang Hsu, Ting Hua, Sung-En Chang, Qiang Lou, Yilin Shen, and Hongxia Jin. 2022a · 2022
Later among the works it cites.
Domain Robust Deep Embedding Learning for Speaker Recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7182–7186
Hang-Rui Hu, Yan Song, Ying Liu, Li-Rong Dai, Ian McLoughlin, and Lin Liu. 2022a · 2022
Later among the works it cites.
Fastdiff: A fast conditional diffusion model for high-quality speech synthesis
Rongjie Huang, Max WY Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao. 2022a · 2022
Later among the works it cites.
Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech
Rongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao. 2022c · 2022
Later among the works it cites.
Meta-TTS: Meta-learning for few-shot speaker adaptive text-to-speech
Sung-Feng Huang, Chyi-Jiunn Lin, Da-Rong Liu, Yi-Chen Chen, and Hung-yi Lee. 2022b · 2022
Later among the works it cites.
A large TV dataset for speech and music activity detection
Yun-Ning Hung, Chih-Wei Wu, Iroro Orife, Aaron Hipple, William Wolcott, and Alexander Lerch. 2022 · 2022
Later among the works it cites.
Large-scale asr domain adaptation using self-and semi-supervised learning. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6627–6631
Dongseong Hwang, Ananya Misra, Zhouyuan Huo, Nikhil Siddhartha, Shefali Garg, David Qiu, Khe Chai Sim, Trevor Strohman, Françoise Beaufays, and Yanzhang He. 2022 · 2022
Later among the works it cites.
Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. In International Conference on Machine Learning . PMLR, 10120–10134
Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz. 2022 · 2022
Later among the works it cites.
TriniTTS: Pitch-controllable End-to-end TTS without External Aligner. In Proc. Interspeech . 16–20
Yooncheol Ju, Ilhwan Kim, Hongsun Yang, Ji-Hoon Kim, Byeongyeol Kim, Soumi Maiti, and Shinji Watanabe. 2022 · 2022
Later among the works it cites.
Deep Learning-Based Speech Emotion Recognition Using Multi-Level Fusion of Concurrent Features
Samuel Kakuba, Alwin Poulose, and Dong Seog Han. 2022 · 2022
Later among the works it cites.
Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings
Naoyuki Kanda, Jian Wu, Yu Wu, Xiong Xiao, Zhong Meng, Xiaofei Wang, Yashesh Gaur, Zhuo Chen, Jinyu Li, and Takuya Yoshioka. 2022a · 2022
Later among the works it cites.
Transcribe-to-diarize: Neural speaker diarization for unlimited number of speakers using end-to-end speaker-attributed asr. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8082–8086
Naoyuki Kanda, Xiong Xiao, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, and Takuya Yoshioka. 2022b · 2022
Later among the works it cites.
iSTFTNet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time Fourier transform. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6207–6211
Takuhiro Kaneko, Kou Tanaka, Hirokazu Kameoka, and Shogo Seki. 2022 · 2022
Later among the works it cites.
Learning continuous representation of audio for arbitrary scale super resolution. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3703–3707
Jaechang Kim, Yunjoo Lee, Seunghoon Hong, and Jungseul Ok. 2022d · 2022
Later among the works it cites.
Squeezeformer: An efficient transformer for automatic speech recognition
Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W Mahoney, and Kurt Keutzer. 2022a · 2022
Later among the works it cites.
Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech with Untranscribed Data
Sungwon Kim, Heeseung Kim, and Sungroh Yoon. 2022c · 2022
Later among the works it cites.
SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping
Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen, and Michiel Bacchiani. 2022 · 2022
Later among the works it cites.
TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8102–8106
Nithin Rao Koluguri, Taejin Park, and Boris Ginsburg. 2022 · 2022
Later among the works it cites.
AudioGen: Textually Guided Audio Generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre D’efossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2022 · 2022
Later among the works it cites.
Multi-scale speaker embedding-based graph attention networks for speaker diarisation. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8367–8371
Youngki Kwon, Hee-Soo Heo, Jee-weon Jung, You Jin Kim, Bong-Jin Lee, and Joon Son Chung. 2022 · 2022
Later among the works it cites.
Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implementation and analysis on standard tasks
Federico Landini, Ján Profant, Mireia Diez, and Lukáš Burget. 2022 · 2022
Later among the works it cites.
Self-supervised Representation Learning for Speech Processing. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Tutorial Abstracts . 8–13
Hung-yi Lee, Abdelrahman Mohamed, Shinji Watanabe, Tara Sainath, Karen Livescu, Shang-Wen Li, Shu-wen Yang, and Katrin Kirchhoff. 2022b · 2022
Later among the works it cites.
HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis
Sang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song, Min-Jae Hwang, and Seong-Whan Lee. 2022a · 2022
Later among the works it cites.
StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation
Jean-Marie Lemercier, Julius Richter, Simon Welker, and Timo Gerkmann. 2022 · 2022
Later among the works it cites.
BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio Synthesis
Yichong Leng, Zehua Chen, Junliang Guo, Haohe Liu, Jiawei Chen, Xu Tan, Danilo Mandic, Lei He, Xiang-Yang Li, Tao Qin, et al · 2022
Later among the works it cites.
Zero-Shot Voice Conditioning for Denoising Diffusion TTS Models
Alon Levkovitch, Eliya Nachmani, and Lior Wolf. 2022 · 2022
Later among the works it cites.
Recent advances in end-to-end automatic speech recognition
Jinyu Li et al · 2022
Later among the works it cites.
EDITnet: A Lightweight Network for Unsupervised Domain Adaptation in Speaker Verification
Jingyu Li, Wei Liu, and Tan Lee. 2022c · 2022
Later among the works it cites.
An efficient encoder-decoder architecture with top-down attention for speech separation
Kai Li, Runxuan Yang, and Xiaolin Hu. 2022d · 2022
Later among the works it cites.
The coral++ algorithm for unsupervised domain adaptation of speaker recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7172–7176
Rongjin Li, Weibin Zhang, and Dongpeng Chen. 2022e · 2022
Later among the works it cites.
StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis
Yinghao Aaron Li, Cong Han, and Nima Mesgarani. 2022b · 2022
Later among the works it cites.
JETS: Jointly training FastSpeech2 and HiFi-GAN for end to end text to speech
Dan Lim, Sunghee Jung, and Eesung Kim. 2022 · 2022
Later among the works it cites.
Towards end-to-end unsupervised speech recognition. In 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 221–228
Alexander H Liu, Wei-Ning Hsu, Michael Auli, and Alexei Baevski. 2023b · 2022
Later among the works it cites.
Simple and effective unsupervised speech synthesis
Alexander H Liu, Cheng-I Jeff Lai, Wei-Ning Hsu, Michael Auli, Alexei Baevskiv, and James Glass. 2022b · 2022
Later among the works it cites.
Neural vocoder is all you need for speech super-resolution
Haohe Liu, Woosung Choi, Xubo Liu, Qiuqiang Kong, Qiao Tian, and DeLiang Wang. 2022a · 2022
Later among the works it cites.
An Improvement to Conformer-Based Model for High-Accuracy Speech Feature Extraction and Learning
Mengzhuo Liu and Yangjie Wei. 2022 · 2022
Later among the works it cites.
Diffgan-tts: High-fidelity and efficient text-to-speech with denoising diffusion gans
Songxiang Liu, Dan Su, and Dong Yu. 2022d · 2022
Later among the works it cites.
Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech
Zhengxi Liu, Qiao Tian, Chenxu Hu, Xudong Liu, Menglin Wu, Yuping Wang, Hang Zhao, and Yuxuan Wang. 2022e · 2022
Later among the works it cites.
Conditional Diffusion Probabilistic Model for Speech Enhancement. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 7402–7406
Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe, Alexander Richard, Cheng Yu, and Yu Tsao. 2022a · 2022
Later among the works it cites.
Conditional diffusion probabilistic model for speech enhancement. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7402–7406
Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe, Alexander Richard, Cheng Yu, and Yu Tsao. 2022b · 2022
Later among the works it cites.
Sepit approaching a single channel speech separation bound
Shahar Lutati, Eliya Nachmani, and Lior Wolf. 2022 · 2022
Later among the works it cites.
Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features
Florian Lux and Ngoc Thang Vu. 2022 · 2022
Later among the works it cites.
Speaking Style Conversion With Discrete Self-Supervised Units
Gallil Maimon and Yossi Adi. 2022 · 2022
Later among the works it cites.
Damage Control During Domain Adaptation for Transducer Based Automatic Speech Recognition. In 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 130–135
Somshubra Majumdar, Shantanu Acharya, Vitaly Lavrukhin, and Boris Ginsburg. 2023 · 2022
Later among the works it cites.
A Kernel-Based View of Language Model Fine-Tuning
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen, and Sanjeev Arora. 2022 · 2022
Later among the works it cites.
Speech quality assessment through MOS using non-matching references
Pranay Manocha and Anurag Kumar. 2022 · 2022
Later among the works it cites.
Neural HMMS Are All You Need (For High-Quality Attention-Free TTS). In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 7457–7461
Shivam Mehta, Eva Szekely, Jonas Beskow, and Gustav Eje Henter. 2022 · 2022
Later among the works it cites.
Using Voice Activity Detection and Deep Neural Networks with Hybrid Speech Feature Extraction for Deceptive Speech Detection
Serban Mihalache and Dragos Burileanu. 2022 · 2022
Later among the works it cites.
Toward a realistic model of speech processing in the brain with self-supervised learning
Juliette Millet, Charlotte Caucheteux, Yves Boubenec, Alexandre Gramfort, Ewan Dunbar, Christophe Pallier, Jean-Remi King, et al · 2022
Later among the works it cites.
Multi-Channel Speech Enhancement using a Minimum Variance Distortionless Response Beamformer based on Graph Convolutional Network
Huu Binh Nguyen, Duong Van Hai, Tien Dat Bui, Hoang Ngoc Chau, and Quoc Cuong Nguyen. 2022c · 2022
Later among the works it cites.
Tunet: A block-online bandwidth extension model based on transformers and self-supervised pretraining. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 161–165
Viet-Anh Nguyen, Anh HT Nguyen, and Andy WH Khong. 2022a · 2022
Later among the works it cites.
Improving Speech-to-Speech Translation Through Unlabeled Text
Xuan-Phi Nguyen, Sravya Popuri, Changhan Wang, Yun Tang, Ilia Kulikov, and Hongyu Gong. 2022b · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022 · 2022
Later among the works it cites.
SRU++: Pioneering Fast Recurrence with Attention for Speech Recognition. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 7872–7876
Jing Pan, Tao Lei, Kwangyoun Kim, Kyu J. Han, and Shinji Watanabe. 2022 · 2022
Later among the works it cites.
Bunched LPCNet2: Efficient Neural Vocoders Covering Devices from Cloud to Edge
Sangjun Park, Kihyun Choo, Joohyung Lee, Anton V Porov, Konstantin Osipov, and June Sig Sung. 2022 · 2022
Later among the works it cites.
Contentvec: An improved self-supervised speech representation by disentangling speakers. In International Conference on Machine Learning . PMLR, 18003–18017
Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David Cox, Mark Hasegawa-Johnson, and Shiyu Chang. 2022 · 2022
Later among the works it cites.
SRTNet: Time Domain Speech Enhancement Via Stochastic Refinement
Zhibin Qiu, Mengfan Fu, Yinfeng Yu, LiLi Yin, Fuchun Sun, and Hao Huang. 2022 · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022 · 2022
Later among the works it cites.
A novel policy for pre-trained deep reinforcement learning for speech emotion recognition
Thejan Rajapakshe, Rajib Rana, Sara Khalifa, Jiajun Liu, and Bjorn Schuller. 2022 · 2022
Later among the works it cites.
Nas-vad: Neural architecture search for voice activity detection
Daniel Rho, Jinhyeok Park, and Jong Hwan Ko. 2022 · 2022
Later among the works it cites.
Keyword spotting in continuous speech using convolutional neural network
Amir Mohammad Rostami, Ali Karimi, and Mohammad Ali Akhaee. 2022 · 2022
Later among the works it cites.
Phase sensitive masking-based single channel speech enhancement using conditional generative adversarial network
Sidheswar Routray and Qirong Mao. 2022 · 2022
Later among the works it cites.
Contextual adapters for personalized speech recognition in neural transducers. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8537–8541
Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Feng-Ju Chang, Jing Liu, Jinru Su, Grant P Strimel, Athanasios Mouchtaris, and Siegfried Kunzmann. 2022 · 2022
Later among the works it cites.
Diffusion-based Generative Speech Source Separation
Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, and Min-Seok Choi. 2022 · 2022
Later among the works it cites.
Universal speech enhancement with score-based diffusion
Joan Serrà, Santiago Pascual, Jordi Pons, R Oguz Araz, and Davide Scaini. 2022 · 2022
Later among the works it cites.
Graph Attentive Feature Aggregation for Text-Independent Speaker Verification. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 7972–7976
Hye-Jin Shim, Jungwoo Heo, Jae-Han Park, Ga-Hui Lee, and Ha-Jin Yu. 2022 · 2022
Later among the works it cites.
Speaker recognition using constrained convolutional neural networks in emotional speech
Nikola Simić, Siniša Suzić, Tijana Nosek, Mia Vujović, Zoran Perić, Milan Savić, and Vlado Delić. 2022 · 2022
Later among the works it cites.
Improved Meta Learning for Low Resource Speech Recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4798–4802
Satwinder Singh, Ruili Wang, and Feng Hou. 2022 · 2022
Later among the works it cites.
WavThruVec: Latent speech representation as intermediate features for neural speech synthesis
Hubert Siuzdak, Piotr Dura, Pol van Rijn, and Nori Jacoby. 2022 · 2022
Later among the works it cites.
Domain Adaptation of low-resource Target-Domain models using well-trained ASR Conformer Models. In 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 295–301
Vrunda N Sukhadia and S Umesh. 2023 · 2022
Later among the works it cites.
Continual self-training with bootstrapped remixing for speech enhancement. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6947–6951
Efthymios Tzinis, Yossi Adi, Vamsi K Ithapu, Buye Xu, and Anurag Kumar. 2022a · 2022
Later among the works it cites.
RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing
Efthymios Tzinis, Yossi Adi, Vamsi K Ithapu, Buye Xu, Paris Smaragdis, and Anurag Kumar. 2022b · 2022
Later among the works it cites.
Neural speech synthesis on a shoestring: Improving the efficiency of lpcnet. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8437–8441
Jean-Marc Valin, Umut Isik, Paris Smaragdis, and Arvindh Krishnaswamy. 2022 · 2022
Later among the works it cites.
Conformer Based Elderly Speech Recognition System for Alzheimer’s Disease Detection
Tianzi Wang, Jiajun Deng, Mengzhe Geng, Zi Ye, Shoukang Hu, Yi Wang, Mingyu Cui, Zengrui Jin, Xunying Liu, and Helen Meng. 2022a · 2022
Later among the works it cites.
Similarity measurement of segment-level speaker embeddings in speaker diarization
Weiqing Wang, Qingjian Lin, Danwei Cai, and Ming Li. 2022b · 2022
Later among the works it cites.
Deep Sparse Conformer for Speech Recognition
Xianchao Wu. 2022 · 2022
Later among the works it cites.
Adaspeech 4: Adaptive text to speech in zero-shot scenarios
Yihan Wu, Xu Tan, Bohan Li, Lei He, Sheng Zhao, Ruihua Song, Tao Qin, and Tie-Yan Liu. 2022 · 2022
Later among the works it cites.
ECAPA-TDNN for Multi-speaker Text-to-speech Synthesis. In 2022 13th International Symposium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 230–234
Jinlong Xue, Yayue Deng, Yichen Han, Ya Li, Jianqing Sun, and Jiaen Liang. 2022 · 2022
Later among the works it cites.
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Jianwei Yu, Helin Wang, Chao Weng, and Yuexian Zou. 2022a · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu. 2022b · 2022
Later among the works it cites.
Data augmentation for speaker verification. In Proceedings of the 2022 6th International Conference on Electronic Information Technology and Computer Engineering . 1247–1251
Shiqing Yang and Min Liu. 2022 · 2022
Later among the works it cites.
Cold Diffusion for Speech Enhancement
Hao Yen, François G Germain, Gordon Wichern, and Jonathan Le Roux. 2022 · 2022
Later among the works it cites.
Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs
Reo Yoneyama, Ryuichi Yamamoto, and Kentaro Tachibana. 2022 · 2022
Later among the works it cites.
Hubert-ee: Early exiting hubert for efficient speech recognition
Ji Won Yoon, Beom Jun Woo, and Nam Soo Kim. 2022 · 2022
Later among the works it cites.
Auxiliary loss of transformer with residual connection for end-to-end speaker diarization. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8377–8381
Yechan Yu, Dongkeon Park, and Hong Kook Kim. 2022 · 2022
Later among the works it cites.
Exploring machine speech chain for domain adaptation. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6757–6761
Fengpeng Yue, Yan Deng, Lei He, Tom Ko, and Yu Zhang. 2022 · 2022
Later among the works it cites.
Towards end-to-end speaker diarization with generalized neural speaker clustering. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8372–8376
Chunlei Zhang, Jiatong Shi, Chao Weng, Meng Yu, and Dong Yu. 2022d · 2022
Later among the works it cites.
Hifidenoise: High-fidelity denoising text to speech with adversarial networks. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7232–7236
Lichao Zhang, Yi Ren, Liqun Deng, and Zhou Zhao. 2022c · 2022
Later among the works it cites.
TDASS: Target Domain Adaptation Speech Synthesis Framework for Multi-speaker Low-Resource TTS. In 2022 International Joint Conference on Neural Networks (IJCNN) . 1–7
Xulong Zhang, Jianzong Wang, Ning Cheng, and Jing Xiao. 2022e · 2022
Later among the works it cites.
BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
Yu Zhang, Daniel S. Park, Wei Han, James Qin, Anmol Gulati, Joel Shor, Aren Jansen, Yuanzhong Xu, Yanping Huang, Shibo Wang, Zongwei Zhou, Bo Li, Min Ma, William Chan, Jiahui Yu, Yongqiang Wang, Liangliang Cao, Khe Chai Sim, Bhuvana Ramabhadran, Tara N. Sainath, Françoise Beaufays, Zhifeng Chen, Quoc V. Le, Chung-Cheng Chiu, Ruoming Pang, and Yonghui Wu. 2022b · 2022
Later among the works it cites.
Tiny-Attention Adapter: Contexts Are More Important Than the Number of Parameters
Hongyu Zhao, Hao Tan, and Hongyuan Mei. 2022 · 2022
Later among the works it cites.
Multi-Source Domain Adaptation and Fusion for Speaker Verification
Donghui Zhu and Ning Chen. 2022 · 2022
Later among the works it cites.
Human–Computer Interaction with a Real-Time Speech Emotion Recognition with Ensembling Techniques 1D Convolution Neural Network and Attention
Waleed Alsabhan. 2023 · 2023
Closest in time.
Audio-Visual Efficient Conformer for Robust Speech Recognition. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 2258–2267
Maxime Burchi and Radu Timofte. 2023 · 2023
Closest in time.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar and Dan Alistarh. 2023 · 2023
Closest in time.
Self-supervised speech representation learning for keyword-spotting with light-weight transformers
Chenyang Gao, Yue Gu, Francesco Caliva, and Yuzong Liu. 2023 · 2023
Closest in time.
Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria. 2023 · 2023
Closest in time.
Language Is Not All You Need: Aligning Perception with Language Models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. 2023 · 2023
Closest in time.
Yingting Li, Ambuj Mehrish, Shuai Zhao, Rishabh Bhardwaj, Amir Zadeh, Navonil Majumder, Rada Mihalcea, and Soujanya Poria. 2023 · 2023
Closest in time.
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Haohe Liu, Zehua Chen, Yiitan Yuan, Xinhao Mei, Xubo Liu, Danilo P. Mandic, Wenwu Wang, and MarkD . Plumbley. 2023a · 2023
Closest in time.
Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation
Shahar Lutati, Eliya Nachmani, and Lior Wolf. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Prosody-TTS: An end-to-end speech synthesis system with prosody control
Giridhar Pamisetty and K Sri Rama Murty. 2023 · 2023
Closest in time.
CTRAN: CNN-Transformer-based Network for Natural Language Understanding
Mehrdad Rafiepour and Javad Salimi Sartakhti. 2023 · 2023
Closest in time.
Analysing Discrete Self Supervised Speech Representation for Spoken Language Modeling
Amitay Sicherman and Yossi Adi. 2023 · 2023
Closest in time.
Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization
Prachi Singh, Amrit Kaul, and Sriram Ganapathy. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aur’elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Closest in time.
Time-Domain Speech Separation Networks With Graph Encoding Auxiliary
Tingting Wang, Zexu Pan, Meng Ge, Zhen Yang, and Haizhou Li. 2023c · 2023
Closest in time.
Meta-Generalization for Domain-Invariant Speaker Verification
Hanyi Zhang, Longbiao Wang, Kong Aik Lee, Meng Liu, Jianwu Dang, and Helen Meng. 2023b · 2023
Closest in time.
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Yu Zhang, Wei Han, James Qin, Yongqiang Wang, Ankur Bapna, Zhehuai Chen, Nanxin Chen, Bo Li, Vera Axelrod, Gary Wang, et al · 2023
Closest in time.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Closest in time.
Shengkui Zhao and Bin Ma. 2023 · 2023
Closest in time.
An Emotion Speech Synthesis Method Based on VITS
Wei Zhao and Zheng Yang. 2023 · 2023
Closest in time.
Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement. In International Conference on Machine Learning . PMLR, 2031–2041
Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, and Shou-De Lin. 2019 · 2041
Closest in time.
Unsupervised speech representation learning using wavenet autoencoders
Jan Chorowski, Ron J Weiss, Samy Bengio, and Aäron Van Den Oord. 2019 · 2053
Closest in time.