Fetching the paper…
Reading the bibliography…
There is a wide variety of speech processing tasks ranging from extracting content information from speech signals to generating speech signals.
Almost unsupervised text to speech and automatic speech recognition
Yi Ren, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 1905
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra · 2001
Earlier work this paper cites.
The cmu arctic speech databases
John Kominek and Alan W Black · 2004
Earlier work this paper cites.
A tandem algorithm for pitch estimation and voiced speech segregation
Guoning Hu and DeLiang Wang · 2010
Earlier work this paper cites.
A short-time objective intelligibility measure for time-frequency weighted noisy speech
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen · 2010
Earlier work this paper cites.
Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks
Zhuo Chen, Shinji Watanabe, Hakan Erdogan, and John R Hershey · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Multi-task recurrent model for speech and speaker recognition
Zhiyuan Tang, Lantian Li, and Dong Wang · 2016
Earlier work this paper cites.
Lukasz Kaiser, Aidan N Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
Non-parallel voice conversion using i-vector plda: Towards unifying speaker verification and transformation
Tomi Kinnunen, Lauri Juvela, Paavo Alku, and Junichi Yamagishi · 2017
Earlier work this paper cites.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Sebastian Ruder · 2017
Earlier work this paper cites.
Listening while speaking: Speech chain by deep learning
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A survey on multi-task learning
Yu Zhang and Qiang Yang · 2017
Cited alongside, same era.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu · 2018
Cited alongside, same era.
Improved accented speech recognition using accent embeddings and multi-task learning
Abhinav Jain, Minali Upreti, and Preethi Jyothi · 2018
Cited alongside, same era.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, Yonghui Wu, et al · 2018
Cited alongside, same era.
Improving sequence-to-sequence voice conversion by adding text-supervision
Jing-Xuan Zhang, Zhen-Hua Ling, Yuan Jiang, Li-Juan Liu, Chen Liang, and Li-Rong Dai · 2019
Later among the works it cites.
Joint training framework for text-to-speech and voice conversion using multi-source tacotron and wavenet
Mingyang Zhang, Xin Wang, Fuming Fang, Haizhou Li, and Junichi Yamagishi · 2019
Later among the works it cites.
Aipnet: Generative adversarial pre-training of accent-invariant networks for end-to-end speech recognition
Yi-Chen Chen, Zhaojun Yang, Ching-Feng Yeh, Mahaveer Jain, and Michael L Seltzer · 2020
Later among the works it cites.
Boosting objective scores of speech enhancement model through metricgan post-processing
Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh, Kuo-Hsuan Hung, Syu-Siang Wang, Cheng Yu, Heng-Cheng Kuo, Ryandhimas E Zezario, You-Jin Li, Shang-Yi Chuang, et al · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Cited alongside, same era.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
Self-attentive speaker embeddings for text-independent speaker verification
Yingke Zhu, Tom Ko, David Snyder, Brian Mak, and Daniel Povey · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factorization
Wei-Ning Hsu, Yu Zhang, Ron J Weiss, Yu-An Chung, Yuxuan Wang, Yonghui Wu, and James Glass · 2019
Cited alongside, same era.
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Later among the works it cites.
Voice transformer network: Sequence-to-sequence voice conversion using transformer with text-to-speech pretraining
Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, and Tomoki Toda · 2020
Later among the works it cites.
Sandesh V Katta, S Umesh, et al · 2020
Later among the works it cites.
T-gsa: Transformer with gaussian-weighted self-attention for speech enhancement
Jaeyoung Kim, Mostafa El-Khamy, and Jungwon Lee · 2020
Later among the works it cites.
Yist Y Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, and Lin-shan Lee · 2020
Later among the works it cites.
Voice conversion with transformer network
Ruolan Liu, Xiao Chen, and Xue Wen · 2020
Later among the works it cites.
Masked multi-head self-attention for causal speech enhancement
Aaron Nicolson and Kuldip K Paliwal · 2020
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text-to-speech
Yi Ren, Chenxu Hu, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2020
Later among the works it cites.
Self-attention encoding and pooling for speaker recognition
Pooyan Safari and Javier Hernando · 2020
Later among the works it cites.
T-vectors: Weakly supervised speaker identification using hierarchical transformer model
Yanpei Shi, Mingjie Chen, Qiang Huang, and Thomas Hain · 2020
Later among the works it cites.
Multi-task learning for dense prediction tasks: A survey
Simon Vandenhende, Stamatios Georgoulis, Wouter Van Gansbeke, Marc Proesmans, Dengxin Dai, and Luc Van Gool · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Later among the works it cites.
Aligntts: Efficient feed-forward text-to-speech system without explicit alignment
Zhen Zeng, Jianzong Wang, Ning Cheng, Tian Xia, and Jing Xiao · 2020
Later among the works it cites.