Fetching the paper…
Reading the bibliography…
We propose SpeechPainter, a model for filling in gaps of up to one second in speech samples by leveraging an auxiliary textual input.
Scene completion using millions of photographs
J. Hays and A. A. Efros · 2007
Earlier work this paper cites.
Audio Inpainting
A. Adler, V. Emiya, M. G. Jafari, M. Elad, R. Gribonval, and M. D. Plumbley · 2012
Earlier work this paper cites.
Self-content-based audio inpainting
Y. Bahat, Y. Y. Schechner, and M. Elad · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Layer Normalization, 2016
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther · 2016
Earlier work this paper cites.
Text-informed speech inpainting via voice conversion
P. Prablanc, A. Ozerov, N. Q. K. Duong, and P. Pérez · 2016
Earlier work this paper cites.
Attention is All you Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Cited alongside, same era.
Audio-Visual Speech Enhancement Using Multimodal Deep Convolutional Neural Networks
J. Hou, S. Wang, Y. Lai, Y. Tsao, H. Chang, and H. Wang · 2018
Cited alongside, same era.
Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous · 2018
Cited alongside, same era.
Hierarchical Generative Modeling for Controllable Speech Synthesis
W.-N. Hsu, Y. Zhang, R. Weiss, H. Zen, Y. Wu, Y. Cao, and Y. Wang · 2019
Cited alongside, same era.
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville · 2019
Cited alongside, same era.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
J. Kong, J. Kim, and J. Bae · 2020
Later among the works it cites.
SEANet: A Multi-modal Speech Enhancement Network
D. Roblek, K. Misiunas, M. Tagliasacchi, and P. Li · 2020
Later among the works it cites.
Perceiver IO: A General Architecture for Structured Inputs & Outputs, 2021
A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, O. Hénaff, M. M. Botvinick, A. Zisserman, O. Vinyals, and J. Carreira · 2021
Later among the works it cites.
Perceiver: General Perception with Iterative Attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Later among the works it cites.
Meta-StyleSpeech: Multi-Speaker Adaptive Text-to-Speech Generation
D. Min, D. B. Lee, E. Yang, and S. J. Hwang · 2021
Later among the works it cites.
Audio-Visual Speech Inpainting with Deep Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. R. Hershey, R. A. Saurous, R. J. Weiss, Y. Jia, and I. L. Moreno · 2019
Cited alongside, same era.
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92), 2019
J. Yamagishi, C. Veaux, and K. MacDonald · 2019
Cited alongside, same era.
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu · 2019
Cited alongside, same era.
G. Morrone, D. Michelsanti, Z.-H. Tan, and J. Jensen · 2021
Later among the works it cites.
SoundStream: An End-to-End Neural Audio Codec, 2021
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2021
Later among the works it cites.