Fetching the paper…
Reading the bibliography…
Understanding and controlling latent representations in deep generative models is a challenging yet important problem for analyzing, transforming and generating various types of data.
On lines and planes of closest fit to systems of points in space
Pearson, K. (1901) · 1901
Earlier work this paper cites.
Supervised determined source separation with multichannel variational autoencoder
Kameoka, H., Li, L., Inoue, S., & Makino, S. (2019) · 1914
Earlier work this paper cites.
The relations of the newer multivariate statistical methods to factor analysis
Hotelling, H. (1957) · 1957
Earlier work this paper cites.
A course in multivariate analysis
Kendall, M. (1957) · 1957
Earlier work this paper cites.
Spectrum envelopes for synthetic vowels
Tappert, C., Martony, J., & Fant, G. (1963) · 1963
Earlier work this paper cites.
Phase vocoder
Flanagan, J. L., & Golden, R. M. (1966) · 1966
Earlier work this paper cites.
Acoustic theory of speech production
Fant, G. (1970) · 1970
Earlier work this paper cites.
Linear prediction: A tutorial review
Makhoul, J. (1975) · 1975
Earlier work this paper cites.
Linear Prediction of Speech
Markel, J. D., & Gray, A. J. (1976) · 1976
Earlier work this paper cites.
A comparative performance study of several pitch detection algorithms
Rabiner, L., Cheng, M., Rosenberg, A., & McGonegal, C. (1976) · 1976
Earlier work this paper cites.
Speech analysis/synthesis based on a sinusoidal representation
McAulay, R., & Quatieri, T. (1986) · 1986
Earlier work this paper cites.
Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones
Moulines, E., & Charpentier, F. (1990) · 1990
Earlier work this paper cites.
Spectral modeling synthesis: A sound analysis/synthesis system based on a deterministic plus stochastic decomposition
Serra, X., & Smith, J. (1990) · 1990
Earlier work this paper cites.
CSR-I (WSJ0) Sennheiser LDC93S6B
Garofalo, J. S., Graff, D., Paul, D., & Pallett, D. (1993) · 1993
Earlier work this paper cites.
TIMIT acoustic phonetic continuous speech corpus
Garofolo, J. S., Lamel, L. F., Fisher, W. M., Fiscus, J. G., Pallett, D. S., Dahlgren, N. L., & Zue, V. (1993) · 1993
Earlier work this paper cites.
HNS: Speech modification based on a harmonic+noise model
Laroche, J., Stylianou, Y., & Moulines, E. (1993) · 1993
Earlier work this paper cites.
Acoustic characteristics of American English vowels
Hillenbrand, J., Getty, L. A., Clark, M. J., & Wheeler, K. (1995) · 1995
Earlier work this paper cites.
Speech analysis/synthesis and modification using an analysis-by-synthesis/overlap-add sinusoidal model
George, E. B., & Smith, M. J. (1997) · 1997
Earlier work this paper cites.
Improved phase vocoder time-scale modification of audio
Laroche, J., & Dolson, M. (1999) · 1999
Earlier work this paper cites.
A comparative study of empirical formulae for estimating vowel-formant bandwidths
Khodai-Joopari, M., & Clermont, F. (2002) · 2002
Earlier work this paper cites.
Time and pitch scale modification of audio signals
Laroche, J. (2002) · 2002
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Bishop, C. M. (2006) · 2006
Earlier work this paper cites.
STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds
Kawahara, H. (2006) · 2006
Earlier work this paper cites.
Implementation of realtime STRAIGHT speech manipulation system: Report on its first implementation
Banno, H., Hata, H., Morise, M., Takahashi, T., Irino, T., & Kawahara, H. (2007) · 2007
Earlier work this paper cites.
A sawtooth waveform inspired pitch estimator for speech and music
Camacho, A., & Harris, J. G. (2008) · 2008
Earlier work this paper cites.
Toronto emotional speech set (TESS)
Dupuis, K., & Pichora-Fuller, M. K. (2010) · 2010
Earlier work this paper cites.
Probing the independence of formant control using altered auditory feedback
MacDonald, E. N., Purcell, D. W., & Munhall, K. G. (2011) · 2011
Earlier work this paper cites.
A pitch tracking corpus with evaluation on multipitch tracking scenario
Pirker, G., Wohlmayr, M., Petrik, S., & Pernkopf, F. (2011) · 2011
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., & Vincent, P. (2013) · 2013
Earlier work this paper cites.
Demand: a collection of multi-channel recordings of acoustic noise in diverse environments
Thiemann, J., Ito, N., & Vincent, E. (2013) · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014) · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
Kingma, D. P., & Welling, M. (2014) · 2014
Cited alongside, same era.
pYIN: A fundamental frequency estimator using probabilistic threshold distributions
Mauch, M., & Dixon, S. (2014) · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., & Wierstra, D. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P., & Ba, J. (2015) · 2015
Cited alongside, same era.
librosa: Audio and music signal analysis in python
McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., & Nieto, O. (2015) · 2015
Cited alongside, same era.
PWLF: a Python library for fitting 1D continuous piecewise linear functions
Jekel, C. F., & Venter, G. (2019) · 2019
Later among the works it cites.
GlotNet—a raw waveform model for the glottal excitation in statistical parametric speech synthesis
Juvela, L., Bollepalli, B., Tsiaras, V., & Alku, P. (2019) · 2019
Later among the works it cites.
Adversarially trained end-to-end Korean singing voice synthesis system
Lee, J., Choi, H.-S., Jeon, C.-B., Koo, J., & Lee, K. (2019) · 2019
Later among the works it cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Schölkopf, B., & Bachem, O. (2019) · 2019
Later among the works it cites.
A statistically principled and computationally efficient approach to speech enhancement using variational autoencoders
Pariente, M., Deleforge, A., & Vincent, E. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Modeling and transforming speech using variational autoencoders
Blaauw, M., & Bonada, J. (2016) · 2016
Cited alongside, same era.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., & Abbeel, P. (2016) · 2016
Cited alongside, same era.
Voice conversion from non-parallel corpora using variational auto-encoder
Hsu, C.-C., Hwang, H.-T., Wu, Y.-C., Tsao, Y., & Wang, H.-M. (2016) · 2016
Cited alongside, same era.
Adversarial autoencoders
Makhzani, A., Shlens, J., Jaitly, N., Goodfellow, I., & Frey, B. (2016) · 2016
Cited alongside, same era.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Morise, M., Yokomori, F., & Ozawa, K. (2016) · 2016
Cited alongside, same era.
Compressed sensing using generative models
Bora, A., Jalal, A., Price, E., & Dimakis, A. G. (2017) · 2017
Cited alongside, same era.
Waveglow: A flow-based generative network for speech synthesis
Prenger, R., Valle, R., & Catanzaro, B. (2019) · 2019
Later among the works it cites.
Semi-supervised multichannel speech enhancement with a deep speech prior
Sekiguchi, K., Bando, Y., Nugraha, A. A., Yoshii, K., & Kawahara, T. (2019) · 2019
Later among the works it cites.
LPCNet: Improving neural speech synthesis through linear prediction
Valin, J.-M., & Skoglund, J. (2019) · 2019
Later among the works it cites.
Neural source-filter waveform models for statistical parametric speech synthesis
Wang, X., Takaki, S., & Yamagishi, J. (2019) · 2019
Later among the works it cites.
GANSpace: Discovering interpretable GAN controls
Härkönen, E., Hertzmann, A., Lehtinen, J., & Paris, S. (2020) · 2020
Later among the works it cites.
Source separation with deep generative priors
Jayaram, V., & Thickstun, J. (2020) · 2020
Later among the works it cites.
A recurrent variational autoencoder for speech enhancement
Leglaive, S., Alameda-Pineda, X., Girin, L., & Horaud, R. (2020) · 2020
Later among the works it cites.
Deep learning based assessment of synthetic speech naturalness
Mittag, G., & Möller, S. (2020) · 2020
Later among the works it cites.
torchcrepe
Morrison, M. (2020) · 2020
Later among the works it cites.
Controlling generative models with continuous factors of variations
Plumerault, A., Borgne, H. L., & Hudelot, C. (2020) · 2020
Later among the works it cites.
Unsupervised speech decomposition via triple information bottleneck
Qian, K., Zhang, Y., Chang, S., Hasegawa-Johnson, M., & Cox, D. (2020) · 2020
Later among the works it cites.
Speech enhancement with stochastic temporal convolutional networks
Richter, J., Carbajal, G., & Gerkmann, T. (2020) · 2020
Later among the works it cites.
Weakly supervised disentanglement with guarantees
Shu, R., Chen, Y., Kumar, A., Ermon, S., & Poole, B. (2020) · 2020
Later among the works it cites.
Disentanglement by nonlinear ICA with general incompressible-flow networks (GIN)
Sorrenson, P., Rother, C., & Köthe, U. (2020) · 2020
Later among the works it cites.
NVAE: A deep hierarchical variational autoencoder
Vahdat, A., & Kautz, J. (2020) · 2020
Later among the works it cites.
Hider-finder-combiner: An adversarial architecture for general speech signal modification
Webber, J. J., Perrotin, O., & King, S. (2020) · 2020
Later among the works it cites.
Praat: doing phonetics by computer [Computer program]
Boersma, P., & Weenink, D. (2021) · 2021
Later among the works it cites.
RAVE: A variational autoencoder for fast and high-quality neural audio synthesis
Caillon, A., & Esling, P. (2021) · 2021
Later among the works it cites.
Guided variational autoencoder for speech enhancement with a supervised classifier
Carbajal, G., Richter, J., & Gerkmann, T. (2021) · 2021
Later among the works it cites.
Neural analysis and synthesis: Reconstructing speech from self-supervised representations
Choi, H.-S., Lee, J., Kim, W., Lee, J. H., Heo, H., & Lee, K. (2021) · 2021
Later among the works it cites.
Variational autoencoder for speech enhancement with a noise-aware encoder
Fang, H., Carbajal, G., Wermter, S., & Gerkmann, T. (2021) · 2021
Later among the works it cites.
Dynamical variational autoencoders: A comprehensive review
Girin, L., Leglaive, S., Bie, X., Diard, J., Hueber, T., & Alameda-Pineda, X. (2021) · 2021
Later among the works it cites.
Neural pitch-shifting and time-stretching with controllable LPCNet
Morrison, M., Jin, Z., Bryan, N. J., Caceres, J.-P., & Pardo, B. (2021) · 2021
Later among the works it cites.
Unsupervised speech enhancement using dynamical variational autoencoders
Bie, X., Leglaive, S., Alameda-Pineda, X., & Girin, L. (2022) · 2022
Closest in time.