Fetching the paper…
Reading the bibliography…
We present a factorized hierarchical variational autoencoder, which learns disentangled and interpretable representations from sequential data without supervision.
The design for the wall street journal-based csr corpus
Douglas B Paul and Janet M Baker · 1992
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1
John S Garofolo, Lori F Lamel, William M Fisher, Jonathon G Fiscus, and David S Pallett · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
An introduction to variational methods for graphical models
Michael I Jordan, Zoubin Ghahramani, Tommi S Jaakkola, and Lawrence K Saul · 1999
Earlier work this paper cites.
Aurora working group: DSR front end LVCSR evaluation AU/384/02
David Pearce · 2002
Earlier work this paper cites.
Automatic speech recognition: An auditory perspective
Nelson Mogran, Hervé Bourlard, and Hynek Hermansky · 2004
Earlier work this paper cites.
Support vector machines versus fast scoring in the low-dimensional total variability space for speaker verification
Najim Dehak, Reda Dehak, Patrick Kenny, Niko Brümmer, Pierre Ouellet, and Pierre Dumouchel · 2009
Earlier work this paper cites.
Front-end factor analysis for speaker verification
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Hybrid speech recognition with deep bidirectional LSTM
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Conditional restricted boltzmann machine for voice conversion
Zhizheng Wu, Eng Siong Chng, and Haizhou Li · 2013
Earlier work this paper cites.
Feature learning in deep neural networks – studies on speech recognition tasks
Dong Yu, Michael Seltzer, Jinyu Li, Jui-Ting Huang, and Frank Seide · 2013
Earlier work this paper cites.
Variational recurrent auto-encoders
Otto Fabius and Joost R van Amersfoort · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Semi-supervised learning with deep generative models
Diederik P Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Hasim Sak, Andrew W Senior, and Françoise Beaufays · 2014
Cited alongside, same era.
An introduction to computational networks and the computational network toolkit
Dong Yu, Adam Eversole, Mike Seltzer, Kaisheng Yao, Zhiheng Huang, Brian Guenter, Oleksii Kuchaiev, Yu Zhang, Frank Seide, Huaming Wang, et al · 2014
Cited alongside, same era.
Sequential neural models with stochastic layers
Marco Fraccaro, Søren Kaae Sønderby, Ulrich Paquet, and Ole Winther · 2016
Later among the works it cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Later among the works it cites.
Composing graphical models with neural networks for structured representations and fast inference
Matthew Johnson, David K Duvenaud, Alex Wiltschko, Ryan P Adams, and Sandeep R Datta · 2016
Later among the works it cites.
Improved variational inference with inverse autoregressive flow
Diederik P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling · 2016
Later among the works it cites.
Non-parallel training in voice conversion using an adaptive restricted boltzmann machine
Toru Nakashika, Tetsuya Takiguchi, Yasuhiro Minami, Toru Nakashika, Tetsuya Takiguchi, and Yasuhiro Minami · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gated feedback recurrent neural networks
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
A recurrent latent variable model for sequential data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio · 2015
Cited alongside, same era.
Deep convolutional inverse graphics network
Tejas D Kulkarni, William F Whitney, Pushmeet Kohli, and Josh Tenenbaum · 2015
Cited alongside, same era.
Autoencoding beyond pixels using a learned similarity metric
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther · 2015
Cited alongside, same era.
Alireza Makhzani, Jonathon Shlens, Navdeep Jaitly, Ian Goodfellow, and Brendan Frey · 2015
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Cited alongside, same era.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Invariant representations for noisy speech recognition
Dmitriy Serdyuk, Kartik Audhkhasi, Philemon Brakel, Bhuvana Ramabhadran, Samuel Thomas, and Yoshua Bengio · 2016
Later among the works it cites.
Adversarial multi-task learning of deep neural networks for robust speech recognition
Yusuke Shunohara · 2016
Later among the works it cites.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Highway long short-term memory RNNs for distant speech recognition
Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yaco, Sanjeev Khudanpur, and James Glass · 2016
Later among the works it cites.
Learning latent representations for speech generation and transformation
Wei-Ning Hsu, Yu Zhang, and James Glass · 2017
Closest in time.
Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentation
Wei-Ning Hsu, Yu Zhang, and James Glass · 2017
Closest in time.
Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov, and Eric P Xing · 2017
Closest in time.
Non-parallel voice conversion using i-vector plda: Towards unifying speaker verification and transformation
Tomi Kinnunen, Lauri Juvela, Paavo Alku, and Junichi Yamagishi · 2017
Closest in time.
A hierarchical latent variable encoder-decoder model for generating dialogues
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Closest in time.