Fetching the paper…
Reading the bibliography…
As two of the five traditional human senses (sight, hearing, taste, smell, and touch), vision and sound are basic sources through which humans understand the world.
Unit selection in a concatenative speech synthesis system using a large speech database
A. Hunt and A. Black · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Simultaneous modeling of spectrum, pitch and duration in hmm-based speech synthesis
T. Yoshimura, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Robust speaker-adaptive hmm-based text-to-speech synthesis
J. Yamagishi, T. Nose, H. Zen, Z. H. Ling, T. Toda, K. Tokuda, S. King, and S. Renals · 2009
Earlier work this paper cites.
Statistical parametric speech synthesis
H. Zen, K. Tokuda, and A. W. Black · 2009
Earlier work this paper cites.
Secrets of optical flow estimation and their principles
D. Sun, S. Roth, and M. J. Black · 2010
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, M. Shah, K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Statistical parametric speech synthesis using deep neural networks
H. Zen, A. Senior, and M. Schuster · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Çaglar Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Samplernn: An unconditional end-to-end neural audio generation model
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. C. Courville, and Y. Bengio · 2016
Later among the works it cites.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. Adelson, and W. Freeman · 2016
Later among the works it cites.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Later among the works it cites.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu · 2016
Later among the works it cites.
Look, listen and learn
R. Arandjelović and A. Zisserman · 2017
Closest in time.
Deep cross-modal audio-visual generation
L. Chen, S. Srivastava, Z. Duan, and C. Xu · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Cited alongside, same era.
Unsupervised learning of spoken language with visual context
D. Harwath, A. Torralba, and J. R. Glass · 2016
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Closest in time.
Char2wav: End-to-end speech synthesis
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio · 2017
Closest in time.
Generative modeling of audible shapes for object perception
Z. Zhang, J. Wu, Q. Li, Z. Huang, J. Traer, J. H. McDermott, J. B. Tenenbaum, and W. T. Freeman · 2017
Closest in time.