Fetching the paper…
Reading the bibliography…
We introduce an approach to convert mono audio recorded by a 360 video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere.
Probable inference, the law of succession, and statistical inference
E. B. Wilson · 1927
Earlier work this paper cites.
Periphony: With-height sound reproduction
M. A. Gerzon · 1973
Earlier work this paper cites.
Blind separation of sources, part I: An adaptive algorithm based on neuromimetic architecture
C. Jutten and J. Herault · 1991
Earlier work this paper cites.
Independent component analysis, a new concept?
P. Comon · 1994
Earlier work this paper cites.
An information-maximization approach to blind separation and blind deconvolution
A. J. Bell and T. J. Sejnowski · 1995
Earlier work this paper cites.
A new learning algorithm for blind signal separation
S. Amari, A. Cichocki, and H. H. Yang · 1996
Earlier work this paper cites.
Towards optimal soundfield representation
G. Dickins and R. Kennedy · 1999
Earlier work this paper cites.
The earth mover’s distance is the mallows distance: Some insights from statistics
E. Levina and P. Bickel · 2001
Earlier work this paper cites.
Joint audio-video object localization and tracking
N. Strobel, S. Spors, and R. Rabenstein · 2001
Earlier work this paper cites.
Real-time sound source localization and separation for robot audition
K. Nakadai, H. G. Okuno, and H. Kitano · 2002
Earlier work this paper cites.
Pixels that sound
E. Kidron, Y. Y. Schechner, and M. Elad · 2005
Earlier work this paper cites.
Sound localization for humanoid robots-building audio-motor maps based on the HRTF
J. Hornstein, M. Lopes, J. Santos-Victor, and F. Lacerda · 2006
Earlier work this paper cites.
Harmony in motion
Z. Barzelay and Y. Y. Schechner · 2007
Earlier work this paper cites.
Cross-modal localization via sparsity
E. Kidron, Y. Y. Schechner, and M. Elad · 2007
Earlier work this paper cites.
Robust localization and tracking of simultaneous moving sound sources using beamforming and particle filtering
J.-M. Valin, F. Michaud, and J. Rouat · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
Intelligent sound source localization and its application to multimodal human tracking
K. Nakamura, K. Nakadai, F. Asano, and G. Ince · 2011
Cited alongside, same era.
Directional perception of distributed sound sources
O. Santala and V. Pulkki · 2011
Cited alongside, same era.
Multimodal analysis for identification and segmentation of moving-sounding objects
H. Izadinia, I. Saleemi, and M. Shah · 2013
Cited alongside, same era.
Learning a deep convolutional network for image super-resolution
C. Dong, C. C. Loy, K. He, and X. Tang · 2014
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Accurate image super-resolution using very deep convolutional networks
J. Kim, J. Kwon Lee, and K. Mu Lee · 2016
Later among the works it cites.
Wavenet: A generative model for raw audio
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Later among the works it cites.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2016
Later among the works it cites.
Suggesting sounds for images from video collections
M. Soler, J.-C. Bazin, O. Wang, A. Krause, and A. Sorkine-Hornung · 2016
Later among the works it cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Later among the works it cites.
CNN architectures for large-scale audio classification
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning for monaural speech separation
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2014
Cited alongside, same era.
Spatial transformations for the alteration of ambisonic recordings
M. Kronlachner · 2014
Cited alongside, same era.
A survey on sound source localization in robotics: From binaural to array processing methods
S. Argentieri, P. Danès, and P. Souères · 2015
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Cited alongside, same era.
Flownet 2.0: Evolution of optical flow estimation with deep networks
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox · 2017
Later among the works it cites.
Audio super resolution using neural networks
V. Kuleshov, S. Z. Enam, and S. Ermon · 2017
Later among the works it cites.
Voxceleb: A large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2017
Later among the works it cites.
The conversation: Deep audio-visual speech enhancement
T. Afouras, J. S. Chung, and A. Zisserman · 2018
Closest in time.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. Freeman, and M. Rubinstein · 2018
Closest in time.
Visual speech enhancement using noise-invariant training
A. Gabbay, A. Shamir, and S. Peleg · 2018
Closest in time.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Closest in time.
Supervised speech separation based on deep learning: An overview
D. Wang and J. Chen · 2018
Closest in time.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Closest in time.