Fetching the paper…
Reading the bibliography…
Human perceives rich auditory experience with distinct sound heard by ears.
Signal estimation from modified short-time Fourier transform
Griffin, D.; and Jae Lim. 1983 · 1983
Earlier work this paper cites.
Learning Joint Statistical Models for Audio-Visual Fusion and Segregation
Fisher III, J. W.; Darrell, T.; Freeman, W. T.; and Viola, P. A. 2001 · 2001
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Ronneberger, O.; P.Fischer; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Visually indicated sounds
Owens, A.; Isola, P.; McDermott, J.; Torralba, A.; Adelson, E. H.; and Freeman, W. T. 2016 · 2016
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Chen, L.; Srivastava, S.; Duan, Z.; and Xu, C. 2017 · 2017
Earlier work this paper cites.
Vid2Speech: speech reconstruction from silent video
Ephrat, A.; and Peleg, S. 2017 · 2017
Earlier work this paper cites.
FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
Ilg, E.; Mayer, N.; Saikia, T.; Keuper, M.; Dosovitskiy, A.; and Brox, T. 2017 · 2017
Earlier work this paper cites.
Visually indicated sound generation by perceptually optimized classification
Chen, K.; Zhang, C.; Fang, C.; Wang, Z.; Bui, T.; and Nevatia, R. 2018 · 2018
Earlier work this paper cites.
Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation
Ephrat, A.; Mosseri, I.; Lang, O.; Dekel, T.; Wilson, K.; Hassidim, A.; Freeman, W. T.; and Rubinstein, M. 2018 · 2018
Cited alongside, same era.
Learning to Separate Object Sounds by Watching Unlabeled Video
Gao, R.; Feris, R.; and Grauman, K. 2018 · 2018
Cited alongside, same era.
Cmcgan: A uniform framework for cross-modal visual-audio mutual generation
Hao, W.; Zhang, Z.; and Guan, H. 2018 · 2018
Cited alongside, same era.
Scene-Aware Audio for 360 Videos
Li, D.; Langlois, T. R.; and Zheng, C. 2018 · 2018
Cited alongside, same era.
Audio-Visual Scene Analysis with Self-Supervised Multisensory Features
Owens, A.; and Efros, A. A. 2018 · 2018
Cited alongside, same era.
Self-Supervised Generation of Spatial Audio for 360+ Video
Pedro Morgado, Nuno Vasconcelos, T. L.; and Wang, O. 2018 · 2018
Immersive Spatial Audio Reproduction for VR/AR Using Room Acoustic Modelling from 360 Images
Kim, H.; Remaggi, L.; Jackson, P. J.; and Hilton, A. 2019 · 2019
Later among the works it cites.
Dual-modality seq2seq network for audio-visual event localization
Lin, Y.-B.; Li, Y.-J.; and Wang, Y.-C. F. 2019 · 2019
Later among the works it cites.
Self-supervised Audio Spatialization with Correspondence Classifier
Lu, Y.-D.; Lee, H.-Y.; Tseng, H.-Y.; and Yang, M.-H. 2019 · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 2019
Later among the works it cites.
Recursive Visual Sound Separation Using Minus-Plus Net
Xu, X.; Dai, B.; and Lin, D. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Audio-Visual Event Localization in Unconstrained Videos
Tian, Y.; Shi, J.; Li, B.; Duan, Z.; and Xu, C. 2018 · 2018
Cited alongside, same era.
The Sound of Pixels
Zhao, H.; Gan, C.; Rouditchenko, A.; Vondrick, C.; McDermott, J.; and Torralba, A. 2018 · 2018
Cited alongside, same era.
Visual to Sound: Generating Natural Sound for Videos in the Wild
Zhou, Y.; Wang, Z.; Fang, C.; Bui, T.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
Self-supervised moving vehicle tracking with stereo sound
Gan, C.; Zhao, H.; Chen, P.; Cox, D.; and Torralba, A. 2019 · 2019
Cited alongside, same era.
Foley music: Learning to generate music from videos
Gan, C.; Huang, D.; Chen, P.; Tenenbaum, J. B.; and Torralba, A. 2020a
Cited in the paper.
Music Gesture for Visual Sound Separation
Gan, C.; Huang, D.; Zhao, H.; Tenenbaum, J. B.; and Torralba, A. 2020b
Cited in the paper.
Zhao, H.; Gan, C.; Ma, W.-C.; and Torralba, A. 2019 · 2019
Later among the works it cites.
Vision-Infused Deep Audio Inpainting
Zhou, H.; Liu, Z.; Xu, X.; Luo, P.; and Wang, X. 2019 · 2019
Later among the works it cites.
Generating visually aligned sound from videos
Chen, P.; Zhang, Y.; Tan, M.; Xiao, H.; Huang, D.; and Gan, C. 2020 · 2020
Later among the works it cites.
Audiovisual Transformer with Instance Attention for Audio-Visual Event Localization
Lin, Y.-B.; and Wang, Y.-C. F. 2020 · 2020
Later among the works it cites.
Unified multisensory perception: weakly-supervised audio-visual video parsing
Tian, Y.; Li, D.; and Xu, C. 2020 · 2020
Later among the works it cites.