Fetching the paper…
Reading the bibliography…
Deep learning based visual to sound generation systems essentially need to be developed particularly considering the synchronicity aspects of visual and audio features with time.
J. Allen, “Short term spectral analysis, synthesis, and modification by discrete fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 25, no. 3, pp. 235–238, 1977
1977
Earlier work this paper cites.
R. Crochiere, “A weighted overlap-add method of short-time fourier analysis/synthesis,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 28, no. 1, pp. 99–102, 1980
1980
Earlier work this paper cites.
B. Shelton and C. Searle, “The influence of vision on the absolute identification of sound-source position,” Perception & Psychophysics , vol. 28, no. 6, pp. 589–596, 1980
1980
Earlier work this paper cites.
W. W. Gaver, “What in the world do we hear?: An ecological approach to auditory event perception,” Ecological psychology , vol. 5, no. 1, pp. 1–29, 1993
1993
Earlier work this paper cites.
D. R. Perrott, J. Cisneros, R. L. Mckinley, and W. R. D’Angelo, “Aurally aided visual search under virtual and free-field listening conditions,” Human factors , vol. 38, no. 4, pp. 702–715, 1996
1996
Earlier work this paper cites.
R. S. Bolia, W. R. D’Angelo, and R. L. McKinley, “Aurally aided visual search in three-dimensional space,” Human factors , vol. 41, no. 4, pp. 664–669, 1999
1999
Earlier work this paper cites.
K. Van Den Doel, P. G. Kry, and D. K. Pai, “Foleyautomatic: physically-based sound effects for interactive simulation and animation,” in Proceedings of the 28th annual conference on Computer graphics and interactive techniques . ACM, 2001, pp. 537–544
2001
Earlier work this paper cites.
P. Majdak, M. J. Goupell, and B. Laback, “3-d localization of virtual sound sources: effects of visual environment, pointing method, and training,” Attention, perception, & psychophysics , vol. 72, no. 2, pp. 454–469, 2010
2010
Earlier work this paper cites.
S. P. L. Prˇsa, Zdenek and, N. Holighaus, C. Wiesmeyr, and P. Balazs, “The large time-frequency analysis toolbox 2.0,” in International Symposium on Computer Music Multidisciplinary Research . Springer, 2013, pp. 419–442
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
E. Richardson and Y. Weiss, “On gans and gmms,” arXiv preprint arXiv:1805.12462 , 2018
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba, “Ambient sound provides supervision for visual learning,” in European conference on computer vision . Springer, 2016, pp. 801–816
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in Advances in neural information processing systems , 2016, pp. 892–900
2016
Earlier work this paper cites.
F. Wang, H. Nagano, K. Kashino, and T. Igarashi, “Visualizing video sounds with sound word animation to enrich user experience,” IEEE Transactions on Multimedia , vol. 19, no. 2, pp. 418–429, 2016
2016
Cited alongside, same era.
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman, “Visually indicated sounds,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2405–2413
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
D. Li, T. R. Langlois, and C. Zheng, “Scene-aware audio for 360 videos,” ACM Transactions on Graphics (TOG) , vol. 37, no. 4, pp. 1–12, 2018
2018
Later among the works it cites.
2019
Later among the works it cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4401–4410
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2017
Cited alongside, same era.
L. Chen, S. Srivastava, Z. Duan, and C. Xu, “Deep cross-modal audio-visual generation,” in Proceedings of the on Thematic Workshops of ACM Multimedia 2017 . ACM, 2017, pp. 349–357
2017
Cited alongside, same era.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 609–617
2017
Cited alongside, same era.
N. Takahashi, M. Gygli, and L. Van Gool, “Aenet: Learning deep audio features for video analysis,” IEEE Transactions on Multimedia , vol. 20, no. 3, pp. 513–524, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
T. Yu, L. Wang, C. Da, H. Gu, S. Xiang, and C. Pan, “Weakly semantic guided action recognition,” IEEE Transactions on Multimedia , 2019
2019
Later among the works it cites.
R. Gao and K. Grauman, “2.5 d visual sound,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 324–333
2019
Later among the works it cites.
H. Zhou, Z. Liu, X. Xu, P. Luo, and X. Wang, “Vision-infused deep audio inpainting,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 283–292
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, “Self-attention generative adversarial networks,” in International conference on machine learning . PMLR, 2019, pp. 7354–7363
2019
Later among the works it cites.
S. Ghose and J. J. Prevost, “Autofoley: Artificial synthesis of synchronized sound tracks for silent videos with deep learning,” IEEE Transactions on Multimedia , pp. 1–1, 2020
2020
Later among the works it cites.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of stylegan,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 8110–8119
2020
Later among the works it cites.
P. Chen, Y. Zhang, M. Tan, H. Xiao, D. Huang, and C. Gan, “Generating visually aligned sound from videos,” IEEE Transactions on Image Processing , vol. 29, pp. 8292–8302, 2020
2020
Later among the works it cites.
S. Ghose and J. J. Prevost, “Enabling an iot system of systems through auto sound synthesis in silent video with dnn,” in 2020 IEEE 15th International Conference of System of Systems Engineering (SoSE) . IEEE, 2020, pp. 563–568
2020
Later among the works it cites.
K. N. Haque, R. Rana, and B. W. Schuller, “High-fidelity audio generation and representation learning with guided adversarial autoencoder,” IEEE Access , vol. 8, pp. 223 509–223 528, 2020
2020
Later among the works it cites.
S. Liu, S. Li, and H. Cheng, “Towards an end-to-end visual-to-raw-audio generation with gan,” IEEE Transactions on Circuits and Systems for Video Technology , 2021
2021
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International conference on machine learning , 2015, pp. 2048–2057
2057
Closest in time.