Fetching the paper…
Reading the bibliography…
In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments.
McGurk, H., MacDonald, J.: Hearing lips and seeing voices. Nature 264
1976
Earlier work this paper cites.
Hershey, J.R., Movellan, J.R.: Audio vision: Using audio-visual synchrony to locate sounds. In: Solla, S.A., Leen, T.K., Müller, K. (eds.) Advances in Neural Information Processing Systems 12, pp. 813–819 (2000)
2000
Earlier work this paper cites.
Godøy, R.I., Leman, M.: Musical gestures: Sound, movement, and meaning. Routledge (2010)
2010
Earlier work this paper cites.
Izadinia, H., Saleemi, I., Shah, M.: Multimodal analysis for identification and segmentation of moving-sounding objects. IEEE Transactions on Multimedia 15
2013
Earlier work this paper cites.
Aytar, Y., Vondrick, C., Torralba, A.: Soundnet: Learning sound representations from unlabeled video. In: Advances in Neural Information Processing Systems. pp. 892–900 (2016)
2016
Earlier work this paper cites.
Owens, A., Isola, P., McDermott, J., Torralba, A., Adelson, E.H., Freeman, W.T.: Visually indicated sounds. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2405–2413 (2016)
2016
Earlier work this paper cites.
Owens, A., Wu, J., McDermott, J.H., Freeman, W.T., Torralba, A.: Ambient sound provides supervision for visual learning. In: European Conference on Computer Vision. pp. 801–816. Springer (2016)
2016
Earlier work this paper cites.
Waite, E., et al.: Generating long-term structure in songs and stories. Web blog post. Magenta 15
2016
Earlier work this paper cites.
Arandjelovic, R., Zisserman, A.: Look, listen and learn. In: 2017 IEEE International Conference on Computer Vision (ICCV). pp. 609–617. IEEE (2017)
2017
Earlier work this paper cites.
Arandjelović, R., Zisserman, A.: Objects that sound. arXiv preprint arXiv:1712.06651 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Computer Vision and Pattern Recognition (CVPR), 2017 IEEE Conference on. pp. 4724–4733. IEEE (2017)
2017
Earlier work this paper cites.
Chen, L., Srivastava, S., Duan, Z., Xu, C.: Deep cross-modal audio-visual generation. In: ACM Multimedia 2017. pp. 349–357 (2017)
2017
Earlier work this paper cites.
Chu, H., Urtasun, R., Fidler, S.: Song from pi: A musically plausible network for pop music generation. ICLR (2017)
2017
Earlier work this paper cites.
Chung, J.S., Senior, A.W., Vinyals, O., Zisserman, A.: Lip reading sentences in the wild. In: CVPR. pp. 3444–3453 (2017)
2017
Earlier work this paper cites.
Hadjeres, G., Pachet, F., Nielsen, F.: Deepbach: a steerable model for bach chorales generation. In: ICML. pp. 1362–1371 (2017)
2017
Earlier work this paper cites.
Karras, T., Aila, T., Laine, S., Herva, A., Lehtinen, J.: Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics (TOG) 36
2017
Earlier work this paper cites.
Oord, A.v.d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., Kavukcuoglu, K.: Wavenet: A generative model for raw audio. ICLR (2017)
2017
Earlier work this paper cites.
Simon, T., Joo, H., Matthews, I., Sheikh, Y.: Hand keypoint detection in single images using multiview bootstrapping. In: CVPR (2017)
2017
Earlier work this paper cites.
Suwajanakorn, S., Seitz, S.M., Kemelmacher-Shlizerman, I.: Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics (TOG) 36
2017
Earlier work this paper cites.
Taylor, S., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A.G., Hodgins, J., Matthews, I.: A deep learning approach for generalized speech animation. ACM Transactions on Graphics (TOG) 36
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: NIPS. pp. 5998–6008 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Albanie, S., Nagrani, A., Vedaldi, A., Zisserman, A.: Emotion recognition in speech using cross-modal transfer in the wild. ACM Multimedia (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Chen, K., Zhang, C., Fang, C., Wang, Z., Bui, T., Nevatia, R.: Visually indicated sound generation by perceptually optimized classification. In: ECCV. vol. 11134, pp. 560–574 (2018)
2018
Cited alongside, same era.
Chen, K., Zhang, C., Fang, C., Wang, Z., Bui, T., Nevatia, R.: Visually indicated sound generation by perceptually optimized classification. In: The European Conference on Computer Vision. pp. 560–574 (2018)
2018
Cited alongside, same era.
2018
Later among the works it cites.
Engel, J.H., Agrawal, K.K., Chen, S., Gulrajani, I., Donahue, C., Roberts, A.: GANSynth: Adversarial neural audio synthesis. In: ICLR (2019)
2019
Later among the works it cites.
Gan, C., Zhao, H., Chen, P., Cox, D., Torralba, A.: Self-supervised moving vehicle tracking with stereo sound. In: ICCV. pp. 7053–7062 (2019)
2019
Later among the works it cites.
Ginosar, S., Bar, A., Kohavi, G., Chan, C., Owens, A., Malik, J.: Learning individual styles of conversational gesture. In: CVPR. pp. 3497–3506 (2019)
2019
Later among the works it cites.
Hawthorne, C., Stasyuk, A., Roberts, A., Simon, I., Huang, C.Z.A., Dieleman, S., Elsen, E., Engel, J., Eck, D.: Enabling factorized piano music modeling and generation with the maestro dataset. ICLR (2019)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W.T., Rubinstein, M.: Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation. ACM Transactions on Graphics (TOG) 37
2018
Cited alongside, same era.
Gao, R., Feris, R., Grauman, K.: Learning to separate object sounds by watching unlabeled video. In: ECCV. pp. 35–53 (2018)
2018
Cited alongside, same era.
Gao, R., Grauman, K.: 2.5 d visual sound. arXiv preprint arXiv:1812.04204 (2018)
2018
Cited alongside, same era.
Huang, C.Z.A., Vaswani, A., Uszkoreit, J., Simon, I., Hawthorne, C., Shazeer, N., Dai, A.M., Hoffman, M.D., Dinculescu, M., Eck, D.: Music transformer: Generating music with long-term structure (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Li, B., Liu, X., Dinesh, K., Duan, Z., Sharma, G.: Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications. IEEE Transactions on Multimedia 21
2018
Cited alongside, same era.
Long, X., Gan, C., De Melo, G., Liu, X., Li, Y., Li, F., Wen, S.: Multimodal keyless attention fusion for video classification. In: AAAI (2018)
2018
Cited alongside, same era.
Long, X., Gan, C., de Melo, G., Wu, J., Liu, X., Wen, S.: Attention clusters: Purely attention based local feature integration for video classification. In: CVPR (2018)
2018
Cited alongside, same era.
2019
Later among the works it cites.
Jamaludin, A., Chung, J.S., Zisserman, A.: You said that?: Synthesising talking faces from audio. International Journal of Computer Vision pp. 1–13 (2019)
2019
Later among the works it cites.
Rouditchenko, A., Zhao, H., Gan, C., McDermott, J., Torralba, A.: Self-supervised audio-visual co-segmentation. In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 2357–2361. IEEE (2019)
2019
Later among the works it cites.
Xu, X., Dai, B., Lin, D.: Recursive visual sound separation using minus-plus net. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 882–891 (2019)
2019
Later among the works it cites.
Zhao, H., Gan, C., Ma, W.C., Torralba, A.: The sound of motions. ICCV (2019)
2019
Later among the works it cites.
Zhao, K., Li, S., Cai, J., Wang, H., Wang, J.: An emotional symbolic music generation system based on lstm networks. In: 2019 IEEE 3rd Information Technology, Networking, Electronic and Automation Control Conference (ITNEC). pp. 2039–2043 (2019)
2019
Later among the works it cites.
Zhou, H., Liu, Z., Xu, X., Luo, P., Wang, X.: Vision-infused deep audio inpainting. In: ICCV. pp. 283–292 (2019)
2019
Later among the works it cites.
Gan, C., Huang, D., Zhao, H., Tenenbaum, J.B., Torralba, A.: Music gesture for visual sound separation. In: CVPR. pp. 10478–10487 (2020)
2020
Closest in time.
2020
Closest in time.
Gan, C., Zhang, Y., Wu, J., Gong, B., Tenenbaum, J.B.: Look, listen, and act: Towards audio-visual embodied navigation. ICRA (2020)
2020
Closest in time.
Gao, R., Oh, T.H., Grauman, K., Torresani, L.: Listen to look: Action recognition by previewing audio. In: CVPR. pp. 10457–10467 (2020)
2020
Closest in time.
Hu, D., Li, X., Mou, L., Jin, P., Chen, D., Jing, L., Zhu, X., Dou, D.: Cross-task transfer for multimodal aerial scene recognition. ECCV (2020)
2020
Closest in time.
Koepke, A.S., Wiles, O., Moses, Y., Zisserman, A.: Sight to sound: An end-to-end approach for visual piano transcription. In: ICASSP. pp. 1838–1842 (2020)
2020
Closest in time.
Peihao, C., Yang, Z., Mingkui, T., Hongdong, X., Deng, H., Chuang, G.: Generating visually aligned sound from videos. IEEE Transactions on Image Processing (October 2020)
2020
Closest in time.
2020
Closest in time.
Submission, A.: At your fingertips: Automatic piano fingering detection. In: ICLR (2020)
2020
Closest in time.
Tian, Y., Krishnan, D., Isola, P.: Contrastive multiview coding. ECCV (2020)
2020
Closest in time.
Zhou, H., Xu, X., Lin, D., Wang, X., Liu, Z.: Sep-stereo: Visually guided stereophonic audio generation by associating source separation. In: ECCV (2020)
2020
Closest in time.