Fetching the paper…
Reading the bibliography…
In this work, we study music/video cross-modal recommendation, i.e.
Inskip, C., Macfarlane, A., Rafferty, P.: Music, Movies and Meaning: Communication in Film-makers’ Search for Pre-existing Music, and the Implications for Music Information Retrieval. In: Proceedings of ISMIR (International Conference on Music Information Retrieval). Philadelphia, PA, USA (2008)
2008
Earlier work this paper cites.
Liao, C., Wang, P.P., Zhang, Y.: Mining Association Patterns between Music and Video Clips in Professional MTV. In: Proceedings of MMM (International Conference on Multimedia Modeling). Sophia Antipolis, France (2009). https://doi.org/10.1007/978-3-540-92892-8_41
2009
Earlier work this paper cites.
Weinberger, K.Q., Saul, L.K.: Distance metric learning for large margin nearest neighbor classification. Journal of Machine Learning Research 10
2009
Earlier work this paper cites.
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., Ng, A.Y.: Multimodal Deep Learning. In: Proceedings of ICML (International Conference on Machine Learning). Bellevue, WA, USA (2011)
2011
Earlier work this paper cites.
Kuo, F.F., Shan, M.K., Lee, S.Y.: Background Music Recommendation for Video Based on Multimodal Latent Semantic Analysis. In: Proceedings of ICME (International Conference on Multimedia and Expo). San Jose, CA, USA (2013)
2013
Earlier work this paper cites.
Oquab, M., Bottou, L., Laptev, I., Sivic, J.: Learning and Transferring Mid-Level Image Representations using Convolutional Neural Networks. In: Proceedings of IEEE CVPR (Conference on Computer Vision and Pattern Recognition). Columbus, OH, USA (2014)
2014
Earlier work this paper cites.
Shah, R.R., Yu, Y., Zimmermann, R.: ADVISOR - Personalized video soundtrack recommendation by late fusion with heuristic rankings. In: Proceedings of ACM Multimedia. Orlando, FL, USA (2014). https://doi.org/10.1145/2647868.2654919
2014
Earlier work this paper cites.
Wang, J., Song, Y., Leung, T., Rosenberg, C., Wang, J., Philbin, J., Chen, B., Wu, Y.: Learning Fine-grained Image Similarity with Deep Ranking. In: Proceedings of IEEE CVPR (Conference on Computer Vision and Pattern Recognition). Columbus, OH, USA (2014)
2014
Earlier work this paper cites.
McFee, B., Raffel, C., Liang, D., Ellis, D.P., McVicar, M., Battenberg, E., Nieto, O.: librosa: Audio and music signal analysis in python. In: Proceedings of Python in Science. Austin, TX, USA (2015)
2015
Earlier work this paper cites.
Sasaki, S., Hirai, T., Ohya, H., Morishima, S.: Affective Music Recommendation System Based on the Mood of Input Video. In: Proceedings of MMM (International Conference on Multimedia Modeling). Sydney, Australia (2015)
2015
Earlier work this paper cites.
Yue, J., Ng, H., Hausknecht, M., Vijayanarasimhan, S., Vinyals, O., Monga, R., Toderici, G.: Beyond Short Snippets: Deep Networks for Video Classification. In: Proceedings of IEEE CVPR (Conference on Computer Vision and Pattern Recognition). Boston, MA, USA (2015)
2015
Earlier work this paper cites.
2016
Cited alongside, same era.
Aytar, Y., Vondrick, C., Torralba, A.: SoundNet: Learning Sound Representations from Unlabeled Video. Advances in neural information processing systems pp. 892–900 (2016)
2016
Cited alongside, same era.
Owens, A., Wu, J., McDermott, J.H., Freeman, W.T., Torralba, A.: Ambient sound provides supervision for visual learning. In: Proceedings of ECCV (European Conference on Computer Vision). Amsterdam, The Netherlands (2016). https://doi.org/10.1007/978-3-319-46448-0_48
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Hong, S., Im, W., Yang, H.S.: CBVMR: Content-based video-music retrieval using soft intra-modal structure constraint. In: Proceedings of ACM ICMR (International Conference on Multimedia Retrieval). Yokohama, Japan (2018)
2018
Later among the works it cites.
Müller, M., Arzt, A., Balke, S., Dorfer, M., Widmer, G.: Cross-Modal Music Retrieval and Applications: An Overview of Key Methodologies. In: Proceedings of IEEE ICASSP (International Conference on Computer Vision). Brighton, UK (2019). https://doi.org/10.1109/MSP.2018.2868887
2018
Later among the works it cites.
Owens, A., Efros, A.A.: Audio-Visual Scene Analysis with Self-Supervised Multisensory Features. In: Proceedings of IEEE CVPR (Conference on Computer Vision and Pattern Recognition). Salt Lake City, UT, USA (2018)
2018
Later among the works it cites.
Tian, Y., Shi, J., Li, B., Duan, Z., Xu, C.: Audio-visual event localization in unconstrained videos. In: Proceedings of ECCV (European Conference on Computer Vision). Munich, Germany (2018)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arandjelovic, R., Zisserman, A.: Look, Listen and Learn. In: Proceedings of IEEE ICASSP (International Conference on Computer Vision). Venice, Italy (2017). https://doi.org/10.1109/ICCV.2017.73
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Gemmeke, J.F., Ellis, D.P.W., Freedman, D., Jansen, A., Lawrence, W., Channing Moore, R., Plakal, M., Ritter, M.: Audio Set: An Ontology and Human-Labeled Dataset for Audio Events. In: Proceedings of IEEE ICASSP (International Conference on Computer Vision). New Orleans, LA, USA (2017)
2017
Cited alongside, same era.
Schindler, A., Rauber, A.: Harnessing music-related visual stereotypes for music information retrieval. ACM Transactions on Intelligent Systems and Technology 8
2017
Cited alongside, same era.
Shin, K.H., Lee, I.K.: Music synchronization with video using emotion similarity. In: Proceedings of IEEE BigComp (International Conference on Big Data and Smart Computing). Jeju Island, South Korea (2017)
2017
Cited alongside, same era.
Arandjelović, R., Zisserman, A.: Objects that Sound. In: Proceedings of ECCV (European Conference on Computer Vision). Munich, Germany (2018). https://doi.org/10.1007/978-3-030-01246-5_27
2018
Cited alongside, same era.
Balntas, V., Li, S., Prisacariu, V.: RelocNet: Continuous Metric Learning Relocalisation using Neural Nets. In: Proceedings of ECCV (European Conference on Computer Vision). Munich, Germany (2018)
2018
Cited alongside, same era.
2018
Later among the works it cites.
Zeng, D., Yu, Y., Oyama, K.: Audio-Visual Embedding for Cross-Modal Music Video Retrieval through Supervised Deep CCA. In: Proceedings of IEEE ISM (International Symposium on Multimedia). Taichung, Taiwan (2018). https://doi.org/10.1109/ism.2018.00-21
2018
Later among the works it cites.
Cramer, J., Wu, H.H., Salamon, J., Bello, J.P.: Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings. In: Proceedings of IEEE ICASSP (International Conference on Computer Vision). Brighton, UK (2019)
2019
Later among the works it cites.
Li, B., Kumar, A.: Query by Video: Cross-Modal Music Retrieval. In: Proceedings of ISMIR (International Conference on Music Information Retrieval). Delft, The Netherlands (2019)
2019
Later among the works it cites.
Parekh, S., Essid, S., Ozerov, A., Duong, N.Q.K., Pérez, P., Richard, G.: Weakly Supervised Representation Learning for Audio-Visual Scene Analysis. IEEE/ACM Transactions on Audio, Speech, and Language Processing (2019)
2019
Later among the works it cites.
Pons, J., Serra, X.: musicnn: Pre-trained convolutional neural networks for music audio tagging. In: Late Breaking Demo, ISMIR (International Conference on Music Information Retrieval). Delft, The Netherlands (2019)
2019
Later among the works it cites.
Schindler, A.: Multi-Modal Music Information Retrieval: Augmenting Audio-Analysis with Visual Computing for Improved Music Video Analysis. Ph.D. thesis, Technische Universität Wien (2019)
2019
Later among the works it cites.
Prétet, L., Richard, G., Peeters, G.: Learning to Rank Music Tracks Using Triplet Loss. In: Proceedings of IEEE ICASSP (International Conference on Computer Vision). Barcelona, Spain (2020)
2020
Later among the works it cites.