Fetching the paper…
Reading the bibliography…
Up to now, only limited research has been conducted on cross-modal retrieval of suitable music for a specified video or vice versa.
B. Shaw, B. Huang, and T. Jebara, “Learning a distance metric from a network,” in Advances in Neural Information Processing Systems , 2011, pp. 1899–1907
1907
Earlier work this paper cites.
J. A. Wegelin et al. , “A survey of partial least squares (pls) methods, with emphasis on the two-block case,” University of Washington, Department of Statistics, Tech. Rep , 2000
2000
Earlier work this paper cites.
I. Jolliffe, Principal component analysis . Wiley Online Library, 2002
2002
Earlier work this paper cites.
E. Brochu, N. De Freitas, and K. Bao, “The sound of an album cover: Probabilistic multimedia and ir,” in Workshop on Artificial Intelligence and Statistics , 2003
2003
Earlier work this paper cites.
X.-S. Hua, L. Lu, and H.-J. Zhang, “Automatic music video generation based on temporal pattern analysis,” in Proceedings of the 12th annual ACM international conference on Multimedia . ACM, 2004, pp. 472–475
2004
Earlier work this paper cites.
J. Shawe-Taylor and N. Cristianini, Kernel methods for pattern analysis . Cambridge university press, 2004
2004
Earlier work this paper cites.
D. R. Hardoon, S. Szedmak, and J. Shawe-Taylor, “Canonical correlation analysis: An overview with application to learning methods,” Neural computation , vol. 16, no. 12, pp. 2639–2664, 2004
2004
Earlier work this paper cites.
D. A. Shamma, B. Pardo, and K. J. Hammond, “Musicstory: a personalized music video creator,” in Proceedings of the 13th annual ACM international conference on Multimedia . ACM, 2005, pp. 563–566
2005
Earlier work this paper cites.
O. Gillet, S. Essid, and G. Richard, “On the correlation of automatic audio and visual segmentations of music videos,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 17, no. 3, pp. 347–355, 2007
2007
Earlier work this paper cites.
R. Mayer, R. Neumayer, and A. Rauber, “Combination of audio and lyrics features for genre classification in digital audio collections,” in Proceedings of the 16th ACM international conference on Multimedia . ACM, 2008, pp. 159–168
2008
Earlier work this paper cites.
J. H. Kim, B. Tomasik, and D. Turnbull, “Using artist similarity to propagate semantic information.” in ISMIR , vol. 9, 2009, pp. 375–380
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
K. Q. Weinberger and L. K. Saul, “Distance metric learning for large margin nearest neighbor classification,” Journal of Machine Learning Research , vol. 10, no. Feb, pp. 207–244, 2009
2009
Earlier work this paper cites.
R. Mayer, “Analysing the similarity of album art with self-organising maps,” in International Workshop on Self-Organizing Maps . Springer, 2011, pp. 357–366
2011
Earlier work this paper cites.
J. Chao, H. Wang, W. Zhou, W. Zhang, and Y. Yu, “Tunesensor: A semantic-driven music recommendation service for digital photo albums,” in Proceedings of the 10th International Semantic Web Conference. ISWC2011 (October 2011) , 2011
2011
Earlier work this paper cites.
J. Libeks and D. Turnbull, “You can judge an artist by an album cover: Using images for music annotation,” IEEE MultiMedia , vol. 18, no. 4, pp. 30–37, 2011
2011
Earlier work this paper cites.
J. Weston, S. Bengio, and N. Usunier, “Wsabie: Scaling up to large vocabulary image annotation,” in IJCAI , vol. 11, 2011, pp. 2764–2770
2011
Earlier work this paper cites.
Z. Fu, G. Lu, K. M. Ting, and D. Zhang, “A survey of audio-based music classification and annotation,” IEEE transactions on multimedia , vol. 13, no. 2, pp. 303–319, 2011
2011
Cited alongside, same era.
M. Müller and S. Ewert, “Chroma toolbox: Matlab implementations for extracting variants of chroma-based audio features,” in Proceedings of the 12th International Conference on Music Information Retrieval (ISMIR), 2011. hal-00727791, version 2-22 Oct 2012 . Citeseer, 2011
2011
Cited alongside, same era.
Y. Yu, Z. Shen, and R. Zimmermann, “Automatic music soundtrack generation for outdoor videos from contextual sensor information,” in Proceedings of the 20th ACM international conference on Multimedia . ACM, 2012, pp. 1377–1378
2012
Cited alongside, same era.
C. Liem, A. Bazzica, and A. Hanjalic, “Musesync: standing on the shoulders of hollywood,” in Proceedings of the 20th ACM international conference on Multimedia . ACM, 2012, pp. 1383–1384
2012
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 1–9
2015
Later among the works it cites.
A. Schindler and A. Rauber, “An audio-visual approach to music genre classification through affective color features,” in European Conference on Information Retrieval . Springer, 2015, pp. 61–67
2015
Later among the works it cites.
S. Sasaki, T. Hirai, H. Ohya, and S. Morishima, “Affective music recommendation system based on the mood of input video,” in International Conference on Multimedia Modeling . Springer, 2015, pp. 299–302
2015
Later among the works it cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 3128–3137
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Bellard, M. Niedermayer et al. , “Ffmpeg,” Availabel from: http://ffmpeg. org , 2012
2012
Cited alongside, same era.
A. Van den Oord, S. Dieleman, and B. Schrauwen, “Deep content-based music recommendation,” in Advances in neural information processing systems , 2013, pp. 2643–2651
2013
Cited alongside, same era.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov et al. , “Devise: A deep visual-semantic embedding model,” in Advances in neural information processing systems , 2013, pp. 2121–2129
2013
Cited alongside, same era.
M. Khadkevich and M. Omologo, “Reassigned spectrum-based feature extraction for gmm-based automatic chord recognition,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2013, no. 1, p. 15, 2013
2013
Cited alongside, same era.
J. Wan, D. Wang, S. C. H. Hoi, P. Wu, J. Zhu, Y. Zhang, and J. Li, “Deep learning for content-based image retrieval: A comprehensive study,” in Proceedings of the 22nd ACM international conference on Multimedia . ACM, 2014, pp. 157–166
2014
Cited alongside, same era.
E. Acar, F. Hopfgartner, and S. Albayrak, “Understanding affective content of music videos through learned representations,” in International Conference on Multimedia Modeling . Springer, 2014, pp. 303–314
2014
Cited alongside, same era.
R. R. Shah, Y. Yu, and R. Zimmermann, “Advisor: Personalized video soundtrack recommendation by late fusion with heuristic rankings,” in Proceedings of the 22nd ACM international conference on Multimedia . ACM, 2014, pp. 607–616
2014
Cited alongside, same era.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision . Springer, 2014, pp. 740–755
2014
Cited alongside, same era.
S. Kum, C. Oh, and J. Nam, “Melody extraction on vocal segments using multi-column deep neural networks,” in The International Society for Music Information Retrieval (ISMIR), 2016 . ISMIR, 2016
2016
Later among the works it cites.
X. Wu, Y. Qiao, X. Wang, and X. Tang, “Bridging music and image via cross-modal ranking analysis,” IEEE Transactions on Multimedia , vol. 18, no. 7, pp. 1305–1318, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
L. Wang, Y. Li, and S. Lazebnik, “Learning deep structure-preserving image-text embeddings,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 5005–5013
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,” Journal of Machine Learning Research , vol. 17, no. 59, pp. 1–35, 2016
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Later among the works it cites.
2016
Later among the works it cites.
2017
Closest in time.
2017
Closest in time.