Fetching the paper…
Reading the bibliography…
The increasing amount of online videos brings several opportunities for training self-supervised neural networks.
“The sound of an album cover: Probabilistic multimedia and information retrieval,”
Eric Brochu, Nando De Freitas, and Kejie Bao, · 2003
Earlier work this paper cites.
“Multimedia content processing through cross-modal association,”
Dongge Li, Nevenka Dimitrova, Mingkun Li, and Ishwar K Sethi, · 2003
Earlier work this paper cites.
“On the correlation of automatic audio and visual segmentations of music videos,”
Olivier Gillet, Slim Essid, and Gal Richard, · 2007
Earlier work this paper cites.
“Cross-modal correlation learning for clustering on image-audio dataset,”
Hong Zhang, Yueting Zhuang, and Fei Wu, · 2007
Earlier work this paper cites.
“Analysing the similarity of album art with self-organising maps,”
Rudolf Mayer, · 2011
Earlier work this paper cites.
“You can judge an artist by an album cover: Using images for music annotation,”
Janis Libeks and Douglas Turnbull, · 2011
Earlier work this paper cites.
“Tunesensor: A semantic-driven music recommendation service for digital photo albums,”
Jiansong Chao, Haofen Wang, Wenlei Zhou, Weinan Zhang, and Yong Yu, · 2011
Earlier work this paper cites.
“Multimodal deep learning,”
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng, · 2011
Cited alongside, same era.
“Devise: A deep visual-semantic embedding model,”
Andrea Frome, Greg Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov, · 2013
Cited alongside, same era.
“Understanding affective content of music videos through learned representations,”
Esra Acar, Frank Hopfgartner, and Sahin Albayrak, · 2014
Cited alongside, same era.
“Unifying visual-semantic embeddings with multimodal neural language models,”
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel, · 2014
Cited alongside, same era.
“An audio-visual approach to music genre classification through affective color features,”
Alexander Schindler and Andreas Rauber, · 2015
Cited alongside, same era.
“Youtube-8m: A large-scale video classification benchmark,”
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan, · 2016
Later among the works it cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,”
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al., · 2016
Later among the works it cites.
“Learning cross-modal embeddings for cooking recipes and food images,”
Amaia Salvador, Nicholas Hynes, Yusuf Aytar, Javier Marin, Ferda Ofli, Ingmar Weber, and Antonio Torralba, · 2017
Later among the works it cites.
“See, hear, and read: Deep aligned representations,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2017
Later among the works it cites.
“Deep learning for content-based, cross-modal retrieval of videos and music,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liwei Wang, Yin Li, and Svetlana Lazebnik, · 2015
Cited alongside, same era.
“Bridging music and image via cross-modal ranking analysis,”
Xixuan Wu, Yu Qiao, Xiaogang Wang, and Xiaoou Tang, · 2016
Cited alongside, same era.
Sungeun Hong, Woobin Im, and Hyun S. Yang, · 2017
Later among the works it cites.
“Cnn architectures for large-scale audio classification,”
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, · 2017
Later among the works it cites.