Fetching the paper…
Reading the bibliography…
We propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language.
Evaluation of algorithms using games: The case of music tagging
Edith Law, Kris West, Michael I Mandel, Mert Bay, and J Stephen Downie · 2009
Earlier work this paper cites.
The million song dataset
Thierry Bertin-Mahieux, Daniel PW Ellis, Brian Whitman, and Paul Lamere · 2011
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Moddrop: adaptive multi-modal gesture recognition
Natalia Neverova, Christian Wolf, Graham Taylor, and Florian Nebout · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Automatic tagging using deep convolutional neural networks
Keunwoo Choi, George Fazekas, and Mark Sandler · 2016
Earlier work this paper cites.
See, hear, and read: Deep aligned representations
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2017
Earlier work this paper cites.
Cbvmr: Content-based video-music retrieval using soft intra-modal structure constraint
Sungeun Hong, Woobin Im, and Hyun Seung Yang · 2017
Earlier work this paper cites.
Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms
Jongpil Lee, Jiyoung Park, Keunhyoung Kim, and Juhan Nam · 2017
Earlier work this paper cites.
Multi-label music genre classification from audio, text and images using deep features
Sergio Oramas, Oriol Nieto, Francesco Barbieri, and Xavier Serra · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Audio-visual embedding for cross-modal music video retrieval through supervised deep cca
Donghuo Zeng, Yi Yu, and Keizo Oyama · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra · 2019
Earlier work this paper cites.
Zero-shot learning for audio-based music classification and tagging
Jeong Choi, Jongpil Lee, Jiyoung Park, and Juhan Nam · 2019
Earlier work this paper cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdel rahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Query by video: Cross-modal music retrieval
Bochen Li and Aparna Kumar · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
musicnn: Pre-trained convolutional neural networks for music audio tagging
Jordi Pons and Xavier Serra · 2019
Cited alongside, same era.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Bigscience large open-science open-access multilingual language model, 2022
BigScience · 2022
Later among the works it cites.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Later among the works it cites.
Mulan: A joint embedding of music audio and natural language
Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, and Daniel P. W. Ellis · 2022
Later among the works it cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, F. Xia, Ted Xiao, Harris Chan, Jacky Liang, Peter R. Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter · 2022
Later among the works it cites.
Neural pipeline for zero-shot data-to-text generation
Zdeněk Kasner and Ondřej Dušek · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DAGA: Data augmentation with a generation approach for low-resource tagging tasks
Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, and Chunyan Miao · 2020
Cited alongside, same era.
Modality dropout for improved performance-driven talking faces
Ahmed Hussen Abdelaziz, Barry-John Theobald, Paul Dixon, Reinhard Knothe, Nicholas Apostoloff, and Sachin Kajareker · 2020
Cited alongside, same era.
Disentangled multidimensional metric learning for music similarity
Jongpil Lee, Nicholas J. Bryan, Justin Salamon, Zeyu Jin, and Juhan Nam · 2020
Cited alongside, same era.
Avlnet: Learning audio-visual language representations from instructional videos
Andrew Rouditchenko, Angie Boggust, David F. Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogério Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, and James R. Glass · 2020
Cited alongside, same era.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Cited alongside, same era.
Evaluation of cnn-based automatic music tagging models
Minz Won, Andres Ferraro, Dmitry Bogdanov, and Xavier Serra · 2020
Cited alongside, same era.
Generative data augmentation for commonsense reasoning
Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey · 2020
Cited alongside, same era.
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Later among the works it cites.
Learning to recognize procedural activities with distant supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani · 2022
Later among the works it cites.
Contrastive audio-language learning for music
Ilaria Manco, Emmanouil Benetos, Elio Quinton, and György Fazekas · 2022
Later among the works it cites.
It’s time for artistic correspondence in music and video
Dídac Surís, Carl Vondrick, Bryan Russell, and Justin Salamon · 2022
Later among the works it cites.
PromDA: Prompt-based data augmentation for low-resource NLU tasks
Yufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, and Daxin Jiang · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello · 2022
Later among the works it cites.
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al · 2022
Later among the works it cites.
ProgPrompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg · 2023
Closest in time.