Fetching the paper…
Reading the bibliography…
Music recommendation for videos attracts growing interest in multi-modal research.
The million song dataset
Thierry Bertin-Mahieux, Daniel P.W. Ellis, Brian Whitman, and Paul Lamere · 2011
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Automatic tagging using deep convolutional neural networks
Keunwoo Choi, George Fazekas, and Mark Sandler · 2016
Earlier work this paper cites.
Brains on beats, 2016
Umut Güçlü, Jordy Thielen, Michael Hanke, and Marcel A. J. van Gerven · 2016
Earlier work this paper cites.
Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Audio-visual embedding for cross-modal music video retrieval through supervised deep cca
Donghuo Zeng, Yi Yu, and Keizo Oyama · 2018
Earlier work this paper cites.
Zero-shot learning for audio-based music classification and tagging
Jeong Choi, Jongpil Lee, Jiyoung Park, and Juhan Nam · 2019
Earlier work this paper cites.
musicnn: Pre-trained convolutional neural networks for music audio tagging
Jordi Pons and Xavier Serra · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Disentangled multidimensional metric learning for music similarity
Jongpil Lee, Nicholas J Bryan, Justin Salamon, Zeyu Jin, and Juhan Nam · 2020
Earlier work this paper cites.
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Earlier work this paper cites.
Evaluation of cnn-based automatic music tagging models
Minz Won, Andres Ferraro, Dmitry Bogdanov, and Xavier Serra · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert, 2020
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
Advances and challenges in conversational recommender systems: A survey
Chongming Gao, Wenqiang Lei, Xiangnan He, Maarten de Rijke, and Tat-Seng Chua · 2021
Cited alongside, same era.
Ast: Audio spectrogram transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
It’s time for artistic correspondence in music and video
Dídac Surís, Carl Vondrick, Bryan Russell, and Justin Salamon · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello · 2022
Later among the works it cites.
S3t: Self-supervised pre-training with swin transformer for music classification
Hang Zhao, Chen Zhang, Bilei Zhu, Zejun Ma, and Kejun Zhang · 2022
Later among the works it cites.
A unified multi-task learning framework for multi-goal conversational recommender systems
Yang Deng, Wenxuan Zhang, Weiwen Xu, Wenqiang Lei, Tat-Seng Chua, and Wai Lam · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey on conversational recommender systems
Dietmar Jannach, Ahtsham Manzoor, Wanling Cai, and Li Chen · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Cross-modal music-video recommendation: A study of design choices
Laure Prétet, Gael Richard, and Geoffroy Peeters · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Semi-supervised music tagging transformer
Minz Won, Keunwoo Choi, and Xavier Serra · 2021
Cited alongside, same era.
Improving conversational recommendation systems’ quality with context-aware item meta information
Bowen Yang, Cong Han, Yu Li, Lei Zuo, and Zhou Yu · 2021
Cited alongside, same era.
Cross-modal variational auto-encoder for content-based micro-video background music recommendation
Jing Yi, Yaochen Zhu, Jiayi Xie, and Zhenzhong Chen · 2021
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms, 2023
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Closest in time.
Toward universal text-to-music retrieval
SeungHeon Doh, Minz Won, Keunwoo Choi, and Juhan Nam · 2023
Closest in time.
Leveraging large language models in conversational recommender systems
Luke Friedman, Sameer Ahuja, David Allen, Terry Tan, Hakim Sidahmed, Changbo Long, Jun Xie, Gabriel Schubiner, Ajay Patel, Harsh Lara, et al · 2023
Closest in time.
Audio–text retrieval based on contrastive learning and collaborative attention mechanism
Tao Hu, Xuyu Xiang, Jiaohua Qin, and Yun Tan · 2023
Closest in time.
Vision transformers are parameter-efficient audio-visual learners
Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, and Gedas Bertasius · 2023
Closest in time.
Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering
Xiulong Liu, Zhikang Dong, and Peng Zhang · 2023
Closest in time.
Language-guided music recommendation for video via prompt analogies
Daniel McKee, Justin Salamon, Josef Sivic, and Bryan Russell · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Rethinking the evaluation for conversational recommendation in the era of large language models
Xiaolei Wang, Xinyu Tang, Wayne Xin Zhao, Jingyuan Wang, and Ji-Rong Wen · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Closest in time.