Fetching the paper…
Reading the bibliography…
Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain.
“Evaluation of algorithms using games: The case of music tagging”
Edith Law et al · 2009
Earlier work this paper cites.
“Inferring Semantic Facets of a Music Folksonomy with Wikipedia”
Mohamed Sordo et al · 2013
Earlier work this paper cites.
“mirdata: Software for Reproducible Usage of Datasets”
Rachel. Bittner et al · 2019
Earlier work this paper cites.
“The MTG-Jamendo Dataset for Automatic Music Tagging”
Dmitry Bogdanov et al · 2019
Earlier work this paper cites.
“STARC: Structured Annotations for Reading Comprehension”
Yevgeni Berzak, Jonathan Malmaud and Roger Levy · 2020
Earlier work this paper cites.
“On the opportunities and risks of foundation models”
Rishi Bommasani et al · 2021
Earlier work this paper cites.
“Datasheets for datasets”
Timnit Gebru et al · 2021
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”
Dan Hendrycks et al · 2021
Earlier work this paper cites.
“Flamingo: a Visual Language Model for Few-Shot Learning”
Jean-Baptiste Alayrac et al · 2022
Earlier work this paper cites.
“Music Representation Learning Based on Editorial Metadata from Discogs”
Pablo Alonso-Jiménez, Xavier Serra and Dmitry Bogdanov · 2022
Earlier work this paper cites.
“Large Language Models are Zero-Shot Reasoners”
Takeshi Kojima et al · 2022
Earlier work this paper cites.
“Chain of Thought Prompting Elicits Reasoning in Large Language Models”
Jason Wei et al · 2022
Earlier work this paper cites.
“Finetuned Language Models are Zero-Shot Learners”
Jason Wei et al · 2022
Earlier work this paper cites.
“Gemini: a family of highly capable multimodal models”
Gemini Team et al · 2023
Earlier work this paper cites.
OpenAI et al · 2023
Earlier work this paper cites.
“Qwen-vl: A frontier large vision-language model with versatile abilities”
Jinze Bai et al · 2023
Earlier work this paper cites.
“Pengi: An Audio Language Model for Audio Tasks”
Soham Deshmukh, Benjamin Elizalde, Rita Singh and Huaming Wang · 2023
Earlier work this paper cites.
“Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models”
Yunfei Chu et al · 2023
Cited alongside, same era.
“VisIT-Bench: A Dynamic Benchmark for Evaluating Instruction-Following Vision-and-Language Models”
Yonatan Bitton et al · 2023
Cited alongside, same era.
“Musiclm: Generating music from text”
Andrea Agostinelli et al · 2023
Cited alongside, same era.
“The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation”
Ilaria Manco et al · 2023
Cited alongside, same era.
“Language-Guided Music Recommendation for Video via Prompt Analogies”
Daniel McKee, Justin Salamon, Josef Sivic and Bryan Russell · 2023
Cited alongside, same era.
“Bootstrapping Vision-Language Learning with Decoupled Language Pre-training”
Yiren Jian, Chongyang Gao and Soroush Vosoughi · 2023
Later among the works it cites.
“Self-Instruct: Aligning Language Models with Self-Generated Instructions”
Yizhong Wang et al · 2023
Later among the works it cites.
“Grounding Multimodal Large Language Models to the World”
Zhiliang Peng et al · 2024
Closest in time.
“LLark: A Multimodal Instruction-Following Language Model for Music”
Josh Gardner, Simon Durand, Daniel Stoller and Rachel Bittner · 2024
Closest in time.
“Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning”
Shansong Liu, Atin Hussain, Chenshuo Sun and Ying Shan · 2024
Closest in time.
“MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“LP-MusicCaps: LLM-Based Pseudo Music Captioning”
SeungHeon Doh, Keunwoo Choi, Jongpil Lee and Juhan Nam · 2023
Cited alongside, same era.
“MARBLE: Music Audio Representation Benchmark for Universal Evaluation”
Ruibin Yuan et al · 2023
Cited alongside, same era.
“mir_ref: A Representation Evaluation Framework for Music Information Retrieval Tasks”
Christos Plachouras, Pablo Alonso-Jiménez and Dmitry Bogdanov · 2023
Cited alongside, same era.
“Agieval: A human-centric benchmark for evaluating foundation models”
Wanjun Zhong et al · 2023
Cited alongside, same era.
“Holistic Evaluation of Language Models”
Percy Liang et al · 2023
Cited alongside, same era.
“Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models”
Aarohi Srivastava et al · 2023
Cited alongside, same era.
“Leveraging Large Language Models for Multiple Choice Question Answering”
Joshua Robinson and David Wingate · 2023
Cited alongside, same era.
Zihao Deng et al · 2024
Closest in time.
“SALMONN: Towards Generic Hearing Abilities for Large Language Models”
Changli Tang et al · 2024
Closest in time.
Zihao Wang et al · 2024
Closest in time.
“Chatmusician: Understanding and generating music intrinsically with llm”
Ruibin Yuan et al · 2024
Closest in time.
Jiajia Li et al · 2024
Closest in time.
“AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension”
Qian Yang et al · 2024
Closest in time.
“MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training”
Yizhi LI et al · 2024
Closest in time.
“LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized Attention”
Renrui Zhang et al · 2024
Closest in time.
“Towards multimodal in-context learning for vision & language models”
Sivan Doveh et al · 2024
Closest in time.
“MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning”
Haozhe Zhao et al · 2024
Closest in time.
“HallusionBench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models”
Tianrui Guan et al · 2024
Closest in time.
“Large Language Models Are Not Robust Multiple Choice Selectors”
Chujie Zheng et al · 2024
Closest in time.