Fetching the paper…
Reading the bibliography…
The current landscape of research leveraging large language models (LLMs) is experiencing a surge.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Fast Zero-shot Image Tagging
Yang Zhang, Boqing Gong, and Mubarak Shah · 2016
Earlier work this paper cites.
Audio Set: An Ontology and Human-labeled Dataset for Audio Events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
CNN Architectures for Large-scale Audio Classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Boosting Image Captioning with Attributes
Ting Yao, Yingwei Pan, Yehao Li, Zhaofan Qiu, and Tao Mei · 2017
Earlier work this paper cites.
Sigmoid-weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2018
Earlier work this paper cites.
TVQA: Localized, Compositional Video Question Answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg · 2018
Earlier work this paper cites.
A Deep Learning based Approach for Precise Video Tagging
Sadia Ilyas and Hafeez Ur Rehman · 2019
Earlier work this paper cites.
Video Summarization with Attention-based Encoder-Decoder Networks
Zhong Ji, Kailin Xiong, Yanwei Pang, and Xuelong Li · 2019
Earlier work this paper cites.
Fréchet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2019
Earlier work this paper cites.
Zero-shot Event Detection via Event-adaptive Concept Relevance Mining
Zhihui Li, Lina Yao, Xiaojun Chang, Kun Zhan, Jiande Sun, and Huaxiang Zhang · 2019
Earlier work this paper cites.
Jukebox: A Generative Model for Music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Earlier work this paper cites.
Temporal Reasoning via Audio Question Answering
Haytham M Fayek and Justin Johnson · 2020
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Multi-modal Dense Video Captioning
Vladimir Iashin and Esa Rahtu · 2020
Earlier work this paper cites.
ViViT: A Video Vision Transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucic, and Cordelia Schmid · 2021
Earlier work this paper cites.
Video Background Music Generation with Controllable Music Transformer
Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan · 2021
Earlier work this paper cites.
Towards Duration Robust Weakly Supervised Sound Event Detection
Heinrich Dinkel, Mengyue Wu, and Kai Yu · 2021
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation
Yuan Gong, Yu-An Chung, and James Glass · 2021
Cited alongside, same era.
MusCaps: Generating Captions for Music Audio
Ilaria Manco, Emmanouil Benetos, Elio Quinton, and György Fazekas · 2021
Cited alongside, same era.
Audio Captioning Transformer
Xinhao Mei, Xubo Liu, Qiushi Huang, Mark D. Plumbley, and Wenwu Wang · 2021
Cited alongside, same era.
Audio Summarization for Podcasts
Aneesh Vartakavi, Amanmeet Garg, and Zafar Rafii · 2021
Cited alongside, same era.
Just Ask: Learning to Answer Questions from Millions of Narrated Videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2021
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, Hongsheng Li, and Yu Qiao · 2023
Closest in time.
LLark: A Multimodal Foundation Model for Music
Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner · 2023
Closest in time.
ImageBind: One Embedding Space To Bind Them All
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Closest in time.
Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Flamingo: A Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning
Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed Elhoseiny · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Cited alongside, same era.
Multi-Modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training
Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim, and Edward Choi · 2022
Cited alongside, same era.
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Cited alongside, same era.
Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, et al · 2023
Closest in time.
InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models
Bing Han, Junyu Dai, Xuchen Song, Weituo Hao, Xinyan He, Dong Guo, Jitong Chen, Yuxuan Wang, and Yanmin Qian · 2023
Closest in time.
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang · 2023
Closest in time.
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Chenyang Lyu, Minghao Wu, Longyue Wang, Xinting Huang, Bingshuai Liu, Zefeng Du, Shuming Shi, and Zhaopeng Tu · 2023
Closest in time.
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Salman Khan Muhammad Maaz, Hanoona Rasheed and Fahad Khan · 2023
Closest in time.
ChatGPT (Mar 14 version) [Large language model], 2023
OpenAI · 2023
Closest in time.
Mo \ \backslash ˆ usai: Text-to-Music Generation with Long-Context Latent Diffusion
Flavio Schneider, Zhijing Jin, and Bernhard Schölkopf · 2023
Closest in time.
PandaGPT: One Model To Instruction-Follow Them All
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai · 2023
Closest in time.
3D-GPT: Procedural 3D Modeling with Large Language Models
Chunyi Sun, Junlin Han, Weijian Deng, Xinlong Wang, Zishan Qin, and Stephen Gould · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs, 2023
MosaicML NLP Team · 2023
Closest in time.
Llama 2: Open Foundation and Fine-tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
TEAL: Tokenize and Embed ALL for Multi-modal Large Language Models
Zhen Yang, Yingxue Zhang, Fandong Meng, and Jie Zhou · 2023
Closest in time.
A Survey on Multimodal Large Language Models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2023
Closest in time.
Learning Video Representations from Large Language Models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Closest in time.
Video Background Music Generation: Dataset, Method and Evaluation
Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang, and Si Liu · 2023
Closest in time.