Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains.
Collecting Highly Parallel Data for Paraphrase Evaluation
Chen, D.; and Dolan, W. 2011 · 2011
Earlier work this paper cites.
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017 · 2017
Earlier work this paper cites.
Audio Set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P. W.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
Dense-Captioning Events in Videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017 · 2017
Earlier work this paper cites.
Audio Visual Scene-Aware Dialog
Alamri, H.; Cartillier, V.; Das, A.; Wang, J.; Cherian, A.; Essa, I.; Batra, D.; Marks, T. K.; Hori, C.; Anderson, P.; Lee, S.; and Parikh, D. 2019 · 2019
Earlier work this paper cites.
AudioCaps: Generating Captions for Audios in The Wild
Kim, C. D.; Kim, B.; Lee, H.; and Kim, G. 2019 · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Yu, Z.; Xu, D.; Yu, J.; Yu, T.; Zhao, Z.; Zhuang, Y.; and Tao, D. 2019 · 2019
Earlier work this paper cites.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Tian, Y.; Li, D.; and Xu, C. 2020 · 2020
Earlier work this paper cites.
Detecting moments and highlights in videos via natural language queries
Lei, J.; Berg, T. L.; and Bansal, M. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Generic event boundary detection: A benchmark for event segmentation
Shou, M. Z.; Lei, S. W.; Wang, W.; Ghadiyaram, D.; and Feiszli, M. 2021 · 2021
Earlier work this paper cites.
Video self-stitching graph network for temporal action localization
Zhao, C.; et al. 2021 · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
Learning to answer questions in dynamic audio-visual scenarios
Li, G.; Wei, Y.; Tian, Y.; Xu, C.; Wen, J.-R.; and Hu, D. 2022 · 2022
Cited alongside, same era.
End-to-end temporal action detection with transformer
Liu, X.; Wang, Q.; Hu, Y.; Tang, X.; Zhang, S.; Bai, S.; and Bai, X. 2022 · 2022
Cited alongside, same era.
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
Tang, Y.; Xu, S.; Wang, T.; Lin, Q.; Lu, Q.; and Zheng, F. 2022 · 2022
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Kosmos-2: Grounding Multimodal Large Language Models to the World
Peng, Z.; Wang, W.; Dong, L.; Hao, Y.; Huang, S.; Ma, S.; and Wei, F. 2023 · 2023
Later among the works it cites.
Audio-Visual LLM for Video Understanding
Shu, F.; Zhang, L.; Jiang, H.; and Xie, C. 2023 · 2023
Later among the works it cites.
Pandagpt: One model to instruction-follow them all
Su, Y.; Lan, T.; Li, H.; Xu, J.; Wang, Y.; and Cai, D. 2023 · 2023
Later among the works it cites.
A Large-scale Dataset for Audio-Language Representation Learning
Sun, L.; Xu, X.; Wu, M.; and Xie, W. 2023 · 2023
Later among the works it cites.
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MISAR: A Multimodal Instructional System with Augmented Reality
Bi, J.; Nguyen, N. M.; Vosoughi, A.; and Xu, C. 2023 · 2023
Cited alongside, same era.
Clap learning audio concepts from natural language supervision
Elizalde, B.; Deshmukh, S.; Al Ismail, M.; and Wang, H. 2023 · 2023
Cited alongside, same era.
Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline
Geng, T.; Wang, T.; Duan, J.; Cong, R.; and Zheng, F. 2023 · 2023
Cited alongside, same era.
Towards Long Form Audio-visual Video Understanding
Hou, W.; li, G.; Tian, Y.; and Hu, D. 2023 · 2023
Cited alongside, same era.
Vtimellm: Empower llm to grasp video moments
Huang, B.; Wang, X.; Chen, H.; Song, Z.; and Zhu, W. 2023 · 2023
Cited alongside, same era.
Single-stage visual query localization in egocentric videos
Jiang, H.; Ramakrishnan, S. K.; and Grauman, K. 2023 · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Cited alongside, same era.
Tang, Y.; Zhang, J.; Wang, X.; Wang, T.; and Zheng, F. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Launchpadgpt: Language model as music visualization designer on launchpad
Xu, S.; Tang, Y.; and Zheng, F. 2023 · 2023
Later among the works it cites.
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
Xuan, S.; Guo, Q.; Yang, M.; and Zhang, S. 2023 · 2023
Later among the works it cites.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Yang, A.; Nagrani, A.; Seo, P. H.; Miech, A.; Pont-Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C. 2023 · 2023
Later among the works it cites.
DNAGPT: A Generalized Pretrained Tool for Multiple DNA Sequence Analysis Tasks
Zhang, D.; Zhang, W.; He, B.; Zhang, J.; Qin, C.; and Yao, J. 2023a · 2023
Later among the works it cites.
Contrastive Positive Sample Propagation Along the Audio-Visual Event Line
Zhou, J.; et al. 2023 · 2023
Later among the works it cites.
UniAV: Unified Audio-Visual Perception for Multi-Task Video Localization
Geng, T.; Wang, T.; Zhang, Y.; Duan, J.; Guan, W.; and Zheng, F. 2024 · 2024
Closest in time.
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
Zhang, D.; Yang, J.; Lyu, H.; Jin, Z.; Yao, Y.; Chen, M.; and Luo, J. 2024 · 2024
Closest in time.