Fetching the paper…
Reading the bibliography…
Recent studies have shown promising results on utilizing large pre-trained image-language models for video question answering.
Torchvision the machine-vision package of torch
S. Marcel and Y. Rodriguez · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
D.-H. Lee et al · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
J. Lei, L. Yu, M. Bansal, and T. L. Berg · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Temporal localization of moments in video collections with natural language, 2019
V. Escorcia, M. Soldan, J. Sivic, B. Ghanem, and B. Russell · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition
W. Wu, D. He, X. Tan, S. Chen, and S. Wen · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Earlier work this paper cites.
What is more likely to happen next? video-and-language future event prediction
J. Lei, L. Yu, T. L. Berg, and M. Bansal · 2020
Earlier work this paper cites.
Tvr: A large-scale dataset for video-subtitle moment retrieval
J. Lei, L. Yu, T. L. Berg, and M. Bansal · 2020
Earlier work this paper cites.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Earlier work this paper cites.
Unifying vision-and-language tasks via text generation
J. Cho, J. Lei, H. Tan, and M. Bansal · 2021
Earlier work this paper cites.
Differentiable patch selection for image recognition
J.-B. Cordonnier, A. Mahendran, A. Dosovitskiy, D. Weissenborn, J. Uszkoreit, and T. Unterthiner · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2021
Earlier work this paper cites.
Clip2video: Mastering video-text retrieval via image clip
H. Fang, P. Xiong, L. Xu, and Y. Chen · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Cited alongside, same era.
Star: A benchmark for situated reasoning in real-world videos
B. Wu, S. Yu, Z. Chen, J. B. Tenenbaum, and C. Gan · 2021
Cited alongside, same era.
Next-qa: Next phase of question-answering to explaining temporal actions
J. Xiao, X. Shang, A. Yao, and T.-S. Chua · 2021
Cited alongside, same era.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer · 2021
Cited alongside, same era.
Zero-shot video question answering via frozen bidirectional language models
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid · 2022
Later among the works it cites.
Filip: fine-grained interactive language-image pre-training
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Later among the works it cites.
Merlot reserve: Neural script knowledge through vision and language and sound
R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid · 2021
Cited alongside, same era.
Florence: A new foundation model for computer vision
L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, et al · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Revisiting the" video" in video-language understanding
S. Buch, C. Eyzaguirre, A. Gaidon, J. Wu, L. Fei-Fei, and J. C. Niebles · 2022
Cited alongside, same era.
Locvtp: Video-text pre-training for temporal localization
M. Cao, T. Yang, J. Weng, C. Zhang, J. Wang, and Y. Zou · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Cited alongside, same era.
Video moment retrieval from text queries via single frame annotation
R. Cui, T. Qian, P. Peng, E. Daskalaki, J. Chen, X. Guo, H. Sun, and Y.-G. Jiang · 2022
Cited alongside, same era.
A large-scale study of spatiotemporal representation learning with a new benchmark on action recognition
A. Deng, T. Yang, and C. Chen · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Closest in time.
Eva: Exploring the limits of masked visual representation learning at scale
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao · 2023
Closest in time.
An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling
T.-J. Fu*, L. Li*, Z. Gan, K. Lin, W. Y. Wang, L. Wang, and Z. Liu · 2023
Closest in time.
Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering
D. Gao, L. Zhou, L. Ji, L. Zhu, Y. Yang, and M. Z. Shou · 2023
Closest in time.
Semi-parametric video-grounded text generation
S. Kim, J.-H. Kim, J. Lee, and M. Seo · 2023
Closest in time.
Revealing single frame bias for video-and-language learning
J. Lei, T. L. Berg, and M. Bansal · 2023
Closest in time.
Towards fast adaptation of pretrained contrastive models for multi-channel video-language retrieval
X. Lin, S. Tiwari, S. Huang, M. Li, M. Z. Shou, H. Ji, and S.-F. Chang · 2023
Closest in time.
Unified-io: A unified model for vision, language, and multi-modal tasks
J. Lu, C. Clark, R. Zellers, R. Mottaghi, and A. Kembhavi · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2023
Closest in time.
Verbs in action: Improving verb understanding in video-language models
L. Momeni, M. Caron, A. Nagrani, A. Zisserman, and C. Schmid · 2023
Closest in time.
Query-dependent video representation for moment retrieval and highlight detection
W. Moon, S. Hyun, S. Park, D. Park, and J.-P. Heo · 2023
Closest in time.
Eva-clip: Improved training techniques for clip at scale
Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
D. Surís, S. Menon, and C. Vondrick · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Contrastive video question answering via video graph transformer
J. Xiao, P. Zhou, A. Yao, Y. Li, R. Hong, S. Yan, and T. S. Chua · 2023
Closest in time.
mplug-2: A modularized multi-modal foundation model across text, image and video
H. Xu, Q. Ye, M. Yan, Y. Shi, J. Ye, Y. Xu, C. Li, B. Bi, Q. Qian, W. Wang, et al · 2023
Closest in time.
CLIP-vip: Adapting pre-trained image-text model to video-language alignment
H. Xue, Y. Sun, B. Liu, J. Fu, R. Song, H. Li, and J. Luo · 2023
Closest in time.
Hitea: Hierarchical temporal-aware video-language pre-training
Q. Ye, G. Xu, M. Yan, H. Xu, Q. Qian, J. Zhang, and F. Huang · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Closest in time.