Fetching the paper…
Reading the bibliography…
In this paper we present a text-conditioned video resampler (TCR) module that uses a pre-trained and frozen visual encoder and large language model (LLM) to process long video sequences for a task.
Dynamic processing allocation in video
D. Chen, M. Bilgic, L. Getoor, and D. Jacobs · 2011
Earlier work this paper cites.
Action recognition and detection by combining motion and appearance features
L. Wang, Y. Qiao, X. Tang, et al · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
The kinetics human action video dataset, 2017
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Earlier work this paper cites.
Scsampler: Sampling salient clips from video for efficient action recognition
B. Korbar, D. Tran, and L. Torresani · 2019
Earlier work this paper cites.
Liteeval: A coarse-to-fine framework for resource efficient video recognition
Z. Wu, C. Xiong, Y.-G. Jiang, and L. S. Davis · 2019
Earlier work this paper cites.
Hacs: Human action clips and segments dataset for recognition and temporal localization
H. Zhao, A. Torralba, L. Torresani, and Z. Yan · 2019
Earlier work this paper cites.
Counting out time: Class agnostic video repetition counting in the wild
D. Dwibedi, Y. Aytar, J. Tompson, P. Sermanet, and A. Zisserman · 2020
Earlier work this paper cites.
Listen to look: Action recognition by previewing audio
R. Gao, T.-H. Oh, K. Grauman, and L. Torresani · 2020
Earlier work this paper cites.
Automatically discovering and learning new visual categories with ranking statistics
K. Han, S.-A. Rebuffi, S. Ehrhardt, A. Vedaldi, and A. Zisserman · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
G. Bertasius, H. Wang, and L. Torresani · 2021
Earlier work this paper cites.
Pix2seq: A language modeling framework for object detection
T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. Hinton · 2021
Earlier work this paper cites.
Smart frame selection for action recognition
S. N. Gowda, M. Rohrbach, and L. Sevilla-Lara · 2021
Earlier work this paper cites.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Earlier work this paper cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Cited alongside, same era.
Adamml: Adaptive multi-modal learning for efficient video recognition
R. Panda, C.-F. R. Chen, Q. Fan, X. Sun, K. Saenko, A. Oliva, and R. Feris · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Only time can tell: Discovering temporal data for temporal modeling
L. Sevilla-Lara, S. Zha, Z. Yan, V. Goswami, M. Feiszli, and L. Torresani · 2021
Cited alongside, same era.
Adaptive focus for efficient video recognition
Y. Wang, Z. Chen, H. Jiang, S. Song, Y. Han, and G. Huang · 2021
Cited alongside, same era.
Tubedetr: Spatio-temporal video grounding with transformers
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid · 2022
Later among the works it cites.
Actionformer: Localizing moments of actions with transformers
C.-L. Zhang, J. Wu, and Y. Li · 2022
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2023
J. Heek, A. Levskaya, A. Oliver, M. Ritter, B. Rondepierre, A. Steiner, and M. van Zee · 2023
Closest in time.
Palm: Predicting actions through language models @ ego4d long-term action anticipation challenge 2023, 2023
D. Huang, O. Hilliges, L. V. Gool, and X. Wang · 2023
Closest in time.
Single-stage visual query localization in egocentric videos
H. Jiang, S. K. Ramakrishnan, and K. Grauman · 2023
Closest in time.
MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Next-qa: Next phase of question-answering to explaining temporal actions
J. Xiao, X. Shang, A. Yao, and T.-S. Chua · 2021
Cited alongside, same era.
MERLOT: Multimodal Neural Script Knowledge Models
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi · 2021
Cited alongside, same era.
Mgsampler: An explainable sampling strategy for video action recognition
Y. Zhi, Z. Tong, L. Wang, and G. Wu · 2021
Cited alongside, same era.
Flamingo: a Visual Language Model for Few-Shot Learning, 2022
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Cited alongside, same era.
Revisiting the" video" in video-language understanding
S. Buch, C. Eyzaguirre, A. Gaidon, J. Wu, L. Fei-Fei, and J. C. Niebles · 2022
Cited alongside, same era.
Internvideo-ego4d: A pack of champion solutions to ego4d challenges, 2022
G. Chen, S. Xing, Z. Chen, Y. Wang, K. Li, Y. Li, Y. Liu, J. Wang, Y.-D. Zheng, B. Huang, Z. Zhao, J. Pan, Y. Huang, Z. Wang, J. Yu, Y. He, H. Zhang, T. Lu, Y. Wang, L. Wang, and Y. Qiao · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Cited alongside, same era.
W. Kuo, A. J. Piergiovanni, D. Kim, X. Luo, B. Caine, W. Li, A. Ogale, L. Zhou, A. Dai, Z. Chen, C. Cui, and A. Angelova · 2023
Closest in time.
Ring attention with blockwise transformers for near-infinite context, 2023
H. Liu, M. Zaharia, and P. Abbeel · 2023
Closest in time.
Video-chatgpt: Towards detailed video understanding via large vision and language models, 2023
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan · 2023
Closest in time.
Egoschema: A diagnostic benchmark for very long-form video language understanding
K. Mangalam, R. Akshulakov, and J. Malik · 2023
Closest in time.
Learning to ground instructional articles in videos through narrations
E. Mavroudi, T. Afouras, and L. Torresani · 2023
Closest in time.
Spotem: Efficient video search for episodic memory
S. K. Ramakrishnan, Z. Al-Halah, and K. Grauman · 2023
Closest in time.
Action sensitivity learning for the ego4d episodic memory challenge 2023
J. Shao, X. Wang, R. Quan, and Y. Yang · 2023
Closest in time.
Moviechat: From dense token to sparse memory for long video understanding, 2023
E. Song, W. Chai, G. Wang, Y. Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y. Zhang, Y. Lu, J.-N. Hwang, and G. Wang · 2023
Closest in time.
Eva-clip: Improved training techniques for clip at scale
Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao · 2023
Closest in time.
Multiscale video pretraining for long-term activity forecasting
R. Tan, M. De Lange, M. Iuzzolino, B. A. Plummer, K. Saenko, K. Ridgeway, and L. Torresani · 2023
Closest in time.
Videomae v2: Scaling video masked autoencoders with dual masking
L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao · 2023
Closest in time.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid · 2023
Closest in time.
Hitea: Hierarchical temporal-aware video-language pre-training
Q. Ye, G. Xu, M. Yan, H. Xu, Q. Qian, J. Zhang, and F. Huang · 2023
Closest in time.
Self-chained image-language model for video localization and question answering
S. Yu, J. Cho, P. Yadav, and M. Bansal · 2023
Closest in time.