Fetching the paper…
Reading the bibliography…
Understanding long-form video content presents significant challenges due to its temporal complexity and the substantial computational resources required.
Temporary suppression of visual processing in an rsvp task: An attentional blink?
J. E. Raymond, K. L. Shapiro, and K. M. Arnell · 1992
Earlier work this paper cites.
Developmental trajectories of regulating attentional selection over time
S. Heim and A. Keil · 2012
Earlier work this paper cites.
Motivated attention: Affect, activation, and action
P. J. Lang, M. M. Bradley, and B. N. Cuthbert · 2013
Earlier work this paper cites.
Too much information, too little time: How the brain separates important from unimportant things in our fast-paced media world
S. Heim and A. Keil · 2017
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
J. Xiao, X. Shang, A. Yao, and T.-S. Chua · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
C.-Y. Wu, Y. Li, K. Mangalam, H. Fan, B. Xiong, J. Malik, and C. Feichtenhofer · 2022
Earlier work this paper cites.
Zero-shot video question answering via frozen bidirectional language models
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
Bytetrack: Multi-object tracking by associating every detection box
Y. Zhang, P. Sun, Y. Jiang, D. Yu, F. Weng, Z. Yuan, P. Luo, W. Liu, and X. Wang · 2022
Earlier work this paper cites.
Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn
D. Gao, L. Ji, L. Zhou, K. Q. Lin, J. Chen, Z. Fan, and M. Z. Shou · 2023
Earlier work this paper cites.
Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering
D. Gao, L. Zhou, L. Ji, L. Zhu, Y. Yang, and M. Z. Shou · 2023
Earlier work this paper cites.
Sas video-qa: Self-adaptive sampling for efficient video question-answering
W. Han, H. Chen, M.-Y. Kan, and S. Poria · 2023
Earlier work this paper cites.
Can llms critique and iterate on their own outputs? evjang. com
E. Jang · 2023
Earlier work this paper cites.
Videochat: Chat-centric video understanding
K. Li, Y. He, Y. Wang, Y. Li, W. Wang, P. Luo, Y. Wang, L. Wang, and Y. Qiao · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
B. Lin, B. Zhu, Y. Ye, M. Ning, P. Jin, and L. Yuan · 2023
Cited alongside, same era.
Llava-plus: Learning to use tools for creating multimodal agents
S. Liu, H. Cheng, H. Liu, H. Zhang, F. Li, T. Ren, X. Zou, J. Yang, H. Su, J. Zhu, et al · 2023
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan · 2023
Cited alongside, same era.
Understanding the capabilities of large language models for automated planning
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Egoschema: A diagnostic benchmark for very long-form video language understanding
K. Mangalam, R. Akshulakov, and J. Malik · 2024
Closest in time.
Paddleocr, 2024
PaddleOCR · 2024
Closest in time.
A simple recipe for contrastively pre-training video-first encoders beyond 16 frames
P. Papalampidi, S. Koppula, S. Pathak, J. Chiu, J. Heyward, V. Patraucean, J. Shen, A. Miech, A. Zisserman, and A. Nematzdeh · 2024
Closest in time.
Mirasol3b: A multimodal autoregressive model for time-aligned and contextual modalities
A. Piergiovanni, I. Noble, D. Kim, M. S. Ryoo, V. Gomes, and A. Angelova · 2024
Closest in time.
Question-instructed visual descriptions for zero-shot video question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Pallagani, B. Muppasani, K. Murugesan, F. Rossi, B. Srivastava, L. Horesh, F. Fabiano, and A. Loreggia · 2023
Cited alongside, same era.
Retrieving-to-answer: Zero-shot video question answering with frozen large language models
J. Pan, Z. Lin, Y. Ge, X. Zhu, R. Zhang, Y. Wang, Y. Qiao, and H. Li · 2023
Cited alongside, same era.
L. Pan, M. Saxon, W. Xu, D. Nathani, X. Wang, and W. Y. Wang · 2023
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2023
Cited alongside, same era.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, Y. Wang, et al · 2023
Cited alongside, same era.
Lifelongmemory: Leveraging llms for answering queries in egocentric videos
Y. Wang, Y. Yang, and M. Ren · 2023
Cited alongside, same era.
A simple llm framework for long-range video question-answering
C. Zhang, T. Lu, M. M. Islam, Z. Wang, S. Yu, M. Bansal, and G. Bertasius · 2023
Cited alongside, same era.
Video-llama: An instruction-tuned audio-visual language model for video understanding
H. Zhang, X. Li, and L. Bing · 2023
Cited alongside, same era.
D. Romero and T. Solorio · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao · 2024
Closest in time.
Moviechat: From dense token to sparse memory for long video understanding
E. Song, W. Chai, G. Wang, Y. Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y. Zhang, et al · 2024
Closest in time.
Videoagent: Long-form video understanding with large language model as agent
X. Wang, Y. Zhang, O. Zohar, and S. Yeung-Levy · 2024
Closest in time.
Videotree: Adaptive tree-based video representation for llm reasoning on long videos
Z. Wang, S. Yu, E. Stengel-Eskin, J. Yoon, F. Cheng, G. Bertasius, and M. Bansal · 2024
Closest in time.
Doraemongpt: Toward understanding dynamic scenes with large language models
Z. Yang, G. Chen, X. Li, W. Wang, and Y. Yang · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Closest in time.
Self-chained image-language model for video localization and question answering
S. Yu, J. Cho, P. Yadav, and M. Bansal · 2024
Closest in time.
Detrs beat yolos on real-time object detection
Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and J. Chen · 2024
Closest in time.
Large language models as commonsense knowledge for large-scale task planning
Z. Zhao, W. S. Lee, and D. Hsu · 2024
Closest in time.