Fetching the paper…
Reading the bibliography…
The increasing abundance of video data enables users to search for events of interest, e.g., emergency incidents.
Language Models are Few-shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
Probabilistic Automata
Rabin, M. O. 1963 · 1963
Earlier work this paper cites.
A fully automated content-based video search engine supporting spatiotemporal queries
Chang, S.; Chen, W.; Meng, H. J.; Sundaram, H.; and Zhong, D. 1998 · 1998
Earlier work this paper cites.
Detection and Classification of Shot Transitions
Porter, S. V.; Mirmehdi, M.; and Thomas, B. T. 2001 · 2001
Earlier work this paper cites.
Recent Advances and Challenges of Semantic Image/Video Search
Chang, S.; Ma, W.; and Smeulders, A. W. M. 2007 · 2007
Earlier work this paper cites.
Actions as Space-Time Shapes
Gorelick, L.; Blank, M.; Shechtman, E.; Irani, M.; and Basri, R. 2007 · 2007
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
Kuehne, H.; Jhuang, H.; Garrote, E.; Poggio, T. A.; and Serre, T. 2011 · 2011
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M. S.; Berg, A. C.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Learning Multi-domain Convolutional Neural Networks for Visual Tracking
Nam, H.; and Han, B. 2016 · 2016
Earlier work this paper cites.
You Only Look Once: Unified, Real-Time Object Detection
Redmon, J.; Divvala, S. K.; Girshick, R. B.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
MovieQA: Understanding Stories in Movies through Question-Answering
Tapaswi, M.; Zhu, Y.; Stiefelhagen, R.; Torralba, A.; Urtasun, R.; and Fidler, S. 2016 · 2016
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2017 · 2017
Earlier work this paper cites.
Symbolic LTLf Synthesis
Zhu, S.; Tabajara, L. M.; Li, J.; Pu, G.; and Vardi, M. Y. 2017 · 2017
Earlier work this paper cites.
Non-Local Neural Networks
Wang, X.; Girshick, R. B.; Gupta, A.; and He, K. 2018 · 2018
Earlier work this paper cites.
Learning-based Probabilistic Modeling and Verifying Driver Behavior using MDP
Bai, X.; Xu, C.; Ao, Y.; Chen, B.; and Du, D. 2019 · 2019
Cited alongside, same era.
SlowFast Networks for Video Recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Cited alongside, same era.
W2VV++: Fully Deep Learning for Ad-hoc Video Search
Li, X.; Xu, C.; Yang, G.; Chen, Z.; and Dong, J. 2019 · 2019
Cited alongside, same era.
Video Classification With Channel-Separated Convolutional Networks
Tran, D.; Wang, H.; Feiszli, M.; and Torresani, L. 2019 · 2019
Cited alongside, same era.
A Graph-Based Framework to Bridge Movies and Synopses
Xiong, Y.; Huang, Q.; Guo, L.; Zhou, H.; Zhou, B.; and Lin, D. 2019 · 2019
Cited alongside, same era.
nuScenes: A Multimodal Dataset for Autonomous Driving
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2020 · 2020
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
Home Action Genome: Cooperative Compositional Action Understanding
Rai, N.; Chen, H.; Ji, J.; Desai, R.; Kozuka, K.; Ishizaka, S.; Adeli, E.; and Niebles, J. C. 2021 · 2021
Later among the works it cites.
Towards Long-Form Video Understanding
Wu, C.; and Krähenbühl, P. 2021 · 2021
Later among the works it cites.
Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
Gu, X.; Lin, T.; Kuo, W.; and Cui, Y. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MovieNet: A Holistic Dataset for Movie Understanding
Huang, Q.; Xiong, Y.; Rao, A.; Wang, J.; and Lin, D. 2020 · 2020
Cited alongside, same era.
Action Genome: Actions As Compositions of Spatio-Temporal Scene Graphs
Ji, J.; Krishna, R.; Fei-Fei, L.; and Niebles, J. C. 2020 · 2020
Cited alongside, same era.
Representation Learning on Visual-Symbolic Graphs for Video Understanding
Mavroudi, E.; Haro, B. B.; and Vidal, R. 2020 · 2020
Cited alongside, same era.
A Primer in BERTology: What We Know About How BERT Works
Rogers, A.; Kovaleva, O.; and Rumshisky, A. 2020 · 2020
Cited alongside, same era.
Fast Template Matching and Update for Video Object Tracking and Segmentation
Sun, M.; Xiao, J.; Lim, E. G.; Zhang, B.; and Zhao, Y. 2020 · 2020
Cited alongside, same era.
Is Space-Time Attention All You Need for Video Understanding?
Bertasius, G.; Wang, H.; and Torresani, L. 2021 · 2021
Cited alongside, same era.
Grounded Language-Image Pre-training
Li, L. H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.; Chang, K.; and Gao, J. 2022 · 2022
Later among the works it cites.
Grounding LTLf Specifications in Images
Umili, E.; Capobianco, R.; and De Giacomo, G. 2022 · 2022
Later among the works it cites.
Privacy-Preserving Deep Action Recognition: An Adversarial Learning Framework and A New Dataset
Wu, Z.; Wang, H.; Wang, Z.; Jin, H.; and Wang, Z. 2022 · 2022
Later among the works it cites.
Automaton-Based Representations of Task Knowledge from Generative Language Models
Yang, Y.; Gaglione, J.-R.; Neary, C.; and Topcu, U. 2022 · 2022
Later among the works it cites.
Deng, A.; Yang, T.; and Chen, C. 2023 · 2023
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; Dollár, P.; and Girshick, R. 2023 · 2023
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Li, C.; Yang, J.; Su, H.; Zhu, J.; et al. 2023 · 2023
Closest in time.
Safe Networked Robotics via Formal Verification
Narasimhan, S. S.; Bhat, S.; and Chinchali, S. P. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.