Fetching the paper…
Reading the bibliography…
Open-Vocabulary Temporal Action Localization (OVTAL) enables a model to recognize any desired action category in videos without the need to explicitly curate training data for all categories.
ActivityNet: A large-scale video benchmark for human activity understanding
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles · 2015
Earlier work this paper cites.
DAPs: Deep action proposals for action understanding
V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem · 2016
Earlier work this paper cites.
Fast temporal activity proposals for efficient detection of human actions in untrimmed videos
F. C. Heilbron, J. C. Niebles, and B. Ghanem · 2016
Earlier work this paper cites.
SST: Single-stream temporal action proposals
S. Buch, V. Escorcia, C. Shen, B. Ghanem, and J. C. Niebles · 2017
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
The THUMOS challenge on action recognition for videos “in the wild”
H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Diagnosing error in temporal action detectors
H. Alwassel, F. C. Heilbron, V. Escorcia, and B. Ghanem · 2018
Earlier work this paper cites.
Rethinking the faster R-CNN architecture for temporal action localization
Y.-W. Chao, S. Vijayanarasimhan, B. Seybold, D. A. Ross, J. Deng, and R. Sukthankar · 2018
Earlier work this paper cites.
BSN: Boundary sensitive network for temporal action proposal generation
T. Lin, X. Zhao, H. Su, C. Wang, and M. Yang · 2018
Earlier work this paper cites.
BMN: Boundary-matching network for temporal action proposal generation
T. Lin, X. Liu, X. Li, E. Ding, and S. Wen · 2019
Earlier work this paper cites.
Multi-granularity generator for temporal action proposal
Y. Liu, L. Ma, Y. Zhang, W. Liu, and S.-F. Chang · 2019
Earlier work this paper cites.
Graph convolutional networks for temporal action localization
R. Zeng, W. Huang, M. Tan, Y. Rong, P. Zhao, J. Huang, and C. Gan · 2019
Earlier work this paper cites.
Hacs: Human action clips and segments dataset for recognition and temporal localization
H. Zhao, A. Torralba, L. Torresani, and Z. Yan · 2019
Earlier work this paper cites.
Boundary content graph neural network for temporal action proposal generation
Y. Bai, Y. Wang, Y. Tong, Y. Yang, Q. Liu, and J. Liu · 2020
Earlier work this paper cites.
Scale matters: Temporal scale aggregation network for precise action localization in untrimmed videos
G. Gong, L. Zheng, and Y. Mu · 2020
Earlier work this paper cites.
How can we know what language models know?
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig · 2020
Cited alongside, same era.
Progressive boundary refinement network for temporal action detection
Q. Liu and Z. Wang · 2020
Cited alongside, same era.
G-TAD: Sub-graph localization for temporal action detection
M. Xu, C. Zhao, D. S. Rojas, A. Thabet, and B. Ghanem · 2020
Cited alongside, same era.
ZSTAD: Zero-shot temporal activity detection
L. Zhang, X. Chang, J. Liu, M. Luo, S. Wang, Z. Ge, and A. Hauptmann · 2020
Cited alongside, same era.
Bottom-up temporal action localization with mutual regularization
P. Zhao, L. Xie, C. Ju, Y. Zhang, Y. Wang, and Q. Tian · 2020
Cited alongside, same era.
Distance-IoU loss: Faster and better learning for bounding box regression
Z. Zheng, P. Wang, W. Liu, J. Li, R. Ye, and D. Ren · 2020
Cited alongside, same era.
TALLFormer: Temporal action localization with a long-memory transformer
F. Cheng and G. Bertasius · 2022
Later among the works it cites.
Prompting visual-language models for efficient video understanding
C. Ju, T. Han, K. Zheng, Y. Zhang, and W. Xie · 2022
Later among the works it cites.
Zero-shot temporal action detection via vision-language prompting
S. Nag, X. Zhu, Y.-Z. Song, and T. Xiang · 2022
Later among the works it cites.
Self-supervised video transformer
Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, and Michael S Ryoo · 2022
Later among the works it cites.
DualCoOp: Fast adaptation to multi-label recognition with limited annotations
X. Sun, P. Hu, and K. Saenko · 2022
Later among the works it cites.
Spatio-temporal relation modeling for few-shot action recognition
Anirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, and Bernard Ghanem · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero-shot detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Cited alongside, same era.
Learning salient boundary feature for anchor-free temporal action localization
C. Lin, C. Xu, D. Luo, Y. Wang, Y. Tai, C. Wang, J. Li, F. Huang, and Y. Fu · 2021
Cited alongside, same era.
Few-shot temporal action localization with query adaptive transformer
S. Nag, X. Zhu, and T. Xiang · 2021
Cited alongside, same era.
D2-net: Weakly-supervised action localization via discriminative embeddings and denoised activations
Sanath Narayan, Hisham Cholakkal, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, et al · 2021
Cited alongside, same era.
Relaxed transformer decoders for direct action proposal generation
J. Tan, J. Tang, L. Wang, and G. Wu · 2021
Cited alongside, same era.
Later among the works it cites.
ActionFormer: Localizing moments of actions with transformers
C.-L. Zhang, J. Wu, and Y. Li · 2022
Later among the works it cites.
RegionCLIP: Region-based language-image pretraining
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, et al · 2022
Later among the works it cites.
Cascade evidential learning for open-world weakly-supervised temporal action localization
M. Chen, J. Gao, and C. Xu · 2023
Later among the works it cites.
Multi-modal classifiers for open-vocabulary object detection
P. Kaul, W. Xie, and A. Zisserman · 2023
Later among the works it cites.
Few-shot common action localization via cross-attentional fusion of context and temporal dynamics
J. Lee, M. Jain, and S. Yun · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig · 2023
Later among the works it cites.
Semantics guided contrastive learning of transformers for zero-shot temporal activity detection
S. Nag, O. Goldstein, and A. K. Roy-Chowdhury · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Later among the works it cites.
ZEETAD: Adapting pretrained vision-language model for zero-shot end-to-end temporal action detection
T. Phan, K. Vo, D. Le, G. Doretto, D. Adjeroh, and N. Le · 2024
Closest in time.