Fetching the paper…
Reading the bibliography…
Weakly-supervised temporal action localization (WTAL) learns to detect and classify action instances with only category labels.
Image-to-word transformation based on dividing and vector quantizing images with words
Yasuhide Mori, Hironobu Takahashi, and Ryuichi Oka · 1999
Earlier work this paper cites.
WSABIE: Scaling up to large vocabulary image annotation
Jason Weston, Samy Bengio, and Nicolas Usunier · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Joint inference of groups, events and human roles in aerial videos
Tianmin Shu, Dan Xie, Brandon Rothrock, Sinisa Todorovic, and Song Chun Zhu · 2015
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Highlight detection with pairwise deep ranking for first-person video summarization
Ting Yao, Tao Mei, and Yong Rui · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Turn tap: Temporal unit regression network for temporal action proposals
Jiyang Gao, Zhenheng Yang, Kan Chen, Chen Sun, and Ram Nevatia · 2017
Earlier work this paper cites.
Single shot temporal action detection
Tianwei Lin, Xu Zhao, and Zheng Shou · 2017
Earlier work this paper cites.
Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang · 2017
Earlier work this paper cites.
Untrimmednets for weakly supervised action recognition and detection
Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool · 2017
Earlier work this paper cites.
R-c3d: Region convolutional 3d network for temporal activity detection
Huijuan Xu, Abir Das, and Kate Saenko · 2017
Earlier work this paper cites.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Earlier work this paper cites.
Rethinking the faster r-cnn architecture for temporal action localization
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar · 2018
Earlier work this paper cites.
Ctap: Complementary temporal action proposal generation
Jiyang Gao, Kan Chen, and Ram Nevatia · 2018
Earlier work this paper cites.
Bsn: Boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Earlier work this paper cites.
Weakly supervised action localization by sparse temporal pooling network
Phuc Nguyen, Ting Liu, Gautam Prasad, and Bohyung Han · 2018
Earlier work this paper cites.
W-talc: Weakly-supervised temporal activity localization and classification
Sujoy Paul, Sourya Roy, and AmitK Roy-Chowdhury · 2018
Earlier work this paper cites.
Autoloc: Weakly-supervised temporal action localization in untrimmed videos
Zheng Shou, Hang Gao, Lei Zhang, Kazuyuki Miyazawa, and Shih-Fu Chang · 2018
Earlier work this paper cites.
Step-by-step erasion, one-by-one collection: A weakly supervised temporal action detector
Jia-Xing Zhong, Nannan Li, Weijie Kong, Tao Zhang, Thomas H Li, and Ge Li · 2018
Earlier work this paper cites.
Video imprint segmentation for temporal action detection in untrimmed videos
Zhanning Gao, Le Wang, Qilin Zhang, Zhenxing Niu, Nanning Zheng, and Gang Hua · 2019
Earlier work this paper cites.
Bmn: Boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Earlier work this paper cites.
Completeness modeling and context separation for weakly supervised temporal action localization
Daochang Liu, Tingting Jiang, and Yizhou Wang · 2019
Earlier work this paper cites.
Multi-granularity generator for temporal action proposal
Yuan Liu, Lin Ma, Yifeng Zhang, Wei Liu, and Shih-Fu Chang · 2019
Earlier work this paper cites.
Weakly supervised temporal action localization through contrast based evaluation networks
Ziyi Liu, Le Wang, Qilin Zhang, Zhanning Gao, Zhenxing Niu, Nanning Zheng, and Gang Hua · 2019
Earlier work this paper cites.
3c-net: Category count and center loss for weakly-supervised action localization
Sanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, and Ling Shao · 2019
Earlier work this paper cites.
Weakly-supervised action localization with background modeling
Phuc Xuan Nguyen, Deva Ramanan, and Charless C Fowlkes · 2019
Earlier work this paper cites.
Segregated temporal assembly recurrent networks for weakly supervised multiple action detection
Yunlu Xu, Chengwei Zhang, Zhanzhan Cheng, Jianwen Xie, Yi Niu, Shiliang Pu, and Fei Wu · 2019
Earlier work this paper cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan · 2019
Earlier work this paper cites.
Boundary content graph neural network for temporal action proposal generation
Yueran Bai, Yingying Wang, Yunhai Tong, Yang Yang, Qiyue Liu, and Junhui Liu · 2020
Earlier work this paper cites.
Chen Ju, Peisen Zhao, Ya Zhang, Yanfeng Wang, and Qi Tian · 2020
Cited alongside, same era.
Background suppression network for weakly-supervised temporal action localization
Pilhyeon Lee, Youngjung Uh, and Hyeran Byun · 2020
Cited alongside, same era.
Background suppression network for weakly-supervised temporal action localization
Pilhyeon Lee, Youngjung Uh, and Hyeran Byun · 2020
Cited alongside, same era.
Fast learning of temporal action proposal via dense boundary generator
Chuming Lin, Jian Li, Yabiao Wang, Ying Tai, Donghao Luo, Zhipeng Cui, Chengjie Wang, Jilin Li, Feiyue Huang, and Rongrong Ji · 2020
Cited alongside, same era.
Progressive boundary refinement network for temporal action detection
Qinying Liu and Zilei Wang · 2020
Cited alongside, same era.
Actionclip: A new paradigm for video action recognition
Mengmeng Wang, Jiazheng Xing, and Yong Liu · 2021
Later among the works it cites.
Uncertainty guided collaborative training for weakly supervised temporal action detection
Wenfei Yang, Tianzhu Zhang, Xiaoyuan Yu, Tian Qi, Yongdong Zhang, and Feng Wu · 2021
Later among the works it cites.
Open-vocabulary object detection using captions
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang · 2021
Later among the works it cites.
Cola: Weakly-supervised temporal action localization with snippet contrastive learning
Can Zhang, Meng Cao, Dongming Yang, Jie Chen, and Yuexian Zou · 2021
Later among the works it cites.
Enriching local and global contexts for temporal action localization
Zixin Zhu, Wei Tang, Le Wang, Nanning Zheng, and Gang Hua · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weakly-supervised action localization with expectation-maximization multi-instance learning
Zhekun Luo, Devin Guillory, Baifeng Shi, Wei Ke, Fang Wan, Trevor Darrell, and Huijuan Xu · 2020
Cited alongside, same era.
Sf-net: Single-frame supervision for temporal action localization
Fan Ma, Linchao Zhu, Yi Yang, Shengxin Zha, Gourab Kundu, Matt Feiszli, and Zheng Shou · 2020
Cited alongside, same era.
Adversarial background-aware loss for weakly-supervised temporal activity localization
Kyle Min and Jason J Corso · 2020
Cited alongside, same era.
Weakly-supervised action localization by generative attention modeling
Baifeng Shi, Qi Dai, Yadong Mu, and Jingdong Wang · 2020
Cited alongside, same era.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem · 2020
Cited alongside, same era.
Revisiting anchor mechanisms for temporal action localization
Le Yang, Houwen Peng, Dingwen Zhang, Jianlong Fu, and Junwei Han · 2020
Cited alongside, same era.
Two-stream consensus network for weakly-supervised temporal action localization
Yuanhao Zhai, Le Wang, Wei Tang, Qilin Zhang, Junsong Yuan, and Gang Hua · 2020
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Closest in time.
Dual-evidential learning for weakly-supervised temporal action localization
Mengyuan Chen, Junyu Gao, Shicai Yang, and Changsheng Xu · 2022
Closest in time.
Tallformer: Temporal action localization with long-memory transformer
Feng Cheng and Gedas Bertasius · 2022
Closest in time.
Vqgan-clip: Open domain image generation and editing with natural language guidance
Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff · 2022
Closest in time.
Fine-grained temporal contrastive learning for weakly-supervised temporal action localization
Junyu Gao, Mengyuan Chen, and Changsheng Xu · 2022
Closest in time.
Asm-loc: Action-aware segment modeling for weakly-supervised temporal action localization
Bo He, Xitong Yang, Le Kang, Zhiyu Cheng, Xin Zhou, and Abhinav Shrivastava · 2022
Closest in time.
Weakly supervised temporal action localization via representative snippet knowledge propagation
Linjiang Huang, Liang Wang, and Hongsheng Li · 2022
Closest in time.
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Closest in time.
Adaptive mutual supervision for weakly-supervised temporal action localization
Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang, Xiaoyun Zhang, and Qi Tian · 2022
Closest in time.
Simple but effective: Clip embeddings for embodied ai
Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Closest in time.
Exploring denoised cross-video contrast for weakly-supervised temporal action localization
Jingjing Li, Tianyu Yang, Wei Ji, Jue Wang, and Li Cheng · 2022
Closest in time.
Gen-vlkt: Simplify association and enhance interaction understanding for hoi detection
Yue Liao, Aixi Zhang, Miao Lu, Yongliang Wang, Xiaobo Li, and Si Liu · 2022
Closest in time.
Frozen clip models are efficient video learners
Ziyi Lin, Shijie Geng, Renrui Zhang, Peng Gao, Gerard de Melo, Xiaogang Wang, Jifeng Dai, Yu Qiao, and Hongsheng Li · 2022
Closest in time.
An empirical study of end-to-end temporal action detection
Xiaolong Liu, Song Bai, and Xiang Bai · 2022
Closest in time.
Temporal action detection with global segmentation mask learning
Sauradip Nag, Xiatian Zhu, Yi-Zhe Song, and Tao Xiang · 2022
Closest in time.
Zero-shot temporal action detection via vision-language prompting
Sauradip Nag, Xiatian Zhu, Yi-Zhe Song, and Tao Xiang · 2022
Closest in time.
Denseclip: Language-guided dense prediction with context-aware prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu · 2022
Closest in time.
Motionclip: Exposing human motion generation to clip space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or · 2022
Closest in time.
Clip-nerf: Text-and-image driven manipulation of neural radiance fields
Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao · 2022
Closest in time.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Closest in time.
Rcl: Recurrent continuous localization for temporal action detection
Qiang Wang, Yanhao Zhang, Yun Zheng, and Pan Pan · 2022
Closest in time.
Temporal action proposal generation with background constraint
Haosen Yang, Wenhao Wu, Lining Wang, Sheng Jin, Boyang Xia, Hongxun Yao, and Hujie Huang · 2022
Closest in time.
Acgnet: Action complement graph network for weakly-supervised temporal action localization
Zichen Yang, Jie Qin, and Di Huang · 2022
Closest in time.
Filip: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Actionformer: Localizing moments of actions with transformers
Chenlin Zhang, Jianxin Wu, and Yin Li · 2022
Closest in time.
Extract free dense labels from clip
Chong Zhou, Chen Change Loy, and Bo Dai · 2022
Closest in time.