Fetching the paper…
Reading the bibliography…
With the increasing demand for video understanding, video moment and highlight detection (MHD) has emerged as a critical research topic.
Improving deep neural networks for LVCSR using rectified linear units and dropout. In 2013 IEEE international conference on acoustics, speech and signal processing . IEEE, 8609–8613
George E. Dahl, Tara N. Sainath, and Geoffrey E. Hinton. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation. In EMNLP . 1532–1543
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding. In CVPR . 961–970
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection. In CVPR . 3707–3715
Wu Liu, Tao Mei, Yongdong Zhang, Cherry Che, and Jiebo Luo. 2015 · 2015
Earlier work this paper cites.
Tvsum: Summarizing web videos using titles. In CVPR . 5179–5187
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks. In ICCV . 4489–4497
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In CVPR . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep networks with stochastic depth. In ECCV . Springer, 646–661
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. 2016 · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding. In ECCV . Springer, 510–526
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016 · 2016
Earlier work this paper cites.
To click or not to click: Automatic selection of beautiful thumbnails from videos. In CIKM . 659–668
Yale Song, Miriam Redi, Jordi Vallmitjana, and Alejandro Jaimes. 2016 · 2016
Earlier work this paper cites.
Video summarization with long short-term memory. In ECCV . Springer, 766–782
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language. In CVPR . 5803–5812
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR . 6299–6308
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, and Paul Natsev. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos. In ICCV . 706–715
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Unsupervised video summarization with adversarial lstm networks. In CVPR . 202–211
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In NeurIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video. In EMNLP . 162–171
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua. 2018 · 2018
Earlier work this paper cites.
Localizing Moments in Video with Temporal Language. In EMNLP
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2018 · 2018
Earlier work this paper cites.
Semantic proposal for activity localization in videos via sentence query. In AAAI , Vol. 33. 8199–8206
Shaoxiang Chen and Yu-Gang Jiang. 2019 · 2019
Earlier work this paper cites.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell. 2019 · 2019
Earlier work this paper cites.
Slowfast networks for video recognition. In CVPR . 6202–6211
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019 · 2019
Earlier work this paper cites.
Hierarchical graph semantic pooling network for multi-modal community question answer matching. In ACM MM . 1157–1165
Jun Hu, Shengsheng Qian, Quan Fang, and Changsheng Xu. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
A tensorized transformer for language modeling
Xindian Ma, Peng Zhang, Shuai Zhang, Nan Duan, Yuexian Hou, Ming Zhou, and Dawei Song. 2019 · 2019
Cited alongside, same era.
Generalized intersection over union: A metric and a loss for bounding box regression. In CVPR . 658–666
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. 2019 · 2019
Cited alongside, same era.
Less is more: Learning highlight detection from video duration. In CVPR . 1258–1267
Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram, and Kristen Grauman. 2019 · 2019
Cited alongside, same era.
Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders
Nicola Messina, Giuseppe Amato, Andrea Esuli, Fabrizio Falchi, Claudio Gennaro, and Stéphane Marchand-Maillet. 2021 · 2021
Later among the works it cites.
Interaction-integrated network for natural language moment localization
Ke Ning, Lingxi Xie, Jianzhuang Liu, Fei Wu, and Qi Tian. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In ICML . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, and Jack Clark. 2021 · 2021
Later among the works it cites.
DORi: Discovering object relationships for moment localization of a natural language query in a video. In WACV . 1079–1088
Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Basura Fernando, Hongdong Li, and Stephen Gould. 2021 · 2021
Later among the works it cites.
Efficient attention: Attention with linear complexities. In WACV . 3531–3539
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multilevel Language and Vision Integration for Text-to-Clip Retrieval. In AAAI , Vol. 33
Huijuan Xu, Kun He, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. 2019 · 2019
Cited alongside, same era.
Semantic conditioned dynamic modulation for temporal sentence grounding in videos. In NeurIPS , Vol. 32
Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. 2019 · 2019
Cited alongside, same era.
MAN: Moment alignment network for natural language moment retrieval via iterative graph adjustment. In CVPR . 1247–1257
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S. Davis. 2019 · 2019
Cited alongside, same era.
End-to-end object detection with transformers. In ECCV . Springer, 213–229
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, and Sylvain Gelly. 2020 · 2020
Cited alongside, same era.
Tvr: A large-scale dataset for video-subtitle moment retrieval. In ECCV . Springer, 447–463
Jie Lei, Licheng Yu, Tamara L. Berg, and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
Sparse and continuous attention mechanisms
André Martins, António Farinhas, Marcos Treviso, Vlad Niculae, Pedro Aguiar, and Mario Figueiredo. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
VLG-Net: Video-language graph matching network for video grounding. In ICCVW . 3224–3234
Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem. 2021 · 2021
Later among the works it cites.
Rethinking and improving relative position encoding for vision transformer. In ICCV . 10033–10041
Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu, and Hongyang Chao. 2021 · 2021
Later among the works it cites.
Cross-category video highlight detection via set-based learning. In CVPR . 7970–7979
Minghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu, Zhenbang Sun, and Changhu Wang. 2021 · 2021
Later among the works it cites.
Multi-modal interaction graph convolutional network for temporal language localization in videos
Zongmeng Zhang, Xianjing Han, Xuemeng Song, Yan Yan, and Liqiang Nie. 2021 · 2021
Later among the works it cites.
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
Xiang Fang, Daizong Liu, Pan Zhou, Zichuan Xu, and Ruixuan Li. 2022 · 2022
Later among the works it cites.
Fine-grained temporal contrastive learning for weakly-supervised temporal action localization. In CVPR . 19999–20009
Junyu Gao, Mengyuan Chen, and Changsheng Xu. 2022 · 2022
Later among the works it cites.
Swinbert: End-to-end transformers with sparse attention for video captioning. In CVPR . 17949–17958
Kevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed, Zhe Gan, Zicheng Liu, Yumao Lu, and Lijuan Wang. 2022 · 2022
Later among the works it cites.
A Survey on Video Moment Localization
Meng Liu, Liqiang Nie, Yunxiao Wang, Meng Wang, and Yong Rui. 2022d · 2022
Later among the works it cites.
Video summarization through reinforcement learning with a 3D spatio-temporal u-net
Tianrui Liu, Qingjie Meng, Jun-Jie Huang, Athanasios Vlontzos, Daniel Rueckert, and Bernhard Kainz. 2022c · 2022
Later among the works it cites.
End-to-end generative pretraining for multimodal video captioning. In CVPR . 17959–17968
Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid. 2022 · 2022
Later among the works it cites.
Sparse MLP for image recognition: Is self-attention really necessary?. In AAAI , Vol. 36. 2344–2351
Chuanxin Tang, Yucheng Zhao, Guangting Wang, Chong Luo, Wenxuan Xie, and Wenjun Zeng. 2022 · 2022
Later among the works it cites.
Language-enhanced object reasoning networks for video moment retrieval with text query
Gongmian Wang, Xun Jiang, Ning Liu, and Xing Xu. 2022 · 2022
Later among the works it cites.
Learning pixel-level distinctions for video highlight detection. In CVPR . 3073–3082
Fanyue Wei, Biao Wang, Tiezheng Ge, Yuning Jiang, Wen Li, and Lixin Duan. 2022 · 2022
Later among the works it cites.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A. Clifton. 2022b · 2022
Later among the works it cites.
HiSA: Hierarchically Semantic Associating for Video Temporal Grounding
Zhe Xu, Da Chen, Kun Wei, Cheng Deng, and Hui Xue. 2022a · 2022
Later among the works it cites.
Video moment retrieval with cross-modal neural architecture search
Xun Yang, Shanshan Wang, Jian Dong, Jianfeng Dong, Meng Wang, and Tat-Seng Chua. 2022 · 2022
Later among the works it cites.
Hierarchical modular network for video captioning. In CVPR . 17939–17948
Hanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang, Qingming Huang, and Ming-Hsuan Yang. 2022 · 2022
Later among the works it cites.
Metaformer is actually what you need for vision. In CVPR . 10819–10829
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. 2022 · 2022
Later among the works it cites.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, and Heung-Yeung Shum. 2022a · 2022
Later among the works it cites.