Fetching the paper…
Reading the bibliography…
Video moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip.
Temporal Localization of Moments in Video Collections with Natural Language
Escorcia, V.; Soldan, M.; Sivic, J.; Ghanem, B.; and Russell, B. C. 2019 · 1907
Earlier work this paper cites.
ExCL: Extractive Clip Localization Using Natural Language Descriptions
Ghosh, S.; Agarwal, A.; Parekh, Z.; and Hauptmann, A. 2019 · 1990
Earlier work this paper cites.
Univl: A unified video and language pre-training model for multimodal understanding and generation
Luo, H.; Ji, L.; Shi, B.; Huang, H.; Duan, N.; Li, T.; Li, J.; Bharti, T.; and Zhou, M. 2020 · 2002
Earlier work this paper cites.
The graph neural network model
Scarselli, F.; Gori, M.; Tsoi, A. C.; Hagenbuchner, M.; and Monfardini, G. 2008 · 2008
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J.; Gulcehre, C.; Cho, K.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Glove: Global Vectors for Word Representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection
Liu, W.; Mei, T.; Zhang, Y.; Che, C.; and Luo, J. 2015 · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan, K.; and Zisserman, A. 2015 · 2015
Earlier work this paper cites.
TVSum: Summarizing web videos using titles
Song, Y.; Vallmitjana, J.; Stent, A.; and Jaimes, A. 2015 · 2015
Earlier work this paper cites.
To Click or Not To Click: Automatic Selection of Beautiful Thumbnails from Videos
Song, Y.; Redi, M.; Vallmitjana, J.; and Jaimes, A. 2016 · 2016
Earlier work this paper cites.
Video Summarization with Long Short-Term Memory
Zhang, K.; Chao, W.; Sha, F.; and Grauman, K. 2016 · 2016
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
TALL: Temporal Activity Localization via Language Query
Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017 · 2017
Earlier work this paper cites.
Audio Set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P. W.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
Words speak for actions: Using text to find video highlights
Kudi, S.; and Namboodiri, A. M. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Dynamic Coattention Networks For Question Answering
Xiong, C.; Zhong, V.; and Socher, R. 2017 · 2017
Earlier work this paper cites.
Localizing Moments in Video with Temporal Language
Hendricks, L. A.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. C. 2018 · 2018
Earlier work this paper cites.
PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation
Molino, A. G. D.; and Gygli, M. 2018 · 2018
Earlier work this paper cites.
Semantic Proposal for Activity Localization in Videos via Sentence Query
Chen, S.; and Jiang, Y. 2019 · 2019
Earlier work this paper cites.
SlowFast Networks for Video Recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Cited alongside, same era.
MAC: Mining Activity Concepts for Language-Based Temporal Localization
Ge, R.; Gao, J.; Chen, K.; and Nevatia, R. 2019 · 2019
Cited alongside, same era.
DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video Localization
Lu, C.; Chen, L.; Tan, C.; Li, X.; and Xiao, J. 2019 · 2019
Cited alongside, same era.
Less Is More: Learning Highlight Detection From Video Duration
Xiong, B.; Kalantidis, Y.; Ghadiyaram, D.; and Grauman, K. 2019 · 2019
Cited alongside, same era.
Multilevel Language and Vision Integration for Text-to-Clip Retrieval
Xu, H.; He, K.; Plummer, B. A.; Sigal, L.; Sclaroff, S.; and Saenko, K. 2019 · 2019
Cited alongside, same era.
Sentence Specified Dynamic Video Thumbnail Generation
Yuan, Y.; Ma, L.; and Zhu, W. 2019 · 2019
Joint Visual and Audio Learning for Video Highlight Detection
Badamdorj, T.; Rochan, M.; Wang, Y.; and Cheng, L. 2021 · 2021
Later among the works it cites.
Detecting Moments and Highlights in Videos via Natural Language Queries
Lei, J.; Berg, T. L.; and Bansal, M. 2021 · 2021
Later among the works it cites.
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Li, J.; Selvaraju, R. R.; Gotmare, A.; Joty, S. R.; Xiong, C.; and Hoi, S. C. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
Cross-category Video Highlight Detection via Set-based Learning
Xu, M.; Wang, H.; Ni, B.; Zhu, R.; Sun, Z.; and Wang, C. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment
Zhang, D.; Dai, X.; Wang, X.; Wang, Y.; and Davis, L. S. 2019 · 2019
Cited alongside, same era.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Cited alongside, same era.
COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
Ging, S.; Zolfaghari, M.; Pirsiavash, H.; and Brox, T. 2020 · 2020
Cited alongside, same era.
Tripping through time: Efficient Localization of Activities in Videos
Hahn, M.; Kadav, A.; Rehg, J. M.; and Graf, H. P. 2020 · 2020
Cited alongside, same era.
MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection
Hong, F.; Huang, X.; Li, W.; and Zheng, W. 2020 · 2020
Cited alongside, same era.
PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
Kong, Q.; Cao, Y.; Iqbal, T.; Wang, Y.; Wang, W.; and Plumbley, M. D. 2020 · 2020
Cited alongside, same era.
Ye, Q.; Shen, X.; Gao, Y.; Wang, Z.; Bi, Q.; Li, P.; and Yang, G. 2021 · 2021
Later among the works it cites.
Natural language video localization: A revisit in span-based question answering framework
Zhang, H.; Sun, A.; Jing, W.; Zhen, L.; Zhou, J. T.; and Goh, R. S. M. 2021 · 2021
Later among the works it cites.
TaoHighlight: Commodity-Aware Multi-Modal Video Highlight Detection in E-Commerce
Guo, Z.; Zhao, Z.; Jin, W.; Wang, D.; Liu, R.; and Yu, J. 2022 · 2022
Later among the works it cites.
Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning
Li, J.; Xie, J.; Qian, L.; Zhu, L.; Tang, S.; Wu, F.; Yang, Y.; Zhuang, Y.; and Wang, X. E. 2022 · 2022
Later among the works it cites.
Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding
Wang, Z.; Wang, L.; Wu, T.; Li, T.; and Wu, G. 2022 · 2022
Later among the works it cites.
Multimodal Learning With Transformers: A Survey
Xu, P.; Zhu, X.; and Clifton, D. A. 2022 · 2022
Later among the works it cites.
Video Moment Retrieval with Hierarchical Contrastive Learning
Zhang, B.; Yang, C.; Jiang, B.; and Zhou, X. 2022 · 2022
Later among the works it cites.
System-Status-Aware Adaptive Network for Online Streaming Video Understanding
Foo, L. G.; Gong, J.; Fan, Z.; and Liu, J. 2023 · 2023
Later among the works it cites.
UniVTG: Towards Unified Video-Language Temporal Grounding
Lin, K. Q.; Zhang, P.; Chen, J.; Pramanick, S.; Gao, D.; Wang, A. J.; Yan, R.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Query-dependent video representation for moment retrieval and highlight detection
Moon, W.; Hyun, S.; Park, S.; Park, D.; and Heo, J.-P. 2023 · 2023
Later among the works it cites.
Dual-Stream Multimodal Learning for Topic-Adaptive Video Highlight Detection
Xiong, Z.; and Wang, H. 2023 · 2023
Later among the works it cites.
MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer
Xu, Y.; Sun, Y.; Li, Y.; Shi, Y.; Zhu, X.; and Du, S. 2023 · 2023
Later among the works it cites.
Multimodal feature fusion based on object relation for video captioning
Yan, Z.; Chen, Y.; Song, J.; and Zhu, J. 2023 · 2023
Later among the works it cites.
Temporal Sentence Grounding in Videos: A Survey and Future Directions
Zhang, H.; Sun, A.; Jing, W.; and Zhou, J. T. 2023 · 2023
Later among the works it cites.