Fetching the paper…
Reading the bibliography…
Video moment retrieval and highlight detection have received attention in the current era of video content proliferation, aiming to localize moments and estimate clip relevances based on user-specific queries.
H. W. Kuhn, “The hungarian method for the assignment problem,” Naval Research Logistics Quarterly , vol. 2, no. 1-2, pp. 83–97, 1955
1955
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in International Conference on Artificial Intelligence and Statistics (AISTATS) , 2010
2010
Earlier work this paper cites.
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele, “Script data for attribute-based recognition of composite activities,” in European Conference on Computer Vision (ECCV) , 2012
2012
Earlier work this paper cites.
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,” Transactions of the Association for Computational Linguistics , vol. 1, pp. 25–36, 2013
2013
Earlier work this paper cites.
M. Sun, A. Farhadi, and S. Seitz, “Ranking domain-specific highlights by analyzing edited videos,” in European Conference on Computer Vision (ECCV) , 2014
2014
Earlier work this paper cites.
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes, “Tvsum: Summarizing web videos using titles,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2015
2015
Earlier work this paper cites.
H. Yang, B. Wang, S. Lin, D. Wipf, M. Guo, and B. Guo, “Unsupervised extraction of video highlights via robust recurrent auto-encoders,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2015
2015
Earlier work this paper cites.
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
M. Gygli, Y. Song, and L. Cao, “Video2gif: Automatic generation of animated gifs from video,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
G. Sigurdsson, G. Varol, X. Wang, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” in European Conference on Computer Vision (ECCV) , 2016
2016
Earlier work this paper cites.
K. Zhang, W.-L. Chao, F. Sha, and K. Grauman, “Video summarization with long short-term memory,” in European Conference on Computer Vision (ECCV) , 2016
2016
Earlier work this paper cites.
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
B. Mahasseni, M. Lam, and S. Todorovic, “Unsupervised video summarization with adversarial lstm networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
M. Liu, X. Wang, L. Nie, Q. Tian, B. Chen, and T.-S. Chua, “Cross-modal moment localization in videos,” in the 26th ACM International Conference on Multimedia (ACM MM) , 2018
2018
Earlier work this paper cites.
J. Chen, X. Chen, L. Ma, Z. Jie, and T.-S. Chua, “Temporally grounding natural sentence in video,” in Empirical Methods in Natural Language Processing (EMNLP) , 2018
2018
Earlier work this paper cites.
H. Xu, K. He, B. A. Plummer, L. Sigal, S. Sclaroff, and K. Saenko, “Multilevel language and vision integration for text-to-clip retrieval,” in AAAI Conference on Artificial Intelligence (AAAI) , 2019
2019
Earlier work this paper cites.
S. Chen and Y.-G. Jiang, “Semantic proposal for activity localization in videos via sentence query,” in AAAI Conference on Artificial Intelligence (AAAI) , 2019
2019
Earlier work this paper cites.
D. Zhang, X. Dai, X. Wang, Y.-F. Wang, and L. S. Davis, “Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Earlier work this paper cites.
B. Xiong, Y. Kalantidis, D. Ghadiyaram, and K. Grauman, “Less is more: Learning highlight detection from video duration,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Advances in Neural Information Processing Systems (NeurIPS) , 2019
2019
Earlier work this paper cites.
W. Wang, Y. Huang, and L. Wang, “Language-driven temporal activity localization: A semantic matching reinforcement learning model,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. Ge, J. Gao, K. Chen, and R. Nevatia, “Mac: Mining activity concepts for language-based temporal localization,” in IEEE Winter Conference on Applications of Computer Vision (WACV) , 2019
2019
Earlier work this paper cites.
B. Jiang, X. Huang, C. Yang, and J. Yuan, “Cross-modal video moment retrieval with spatial and language-temporal attention,” in International Conference on Multimedia Retrieval (ICMR) , 2019
2019
Earlier work this paper cites.
J. Mun, M. Cho, and B. Han, “Local-global video-text interactions for temporal grounding,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020
2020
Cited alongside, same era.
J. Wu, G. Li, S. Liu, and L. Lin, “Tree-structured policy based progressive reinforcement learning for temporally language grounding in video,” in AAAI Conference on Artificial Intelligence (AAAI) , 2020
2020
Cited alongside, same era.
F.-T. Hong, X. Huang, W.-H. Li, and W.-S. Zheng, “Mini-net: Multiple instance ranking network for video highlight detection,” in European Conference on Computer Vision (ECCV) , 2020
2020
Cited alongside, same era.
J. Lei, L. Yu, T. L. Berg, and M. Bansal, “Tvr: A large-scale dataset for video-subtitle moment retrieval,” in European Conference on Computer Vision (ECCV) , 2020
2020
Cited alongside, same era.
W. Pan, Z. Zhao, W. Huang, Z. Zhang, L. Fu, Z. Pan, J. Yu, and F. Wu, “Video moment retrieval with noisy labels,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–13, 2022
2022
Later among the works it cites.
Y. Liu, S. Li, Y. Wu, C.-W. Chen, Y. Shan, and X. Qie, “Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2022
Later among the works it cites.
Y. Yuan, L. Ma, J. Wang, W. Liu, and W. Zhu, “Semantic conditioned dynamic modulation for temporal sentence grounding in videos,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 5, pp. 2725–2741, 2022
2022
Later among the works it cites.
X. Ding, N. Wang, S. Zhang, Z. Huang, X. Li, M. Tang, T. Liu, and X. Gao, “Exploring language hierarchy for video grounding,” IEEE Transactions on Image Processing , vol. 31, pp. 4693–4706, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
H. Zhang, A. Sun, W. Jing, and J. T. Zhou, “Span-based localizing network for natural language video localization,” in Annual Meeting of the Association for Computational Linguistics (ACL) , 2020
2020
Cited alongside, same era.
Z. Lin, Z. Zhao, Z. Zhang, Z. Zhang, and D. Cai, “Moment retrieval via cross-modal interaction networks with query reconstruction,” IEEE Transactions on Image Processing , vol. 29, pp. 3750–3762, 2020
2020
Cited alongside, same era.
D. Liu, X. Qu, X.-Y. Liu, J. Dong, P. Zhou, and Z. Xu, “Jointly cross- and self-modal graph attention network for query-based moment localization,” in the 28th ACM International Conference on Multimedia (ACM MM) , 2020
2020
Cited alongside, same era.
X. Qu, P. Tang, J. Dong, P. Zhou, and Z. Xu, “Fine-grained iterative attention network for temporal language localization in videos,” in the 28th ACM International Conference on Multimedia (ACM MM) , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
C. Chen and X. Gu, “Semantic modulation based residual network for temporal language queries grounding in video,” in International Symposium on Neural Networks (ISNN) , 2020, pp. 119–129
2020
Cited alongside, same era.
L. Wang, D. Liu, R. Puri, and D. N. Metaxas, “Learning trailer moments in full-length movies with co-contrastive attention,” in European Conference on Computer Vision (ECCV) , 2020
2020
Cited alongside, same era.
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. Mcgrew, I. Sutskever, and M. Chen, “GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models,” in Proceedings of the 39th International Conference on Machine Learning (ICML) , 2022
2022
Later among the works it cites.
R. Huang, Z. Zhao, H. Liu, J. Liu, C. Cui, and Y. Ren, “Prodiff: Progressive fast diffusion model for high-quality text-to-speech,” in the 30th ACM International Conference on Multimedia (ACM MM) , 2022
2022
Later among the works it cites.
X. Li, J. Thickstun, I. Gulrajani, P. S. Liang, and T. B. Hashimoto, “Diffusion-lm improves controllable text generation,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Later among the works it cites.
D. Baranchuk, A. Voynov, I. Rubachev, V. Khrulkov, and A. Babenko, “Label-efficient semantic segmentation with diffusion models,” in International Conference on Learning Representations (ICLR) , 2022
2022
Later among the works it cites.
J. Wolleb, R. Sandkühler, P. Valmaggia, and P. Cattin, “Diffusion models for implicit image segmentation ensembles,” in International Conference on Medical Imaging with Deep Learning (MIDL) , 2022
2022
Later among the works it cites.
H. Zhang, A. Sun, W. Jing, and J. T. Zhou, “Temporal sentence grounding in videos: A survey and future directions,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 8, pp. 10 443–10 465, 2023
2023
Closest in time.
W. Moon, S. Hyun, S. Park, D. Park, and J.-P. Heo, “Query-dependent video representation for moment retrieval and highlight detection,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Closest in time.
K. Lin, P. Zhang, J. Chen, S. Pramanick, D. Gao, and M. Shou, “Univtg: Towards unified video-language temporal grounding,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Closest in time.
N. Liu, X. Sun, H. Yu, F. Yao, G. Xu, and K. Fu, “M 2 dcapsn: Multimodal, multichannel, and dual-step capsule network for natural language moment localization,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–15, 2023
2023
Closest in time.
H. Wang, Z.-J. Zha, L. Li, X. Chen, and J. Luo, “Context-aware proposal–boundary network with structural consistency for audiovisual event localization,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–11, 2023
2023
Closest in time.
G. Li, D. Cheng, X. Ding, N. Wang, J. Li, and X. Gao, “Weakly supervised temporal action localization with bidirectional semantic consistency constraint,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–14, 2023
2023
Closest in time.
X.-Y. Zhang, C. Li, H. Shi, X. Zhu, P. Li, and J. Dong, “Adapnet: Adaptability decomposing encoder–decoder network for weakly supervised action recognition and localization,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1852–1863, 2023
2023
Closest in time.
S. Chen, P. Sun, Y. Song, and P. Luo, “Diffusiondet: Diffusion model for object detection,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Closest in time.
A. O. Tur, N. Dall’Asen, C. Beyan, and E. Ricci, “Exploring diffusion models for unsupervised video anomaly detection,” in 2023 IEEE International Conference on Image Processing (ICIP) , 2023
2023
Closest in time.
P. Jin, H. Li, Z. Cheng, K. Li, X. Ji, C. Liu, L. Yuan, and J. Chen, “Diffusionret: Generative text-video retrieval with diffusion model,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Closest in time.
2023
Closest in time.
R. Feng, Y. Gao, E. Tse, X. Ma, and H. J. Chang, “Diffpose: Spatiotemporal diffusion model for video-based human pose estimation,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Closest in time.
S. Nag, X. Zhu, J. Deng, Y.-Z. Song, and T. Xiang, “Difftad: Temporal action detection with proposal denoising diffusion,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Closest in time.
D. Liu, Q. Li, A.-D. Dinh, T. Jiang, M. Shah, and C. Xu, “Diffusion action segmentation,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Closest in time.
P. Li, C.-W. Xie, H. Xie, L. Zhang, Y. Zheng, D. Zhao, and Y. Zhang, “Momentdiff: Generative video moment retrieval from random to real,” in Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Closest in time.
2023
Closest in time.
N. Huang, Y. Zhang, F. Tang, W. Dong, and C. Xu, “Diffstyler: Controllable dual diffusion for text-driven image stylization,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–14, 2024
2024
Closest in time.