Fetching the paper…
Reading the bibliography…
Video Referring Expression Comprehension (REC) aims to localize a target object in video frames referred by the natural language expression.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
A large-scale hierarchical image database
Jia Deng · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
An improved non-monotonic transition system for dependency parsing
Matthew Honnibal and Mark Johnson · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Netvlad: Cnn architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Modeling relationships in referential expressions with compositional modular networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
A joint speaker-listener-reinforcer model for referring expressions
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Earlier work this paper cites.
Finding” it”: Weakly-supervised reference-aware visual grounding in instructional videos
De-An Huang, Shyamal Buch, Lucio Dery, Animesh Garg, Li Fei-Fei, and Juan Carlos Niebles · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Object referring in videos with language and human gaze
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
Weakly-supervised video object grounding from text by loss weighting and object interaction
Luowei Zhou, Nathan Louis, and Jason J Corso · 2018
Cited alongside, same era.
Gisca: Gradient-inductive segmentation network with contextual attention for scene text detection
Meng Cao, Yuexian Zou, Dongming Yang, and Chao Liu · 2019
Cited alongside, same era.
Weakly-supervised spatio-temporally grounding natural sentence in video
Zhenfang Chen, Lin Ma, Wenhan Luo, and Kwan-Yee K Wong · 2019
Cited alongside, same era.
Dual encoding for zero-example video retrieval
Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, and Xun Wang · 2019
Cited alongside, same era.
Learning 2d temporal adjacent networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Later among the works it cites.
Where does it exist: Spatio-temporal video grounding for multi-form sentences
Zhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang, Huasheng Liu, and Lianli Gao · 2020
Later among the works it cites.
On pursuit of designing multi-modal transformer for video grounding
Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou · 2021
Later among the works it cites.
Unifacegan: A unified framework for temporally consistent facial video editing
Meng Cao, Haozhi Huang, Hao Wang, Xuan Wang, Li Shen, Sheng Wang, Linchao Bao, Zhifeng Li, and Jiebo Luo · 2021
Later among the works it cites.
All you need is a second look: Towards arbitrary-shaped text detection
Meng Cao, Can Zhang, Dongming Yang, and Yuexian Zou · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Siamrpn++: Evolution of siamese visual tracking with very deep networks
Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Generalized intersection over union: A metric and a loss for bounding box regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Cited alongside, same era.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Cited alongside, same era.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Cited alongside, same era.
Cross-modal relationship inference for grounding referring expressions
Sibei Yang, Guanbin Li, and Yizhou Yu · 2019
Cited alongside, same era.
Yi-Wen Chen, Yi-Hsuan Tsai, and Ming-Hsuan Yang · 2021
Later among the works it cites.
Siamese natural language tracker: Tracking by natural language descriptions with siamese trackers
Qi Feng, Vitaly Ablavsky, Qinxun Bai, and Stan Sclaroff · 2021
Later among the works it cites.
Decoupled spatial temporal graphs for generic visual grounding
Qianyu Feng, Yunchao Wei, Mingming Cheng, and Yi Yang · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Later among the works it cites.
Conditional detr for fast training convergence
Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Co-grounding networks with semantic attention for referring expression comprehension in videos
Sijie Song, Xudong Lin, Jiaying Liu, Zongming Guo, and Shih-Fu Chang · 2021
Later among the works it cites.
Anchor detr: Query design for transformer-based detector
Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun · 2021
Later among the works it cites.
Rr-net: Relation reasoning for end-to-end human-object interaction detection
Dongming Yang, Yuexian Zou, Can Zhang, Meng Cao, and Jie Chen · 2021
Later among the works it cites.
Synergic learning for noise-insensitive webly-supervised temporal action localization
Can Zhang, Meng Cao, Dongming Yang, Ji Jiang, and Yuexian Zou · 2021
Later among the works it cites.
Correspondence matters for video referring expression comprehension
Meng Cao, Ji Jiang, Long Chen, and Yuexian Zou · 2022
Closest in time.
Locvtp: Video-text pre-training for temporal localization
Meng Cao, Tianyu Yang, Junwu Weng, Can Zhang, Jue Wang, and Yuexian Zou · 2022
Closest in time.
Deep motion prior for weakly-supervised temporal action localization
Meng Cao, Can Zhang, Long Chen, Mike Zheng Shou, and Yuexian Zou · 2022
Closest in time.
Visual relation-aware unsupervised video captioning
Puzhao Ji, Meng Cao, and Yuexian Zou · 2022
Closest in time.
Dab-detr: Dynamic anchor boxes are better queries for detr
Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, and Lei Zhang · 2022
Closest in time.
Tubedetr: Spatio-temporal video grounding with transformers
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Closest in time.