Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of temporal video grounding (TVG), which aims to predict the starting/ending time points of moments described by a text sentence within a long untrimmed video.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Adversarial reprogramming of neural networks
Gamaleldin F Elsayed, Ian Goodfellow, et al · 2018
Earlier work this paper cites.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
Cross-modal moment localization in videos
Meng Liu, Xiang Wang, Liqiang Nie, Qi Tian, Baoquan Chen, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
Semantic proposal for activity localization in videos via sentence query
Shaoxiang Chen and Yu-Gang Jiang · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Excl: Extractive clip localization using natural language descriptions
Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexander Hauptmann · 2019
Earlier work this paper cites.
Tripping through time: Efficient localization of activities in videos
Meera Hahn, Asim Kadav, James M Rehg, and Hans Peter Graf · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Debug: A dense bottom-up grounding approach for natural language video localization
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Cited alongside, same era.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Cited alongside, same era.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
Cited alongside, same era.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, Tao Mei, and Wenwu Zhu · 2019
A survey on temporal sentence grounding in videos
Xiaohan Lan, Yitian Yuan, Xin Wang, Zhi Wang, and Wenwu Zhu · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
Proposal-free video grounding with contextual pyramid network
Kun Li, Dan Guo, and Meng Wang · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Cited alongside, same era.
Appearance-preserving 3d convolution for video-based person re-identification
Xinqian Gu, Hong Chang, Bingpeng Ma, Hongkai Zhang, and Xilin Chen · 2020
Cited alongside, same era.
Tensor fista-net for real-time snapshot compressive imaging
Xiaochen Han, Bo Wu, Zheng Shou, Xiao-Yang Liu, Yimeng Zhang, and Linghe Kong · 2020
Cited alongside, same era.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu · 2020
Cited alongside, same era.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen · 2020
Cited alongside, same era.
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang · 2021
Later among the works it cites.
Natural language video localization with learnable moment proposals
Shaoning Xiao, Long Chen, Jian Shao, Yueting Zhuang, and Jun Xiao · 2021
Later among the works it cites.
Boundary proposal network for two-stage natural language video localization
Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao · 2021
Later among the works it cites.
Voice2series: Reprogramming acoustic models for time series classification
Chao-Han Huck Yang, Yun-Yun Tsai, et al · 2021
Later among the works it cites.
Cpt: Colorful prompt tuning for pre-trained vision-language models
Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models
Ofir Zafrir, Ariel Larey, Guy Boudoukh, Haihao Shen, and Moshe Wasserblat · 2021
Later among the works it cites.
Multi-modal relational graph for cross-modal video moment retrieval
Yawen Zeng, Da Cao, Xiaochi Wei, Meng Liu, Zhou Zhao, and Zheng Qin · 2021
Later among the works it cites.
Why adversarial reprogramming works, when it fails, and how to tell the difference
Yang Zheng, Xiaoyi Feng, et al · 2021
Later among the works it cites.
Exploring visual prompts for adapting large-scale models
H Bahng, A Jahanian, S Sankaranarayanan, and P Isola · 2022
Later among the works it cites.
Visual prompting for adversarial robustness
Aochuan Chen, Peter Lorenz, Yuguang Yao, Pin-Yu Chen, and Sijia Liu · 2022
Later among the works it cites.
Understanding and improving visual prompting: A label-mapping perspective
Aochuan Chen, Yuguang Yao, Pin-Yu Chen, Yihua Zhang, and Sijia Liu · 2022
Later among the works it cites.
Model reprogramming: Resource-efficient cross-domain machine learning
Pin-Yu Chen · 2022
Later among the works it cites.
On the robustness of deep learning-based mri reconstruction to image transformations
Jinghan Jia, Mingyi Hong, Yimeng Zhang, Mehmet Akçakaya, and Sijia Liu · 2022
Later among the works it cites.
Maple: Multi-modal prompt learning
Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan · 2022
Later among the works it cites.
Bridge-prompt: Towards ordinal action understanding in instructional videos
Muheng Li, Lei Chen, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
Yanyu Li, Pu Zhao, Geng Yuan, Xue Lin, Yanzhi Wang, and Xin Chen · 2022
Later among the works it cites.
Guanhua Zhang, Yihua Zhang, Yang Zhang, Wenqi Fan, Qing Li, Sijia Liu, and Shiyu Chang · 2022
Later among the works it cites.
The elements of temporal sentence grounding in videos: A survey and future directions
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2022
Later among the works it cites.
How to robustify black-box ml models? a zeroth-order optimization perspective
Yimeng Zhang, Yuguang Yao, Jinghan Jia, Jinfeng Yi, Mingyi Hong, Shiyu Chang, and Sijia Liu · 2022
Later among the works it cites.
Advancing model pruning via bi-level optimization
Yihua Zhang, Yuguang Yao, Parikshit Ram, Pu Zhao, Tianlong Chen, Mingyi Hong, Yanzhi Wang, and Sijia Liu · 2022
Later among the works it cites.
Revisiting and advancing fast adversarial training through the lens of bi-level optimization
Yihua Zhang, Guanhua Zhang, Prashant Khanduri, Mingyi Hong, Shiyu Chang, and Sijia Liu · 2022
Later among the works it cites.
Data-model-circuit tri-design for ultra-light video intelligence on edge devices
Yimeng Zhang, Akshay Karkal Kamath, Qiucheng Wu, Zhiwen Fan, Wuyang Chen, Zhangyang Wang, Shiyu Chang, Sijia Liu, and Cong Hao · 2023
Closest in time.