Fetching the paper…
Reading the bibliography…
Localizing moments in untrimmed videos via language queries is a new and interesting task that requires the ability to accurately ground language into video.
Recognizing realistic actions from videos “in the wild”
Jingen Liu, Jiebo Luo, and Mubarak Shah · 2009
Earlier work this paper cites.
Towards surveillance video search by natural language query
Stefanie Tellex and Deb Roy · 2009
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
C3d: generic features for video analysis
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2014
Earlier work this paper cites.
Weakly-supervised alignment of video with text
Piotr Bojanowski, Rémi Lajugie, Edouard Grave, Francis Bach, Ivan Laptev, Jean Ponce, and Cordelia Schmid · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
Ozan Sener, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Gated-attention readers for text comprehension
Bhuwan Dhingra, Hanxiao Liu, Zhilin Yang, William W Cohen, and Ruslan Salakhutdinov · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Cited alongside, same era.
Asynchronous temporal fields for action recognition
Gunnar A Sigurdsson, Santosh Divvala, Ali Farhadi, and Abhinav Gupta · 2016
Cited alongside, same era.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Cited alongside, same era.
Temporal action localization with pyramid of score distribution features
Jun Yuan, Bingbing Ni, Xiaokang Yang, and Ashraf A Kassim · 2016
Cited alongside, same era.
Yunlong Bian, Chuang Gan, Xiao Liu, Fu Li, Xiang Long, Yandong Li, Heng Qi, Jie Zhou, Shilei Wen, and Yuanqing Lin · 2017
Cited alongside, same era.
Rethinking the faster r-cnn architecture for temporal action localization
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar · 2018
Later among the works it cites.
Gated-attention architectures for task-oriented language grounding
Devendra Singh Chaplot, Kanthashree Mysore Sathyendra, Rama Kumar Pasumarthi, Dheeraj Rajagopal, and Ruslan Salakhutdinov · 2018
Later among the works it cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Later among the works it cites.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Later among the works it cites.
Iqa: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi · 2018
Later among the works it cites.
Learning to localize and align fine-grained actions to sparse instructions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sst: Single-stream temporal action proposals
Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Temporal context network for activity localization in videos
Xiyang Dai, Bharat Singh, Guyue Zhang, Larry S Davis, and Yan Qiu Chen · 2017
Cited alongside, same era.
Scc: Semantic context cascade for efficient action detection
Fabian Caba Heilbron, Wayner Barrios, Victor Escorcia, and Bernard Ghanem · 2017
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Cited alongside, same era.
Meera Hahn, Nataniel Ruiz, Jean-Baptiste Alayrac, Ivan Laptev, and James M Rehg · 2018
Later among the works it cites.
Attentive moment retrieval in videos
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua · 2018
Later among the works it cites.
Attend and interact: Higher-order object interactions for video understanding
Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan AlRegib, and Hans Peter Graf · 2018
Later among the works it cites.
Val: Visual-attention action localizer
Xiaomeng Song and Yahong Han · 2018
Later among the works it cites.
Text-to-clip video retrieval with early fusion and re-captioning
Huijuan Xu, Kun He, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2018
Later among the works it cites.
Yitian Yuan, Tao Mei, and Wenwu Zhu · 2018
Later among the works it cites.
Mac: Mining activity concepts for language-based temporal localization
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia · 2019
Closest in time.
Excl: Extractive clip localization using natural language descriptions
Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexander Hauptmann · 2019
Closest in time.
Dongliang He, Xiang Zhao, Jizhou Huang, Fu Li, Xiao Liu, and Shilei Wen · 2019
Closest in time.
Language-driven temporal activity localization: A semantic matching reinforcement learning model
Weining Wang, Yan Huang, and Liang Wang · 2019
Closest in time.
Proposal-free temporal moment localization of a natural-language query in video using guided attention
Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li, and Stephen Gould · 2020
Closest in time.