Fetching the paper…
Reading the bibliography…
We propose a novel Siamese Natural Language Tracker (SNLT), which brings the advancements in visual tracking to the tracking by natural language (NL) descriptions task.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman et al · 1960
Earlier work this paper cites.
g–h and g–h–k filters
Eli Brookner · 1998
Earlier work this paper cites.
Learning to recognize objects
Linda B Smith · 2003
Earlier work this paper cites.
Multiple hypothesis tracking for multiple target tracking
Samuel S Blackman · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2011
Earlier work this paper cites.
Symbolic play connects to language through visual object recognition
Linda B Smith and Susan S Jones · 2011
Earlier work this paper cites.
Tracking-learning-detection
Zdenek Kalal, Krystian Mikolajczyk, and Jiri Matas · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Online object tracking: A benchmark
Yi Wu, Jongwoo Lim, and Ming-Hsuan Yang · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Cited alongside, same era.
Fully-convolutional siamese networks for object tracking
Luca Bertinetto, Jack Valmadre, João F Henriques, Andrea Vedaldi, and Philip HS Torr · 2016
Cited alongside, same era.
Beyond correlation filters: Learning continuous convolution operators for visual tracking
Martin Danelljan, Andreas Robinson, Fahad Shahbaz Khan, and Michael Felsberg · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Cited alongside, same era.
Vital: Visual tracking via adversarial learning
Yibing Song, Chao Ma, Xiaohe Wu, Lijun Gong, Linchao Bao, Wangmeng Zuo, Chunhua Shen, Rynson WH Lau, and Ming-Hsuan Yang · 2018
Later among the works it cites.
Distractor-aware siamese networks for visual object tracking
Zheng Zhu, Qiang Wang, Bo Li, Wei Wu, Junjie Yan, and Weiming Hu · 2018
Later among the works it cites.
Learning discriminative model prediction for tracking
Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte · 2019
Closest in time.
Language features matter: Effective language representations for vision-language tasks
Andrea Burns, Reuben Tan, Kate Saenko, Stan Sclaroff, and Bryan A Plummer · 2019
Closest in time.
Atom: Accurate tracking by overlap maximization
Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg · 2019
Closest in time.
Lasot: A high-quality benchmark for large-scale single object tracking
Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hyeonseob Nam and Bohyung Han · 2016
Cited alongside, same era.
Eco: Efficient convolution operators for tracking
Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg · 2017
Cited alongside, same era.
Learning policies for adaptive tracking with deep feature cascades
Chen Huang, Simon Lucey, and Deva Ramanan · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Tracking by natural language specification
Zhenyang Li, Ran Tao, Efstratios Gavves, Cees GM Snoek, and Arnold WM Smeulders · 2017
Cited alongside, same era.
Spatially supervised recurrent convolutional neural networks for visual object tracking
Guanghan Ning, Zhi Zhang, Chen Huang, Xiaobo Ren, Haohong Wang, Canhui Cai, and Zhihai He · 2017
Cited alongside, same era.
Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video
Esteban Real, Jonathon Shlens, Stefano Mazzocchi, Xin Pan, and Vincent Vanhoucke · 2017
Cited alongside, same era.
Closest in time.
Learning to compose and reason with language tree structures for visual grounding
Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He, and Hanwang Zhang · 2019
Closest in time.
Siamrpn++: Evolution of siamese visual tracking with very deep networks
Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan · 2019
Closest in time.
Tracking the known and the unknown by leveraging semantic information
Ardhendu Shekhar Tripathi, Martin Danelljan, Luc Van Gool, and Radu Timofte · 2019
Closest in time.
A fast and accurate one-stage approach to visual grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo · 2019
Closest in time.
Probabilistic regression for visual tracking
Martin Danelljan, Luc Van Gool, and Radu Timofte · 2020
Closest in time.
Real-time visual object tracking with natural language description
Qi Feng, Vitaly Ablavsky, Qinxun Bai, Guorong Li, and Stan Sclaroff · 2020
Closest in time.
Siam R-CNN: Visual tracking by re-detection
Paul Voigtlaender, Jonathon Luiten, Philip HS Torr, and Bastian Leibe · 2020
Closest in time.