Fetching the paper…
Reading the bibliography…
In static monitoring cameras, useful contextual information can stretch far beyond the few seconds typical video understanding models might see: subjects may exhibit similar behavior over multiple days, and background objects remain static.
Privacy preserving crowd monitoring: Counting people without people models or tracking
Antoni B Chan, Zhang-Sheng John Liang, and Nuno Vasconcelos · 2008
Earlier work this paper cites.
Automated identification of animal species in camera trap images
Xiaoyuan Yu, Jiangping Wang, Roland Kays, Patrick A Jansen, Tianjiang Wang, and Thomas Huang · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Finding action tubes
Georgia Gkioxari and Jitendra Malik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Snapshot serengeti, high-frequency annotated camera trap images of 40 mammalian species in an african savanna
Alexandra Swanson, Margaret Kosmala, Chris Lintott, Robert Simpson, Arfon Smith, and Craig Packer · 2015
Earlier work this paper cites.
Counting in the wild
Carlos Arteta, Victor Lempitsky, and Andrew Zisserman · 2016
Earlier work this paper cites.
R-fcn: Object detection via region-based fully convolutional networks
Jifeng Dai, Yi Li, Kaiming He, and Jian Sun · 2016
Earlier work this paper cites.
Seq-nms for video object detection
Wei Han, Pooya Khorrami, Tom Le Paine, Prajit Ramachandran, Mohammad Babaeizadeh, Honghui Shi, Jianan Li, Shuicheng Yan, and Thomas S Huang · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Ssd: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg · 2016
Earlier work this paper cites.
Finding areas of motion in camera trap images
Agnieszka Miguel, Sara Beery, Erica Flores, Loren Klemesrud, and Rana Bayrakcismith · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Animal detection from highly cluttered natural scenes using spatiotemporal object region proposals and patch verification
Zhi Zhang, Zhihai He, Guitao Cao, and Wenming Cao · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Detect to track and track to detect
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Speed/accuracy trade-offs for modern convolutional object detectors
Jonathan Huang, Vivek Rathod, Chen Sun, Menglong Zhu, Anoop Korattikara, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song, Sergio Guadarrama, et al · 2017
Earlier work this paper cites.
Object detection in videos with tubelet proposal networks
Kai Kang, Hongsheng Li, Tong Xiao, Wanli Ouyang, Junjie Yan, Xihui Liu, and Xiaogang Wang · 2017
Cited alongside, same era.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Cited alongside, same era.
Yolo9000: better, faster, stronger
Joseph Redmon and Ali Farhadi · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Towards automatic wild animal monitoring: Identification of animal species in camera-trap images using very deep convolutional neural networks
Alexander Gomez Villa, Augusto Salazar, and Francisco Vargas · 2017
Cited alongside, same era.
Spatiotemporal modeling for crowd counting in videos
Feng Xiong, Xingjian Shi, and Dit-Yan Yeung · 2017
Cadp: A novel dataset for cctv traffic camera based accident analysis
Ankit Parag Shah, Jean-Bapstite Lamare, Tuan Nguyen-Anh, and Alexander Hauptmann · 2018
Later among the works it cites.
The inaturalist species classification and detection dataset
Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie · 2018
Later among the works it cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Later among the works it cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Later among the works it cites.
Adversarial multiple source domain adaptation
Han Zhao, Shanghang Zhang, Guanhang Wu, José MF Moura, Joao P Costeira, and Geoffrey J Gordon · 2018
Later among the works it cites.
Towards high performance video object detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast human-animal detection from highly cluttered camera-trap images using joint background modeling and deep learning classification
Hayder Yousif, Jianhe Yuan, Roland Kays, and Zhihai He · 2017
Cited alongside, same era.
Fcn-rlstm: Deep spatio-temporal neural networks for vehicle counting in city cameras
Shanghang Zhang, Guanhang Wu, Joao P Costeira, and José MF Moura · 2017
Cited alongside, same era.
Understanding traffic density from large-scale web camera data
Shanghang Zhang, Guanhang Wu, Joao P Costeira, and Jose MF Moura · 2017
Cited alongside, same era.
Flow-guided feature aggregation for video object detection
Xizhou Zhu, Yujie Wang, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Cited alongside, same era.
Deep feature flow for video recognition
Xizhou Zhu, Yuwen Xiong, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Cited alongside, same era.
Recognition in terra incognita
Sara Beery, Grant Van Horn, and Pietro Perona · 2018
Cited alongside, same era.
Xizhou Zhu, Jifeng Dai, Lu Yuan, and Yichen Wei · 2018
Later among the works it cites.
http://lila.science/
Lila.science · 2019
Closest in time.
Synthetic examples improve generalization for rare classes
Sara Beery, Yang Liu, Dan Morris, Jim Piavis, Ashish Kapoor, Markus Meister, and Pietro Perona · 2019
Closest in time.
Efficient pipeline for automating species id in new camera trap projects
Sara Beery and Dan Morris · 2019
Closest in time.
Object guided external memory network for video object detection
Hanming Deng, Yang Hua, Tao Song, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson, and Haibing Guan · 2019
Closest in time.
Scale mlperf-0.6 models on google tpu-v3 pods
Sameer Kumar, Victor Bitorff, Dehao Chen, Chiachen Chou, Blake Hechtman, HyoukJoong Lee, Naveen Kumar, Peter Mattson, Shibo Wang, Tao Wang, et al · 2019
Closest in time.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Closest in time.
Leveraging long-range temporal relationships between proposals for video object detection
Mykhailo Shvets, Wei Liu, and Alexander C Berg · 2019
Closest in time.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Closest in time.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Closest in time.
Fcos: Fully convolutional one-stage object detection
Zhi Tian, Chunhua Shen, Hao Chen, and Tong He · 2019
Closest in time.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Closest in time.
Sequence level semantics aggregation for video object detection
Haiping Wu, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang · 2019
Closest in time.
Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl · 2019
Closest in time.