Fetching the paper…
Reading the bibliography…
We present Long Short-term TRansformer (LSTR), a temporal modeling algorithm for online action detection, which employs a long- and short-term memory mechanism to model prolonged sequence data.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
Attention and memory: An integrated framework
Nelson Cowan · 1998
Earlier work this paper cites.
Visual working memory as visual attention sustained internally over time
Marvin M Chun · 2011
Earlier work this paper cites.
Working memory as internal attention: Toward an integrative account of internal and external selection processes
Anastasia Kiyonaga and Tobias Egner · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Max-margin early event detectors
Minh Hoai and Fernando De la Torre · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Online action detection
Roeland De Geest, Efstratios Gavves, Amir Ghodrati, Zhenyang Li, Cees Snoek, and Tinne Tuytelaars · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning activity progression in lstms for activity detection and early detection
Shugao Ma, Leonid Sigal, and Stan Sclaroff · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
SST: Single-stream temporal action proposals
Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the Kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Spatiotemporal multiplier networks for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Richard P Wildes · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
The THUMOS challenge on action recognition for videos “in the wild”
Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah · 2017
Earlier work this paper cites.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Cited alongside, same era.
CDC: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Cited alongside, same era.
BSN: Boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Cited alongside, same era.
Openmmlab’s next generation video understanding toolbox and benchmark
MMAction2 Contributors · 2020
Later among the works it cites.
Learning to discriminate information for online action detection
Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, and Changick Kim · 2020
Later among the works it cites.
LAP-Net: Adaptive features sampling via learning action progression for online action detection
Sanqing Qu, Guang Chen, Dan Xu, Jinhu Dong, Fan Lu, and Alois Knoll · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Online action detection in untrimmed, streaming videos-modeling and evaluation
Zheng Shou, Junting Pan, Jonathan Chan, Kazuyuki Miyazawa, Hassan Mansour, Anthony Vetro, Xavier Giro-i Nieto, and Shih-Fu Chang · 2018
Cited alongside, same era.
PWC-Net: Cnns for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Cited alongside, same era.
Non-local netvlad encoding for video classification
Yongyi Tang, Xing Zhang, Lin Ma, Jingwen Wang, Shaoxiang Chen, and Yu-Gang Jiang · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
denseflow
Shiguang Wang*, Zhizhong Li*, Yue Zhao, Yuanjun Xiong, Limin Wang, and Dahua Lin · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
Privileged knowledge distillation for online action detection
Peisen Zhao, Jiajie Wang, Lingxi Xie, Ya Zhang, Yanfeng Wang, and Qi Tian · 2020
Later among the works it cites.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Closest in time.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Closest in time.
Temporal filtering networks for online action detection
Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, and Changick Kim · 2021
Closest in time.
WOAD: Weakly supervised online action detection in untrimmed videos
Mingfei Gao, Yingbo Zhou, Ran Xu, Richard Socher, and Caiming Xiong · 2021
Closest in time.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Closest in time.
Temporally smooth online action detection using cycle-consistent future anticipation
Young Hwi Kim, Seonghyeon Nam, and Seon Joo Kim · 2021
Closest in time.
Vidtr: Video transformer without convolutions
Xinyu Li, Yanyi Zhang, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Closest in time.
Activity graph transformer for temporal action localization
Megha Nawhal and Greg Mori · 2021
Closest in time.
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.
An image is worth 16x16 words, what is a video worth?
Gilad Sharir, Asaf Noy, and Lihi Zelnik-Manor · 2021
Closest in time.
Relaxed transformer decoders for direct action proposal generation
Jing Tan, Jiaqi Tang, Limin Wang, and Gangshan Wu · 2021
Closest in time.
TTPP: Temporal transformer with progressive prediction for efficient action anticipation
Wen Wang, Xiaojiang Peng, Yanzhou Su, Yu Qiao, and Jian Cheng · 2021
Closest in time.
Towards long-form video understanding
Chao-Yuan Wu and Philipp Krähenbühl · 2021
Closest in time.