Fetching the paper…
Reading the bibliography…
Action classification has made great progress, but segmenting and recognizing actions from long untrimmed videos remains a challenging problem.
Low-rank random tensor for bilinear pooling
Zhang, Y.; Muandet, K.; Ma, Q.; Neumann, H.; and Tang, S. 2019 · 1906
Earlier work this paper cites.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Singh, B.; Marks, T. K.; Jones, M.; Tuzel, O.; and Shao, M. 2016 · 1970
Earlier work this paper cites.
Recurrent networks and NARMA modeling
Connor, J.; Atlas, L. E.; and Martin, D. R. 1992 · 1992
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
Fathi, A.; Ren, X.; and Rehg, J. M. 2011 · 2011
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Rohrbach, M.; Amin, S.; Andriluka, M.; and Schiele, B. 2012 · 2012
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Stein, S.; and McKenna, S. J. 2013 · 2013
Earlier work this paper cites.
Recognition of complex events: Exploiting temporal dynamics between underlying concepts
Bhattacharya, S.; Kalayeh, M. M.; Sukthankar, R.; and Shah, M. 2014 · 2014
Earlier work this paper cites.
Temporal sequence modeling for video event detection
Cheng, Y.; Fan, Q.; Pankanti, S.; and Choudhary, A. 2014 · 2014
Earlier work this paper cites.
Fast saliency based pooling of fisher encoded dense trajectories
Karaman, S.; Seidenari, L.; and Del Bimbo, A. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Kuehne, H.; Arslan, A.; and Serre, T. 2014 · 2014
Earlier work this paper cites.
From stochastic grammar to bayes network: Probabilistic parsing of complex activity
Vo, N. N.; and Bobick, A. F. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
An end-to-end generative framework for video segmentation and recognition
Kuehne, H.; Gall, J.; and Serre, T. 2016 · 2016
Earlier work this paper cites.
Segmental spatiotemporal cnns for fine-grained action segmentation
Lea, C.; Reiter, A.; Vidal, R.; and Hager, G. D. 2016 · 2016
Earlier work this paper cites.
Temporal action detection using a statistical language model
Richard, A.; and Gall, J. 2016 · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Ulyanov, D.; Vedaldi, A.; and Lempitsky, V. 2016 · 2016
Cited alongside, same era.
WaveNet: A generative model for raw audio
Van Den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A. W.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Cited alongside, same era.
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2020
Later among the works it cites.
Improving Action Segmentation via Graph-Based Temporal Reasoning
Huang, Y.; Sugano, Y.; and Sato, Y. 2020 · 2020
Later among the works it cites.
Rethinking Positional Encoding in Language Pre-training
Ke, G.; He, D.; and Liu, T.-Y. 2020 · 2020
Later among the works it cites.
MS-TCN++: Multi-Stage Temporal Convolutional Network for Action Segmentation
Li, S.-J.; AbuFarha, Y.; Liu, Y.; Cheng, M.-M.; and Gall, J. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weakly supervised learning of actions from transcripts
Kuehne, H.; Richard, A.; and Gall, J. 2017 · 2017
Cited alongside, same era.
Temporal convolutional networks for action segmentation and detection
Lea, C.; Flynn, M. D.; Vidal, R.; Reiter, A.; and Hager, G. D. 2017 · 2017
Cited alongside, same era.
Weakly supervised action learning with rnn based fine-to-coarse modeling
Richard, A.; Kuehne, H.; and Gall, J. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Bai, S.; Kolter, J. Z.; and Koltun, V. 2018 · 2018
Cited alongside, same era.
Temporal deformable residual networks for action segmentation in videos
Lei, P.; and Todorovic, S. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Shaw, P.; Uszkoreit, J.; and Vaswani, A. 2018 · 2018
Cited alongside, same era.
Wang, Z.; Gao, Z.; Wang, L.; Li, Z.; and Wu, G. 2020 · 2020
Later among the works it cites.
Swin-unet: Unet-like pure transformer for medical image segmentation
Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; and Wang, M. 2021 · 2021
Later among the works it cites.
Global2local: Efficient structure search for video action segmentation
Gao, S.-H.; Han, Q.; Li, Z.-Y.; Peng, P.; Wang, L.; and Cheng, M.-M. 2021 · 2021
Later among the works it cites.
Alleviating over-segmentation errors by detecting action boundaries
Ishikawa, Y.; Kasai, S.; Aoki, Y.; and Kataoka, H. 2021 · 2021
Later among the works it cites.
Multi-compound transformer for accurate biomedical image segmentation
Ji, Y.; Zhang, R.; Wang, H.; Li, Z.; Wu, L.; Zhang, S.; and Luo, P. 2021 · 2021
Later among the works it cites.
Coarse to fine multi-resolution temporal convolutional network
Singhania, D.; Rahaman, R.; and Yao, A. 2021 · 2021
Later among the works it cites.
End-to-end video instance segmentation with transformers
Wang, Y.; Xu, Z.; Wang, X.; Shen, C.; Cheng, B.; Shen, H.; and Xia, H. 2021 · 2021
Later among the works it cites.
SegFormer: Simple and efficient design for semantic segmentation with transformers
Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; and Luo, P. 2021 · 2021
Later among the works it cites.
Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy
Xu, J.; Wu, H.; Wang, J.; and Long, M. 2021 · 2021
Later among the works it cites.
Learning spatio-temporal transformer for visual tracking
Yan, B.; Peng, H.; Fu, J.; Wang, D.; and Lu, H. 2021 · 2021
Later among the works it cites.
ASFormer: Transformer for Action Segmentation
Yi, F.; Wen, H.; and Jiang, T. 2021 · 2021
Later among the works it cites.
Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; and Zhang, W. 2021 · 2021
Later among the works it cites.