Fetching the paper…
Reading the bibliography…
Algorithms for the action segmentation task typically use temporal models to predict what action is occurring at each frame for a minute-long daily activity.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Deformable DETR: deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2010
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Deformable DETR: deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2010
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
Alireza Fathi, Xiaofeng Ren, and James M. Rehg · 2011
Earlier work this paper cites.
Human action segmentation and recognition using discriminative semi-Markov models
Qinfeng Shi, Li Cheng, Li Wang, and Alex Smola · 2011
Earlier work this paper cites.
End-to-end video instance segmentation with transformers
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia · 2011
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
Alireza Fathi, Xiaofeng Ren, and James M. Rehg · 2011
Earlier work this paper cites.
Human action segmentation and recognition using discriminative semi-Markov models
Qinfeng Shi, Li Cheng, Li Wang, and Alex Smola · 2011
Earlier work this paper cites.
End-to-end video instance segmentation with transformers
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia · 2011
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele · 2012
Earlier work this paper cites.
Learning latent temporal structure for complex event detection
Kevin Tang, Li Fei-Fei, and Daphne Koller · 2012
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele · 2012
Earlier work this paper cites.
Learning latent temporal structure for complex event detection
Kevin Tang, Li Fei-Fei, and Daphne Koller · 2012
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2012
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Sebastian Stein and Stephen J. McKenna · 2013
Earlier work this paper cites.
Surgical gesture segmentation and recognition
Lingling Tao, Luca Zappella, Gregory Hager, and René Vidal · 2013
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Sebastian Stein and Stephen J. McKenna · 2013
Earlier work this paper cites.
Surgical gesture segmentation and recognition
Lingling Tao, Luca Zappella, Gregory Hager, and René Vidal · 2013
Earlier work this paper cites.
Fast saliency based pooling of fisher encoded dense trajectories
Svebor Karaman, Lorenzo Seidenari, and Alberto Del Bimbo · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Parsing videos of actions with segmental grammars
Hamed Pirsiavash and Deva Ramanan · 2014
Earlier work this paper cites.
Fast saliency based pooling of fisher encoded dense trajectories
Svebor Karaman, Lorenzo Seidenari, and Alberto Del Bimbo · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Parsing videos of actions with segmental grammars
Hamed Pirsiavash and Deva Ramanan · 2014
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
An end-to-end generative framework for video segmentation and recognition
Hilde Kuehne, Juergen Gall, and Thomas Serre · 2016
Earlier work this paper cites.
Segmental spatiotemporal cnns for fine-grained action segmentation
Colin Lea, Austin Reiter, René Vidal, and Gregory D. Hager · 2016
Earlier work this paper cites.
SSD: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg · 2016
Earlier work this paper cites.
Temporal action detection using a statistical language model
Alexander Richard and Juergen Gall · 2016
Cited alongside, same era.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Bharat Singh, Tim K. Marks, Michael Jones, Oncel Tuzel, and Ming Shao · 2016
Cited alongside, same era.
An end-to-end generative framework for video segmentation and recognition
Hilde Kuehne, Juergen Gall, and Thomas Serre · 2016
Cited alongside, same era.
Segmental spatiotemporal cnns for fine-grained action segmentation
Colin Lea, Austin Reiter, René Vidal, and Gregory D. Hager · 2016
Cited alongside, same era.
SSD: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg · 2016
Cited alongside, same era.
Temporal action detection using a statistical language model
Alexander Richard and Juergen Gall · 2016
Timeception for complex action recognition
Noureldien Hussein, Efstratios Gavves, and Arnold W.M. Smeulders · 2019
Later among the works it cites.
Self-attention temporal convolutional network for long-term daily living activity detection
Rui Dai, Luca Minciullo, Lorenzo Garattoni, Gianpiero Francesca, and François Bremond · 2019
Later among the works it cites.
MS-TCN: multi-stage temporal convolutional network for action segmentation
Yazan Abu Farha and Jurgen Gall · 2019
Later among the works it cites.
Timeception for complex action recognition
Noureldien Hussein, Efstratios Gavves, and Arnold W.M. Smeulders · 2019
Later among the works it cites.
Action segmentation with mixed temporal domain adaptation
Min-Hung Chen, Baopu Li, Yingze Bao, and Ghassan AlRegib · 2020
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Bharat Singh, Tim K. Marks, Michael Jones, Oncel Tuzel, and Ming Shao · 2016
Cited alongside, same era.
Quo Vadis, action recognition? a new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Trevor Darrell · 2017
Cited alongside, same era.
Weakly supervised learning of actions from transcripts
Hilde Kuehne, Alexander Richard, and Juergen Gall · 2017
Cited alongside, same era.
Temporal convolutional networks for action segmentation and detection
Colin Lea, Michael D. Flynn, René Vidal, Austin Reiter, and Gregory D. Hager · 2017
Cited alongside, same era.
Single shot temporal action detection
Tianwei Lin, Xu Zhao, and Zheng Shou · 2017
Cited alongside, same era.
Later among the works it cites.
Fine-grained action segmentation using the semi-supervised action GAN
Harshala Gammulle, Simon Denman, Sridha Sridharan, and Clinton Fookes · 2020
Later among the works it cites.
Improving action segmentation via graph-based temporal reasoning
Yifei Huang, Yusuke Sugano, and Yoichi Sato · 2020
Later among the works it cites.
Alleviating over-segmentation errors by detecting action boundaries
Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki, and Hirokatsu Kataoka · 2020
Later among the works it cites.
Deep concept-wise temporal convolutional networks for action localization
Xin Li, Tianwei Lin, Xiao Liu, Wangmeng Zuo, Chao Li, Xiang Long, Dongliang He, Fu Li, Shilei Wen, and Chuang Gan · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
Action segmentation with mixed temporal domain adaptation
Min-Hung Chen, Baopu Li, Yingze Bao, and Ghassan AlRegib · 2020
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2020
Later among the works it cites.
Fine-grained action segmentation using the semi-supervised action GAN
Harshala Gammulle, Simon Denman, Sridha Sridharan, and Clinton Fookes · 2020
Later among the works it cites.
Improving action segmentation via graph-based temporal reasoning
Yifei Huang, Yusuke Sugano, and Yoichi Sato · 2020
Later among the works it cites.
Alleviating over-segmentation errors by detecting action boundaries
Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki, and Hirokatsu Kataoka · 2020
Later among the works it cites.
Deep concept-wise temporal convolutional networks for action localization
Xin Li, Tianwei Lin, Xiao Liu, Wangmeng Zuo, Chao Li, Xiang Long, Dongliang He, Fu Li, Shilei Wen, and Chuang Gan · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chun-Fu Chen, Quanfu Fan, and Rameswar Panda · 2021
Closest in time.
Global2local: Efficient structure search for video action segmentation
Shang-Hua Gao1, Qi Han, Zhong-Yu Li, Pai Peng, Liang Wang, and Ming-Ming Cheng · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Coarse to fine multi-resolution temporal convolutional network
Dipika Singhania, Rahul Rahaman, and Angela Yao · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan · 2021
Closest in time.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H.S. Torr, and Li Zhang · 2021
Closest in time.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chun-Fu Chen, Quanfu Fan, and Rameswar Panda · 2021
Closest in time.
Global2local: Efficient structure search for video action segmentation
Shang-Hua Gao1, Qi Han, Zhong-Yu Li, Pai Peng, Liang Wang, and Ming-Ming Cheng · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Coarse to fine multi-resolution temporal convolutional network
Dipika Singhania, Rahul Rahaman, and Angela Yao · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan · 2021
Closest in time.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H.S. Torr, and Li Zhang · 2021
Closest in time.