Fetching the paper…
Reading the bibliography…
Vision transformers are emerging as a powerful tool to solve computer vision problems.
Recurrent neural networks
Larry R Medsker and LC Jain · 2001
Earlier work this paper cites.
Human activity recognition for video surveillance
Weiyao Lin, Ming-Ting Sun, Radha Poovandran, and Zhengyou Zhang · 2008
Earlier work this paper cites.
Software survey: Vosviewer, a computer program for bibliometric mapping
Nees Van Eck and Ludo Waltman · 2010
Earlier work this paper cites.
Real-time human pose recognition in parts from single depth images
Jamie Shotton, Andrew Fitzgibbon, Mat Cook, Toby Sharp, Mark Finocchio, Richard Moore, Alex Kipman, and Andrew Blake · 2011
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
Alireza Fathi, Xiaofeng Ren, and James M Rehg · 2011
Earlier work this paper cites.
A smart local moving algorithm for large-scale modularity-based community detection
Ludo Waltman and Nees Jan Van Eck · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Visualizing bibliometric networks
Nees Jan Van Eck and Ludo Waltman · 2014
Earlier work this paper cites.
Action recognition in realistic sports videos
Khurram Soomro and Amir R Zamir · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
A critical review of recurrent neural networks for sequence learning
Zachary C Lipton, John Berkowitz, and Charles Elkan · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
An introduction to convolutional neural networks
Keiron O’Shea and Ryan Nash · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Hierarchical attention networks for document classification
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Ntu rgb+ d: A large scale dataset for 3d human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Earlier work this paper cites.
Online action detection
Roeland De Geest, Efstratios Gavves, Amir Ghodrati, Zhenyang Li, Cees Snoek, and Tinne Tuytelaars · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Recognising complex activities with histograms of relative tracklets
Sebastian Stein and Stephen J McKenna · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al · 2018
Earlier work this paper cites.
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Earlier work this paper cites.
Self-attention generative adversarial networks
Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena · 2019
Earlier work this paper cites.
Understanding lstm–a tutorial into long short-term memory recurrent neural networks
Ralf C Staudemeyer and Eric Rothstein Morris · 2019
Earlier work this paper cites.
Learning video representations using contrastive bidirectional transformer
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Earlier work this paper cites.
Deep high-resolution representation learning for human pose estimation
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang · 2019
Earlier work this paper cites.
A short note on the kinetics-700 human action dataset
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman · 2019
Earlier work this paper cites.
Moments in time dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al · 2019
Earlier work this paper cites.
Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot · 2019
Earlier work this paper cites.
A survey of deep learning techniques for neural machine translation
Shuoheng Yang, Yuxin Wang, and Xiaowen Chu · 2020
Earlier work this paper cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Late temporal modeling in 3d cnn architectures with bert for action recognition
M Kalfaoglu, Sinan Kalkan, and A Aydin Alatan · 2020
Earlier work this paper cites.
Lightweight network architecture for real-time action recognition
Alexander Kozlov, Vadim Andronov, and Yana Gritsenko · 2020
Earlier work this paper cites.
Actor-transformers for group activity recognition
Kirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, and Cees GM Snoek · 2020
Earlier work this paper cites.
Driver action recognition using deformable and dilated faster r-cnn with optimized region proposals
Mingqi Lu, Yaocong Hu, and Xiaobo Lu · 2020
Cited alongside, same era.
Speech2action: Cross-modal supervision for action recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, and Andrew Zisserman · 2020
Cited alongside, same era.
Sct: Set constrained temporal transformer for set supervised action segmentation
Mohsen Fayyaz and Jurgen Gall · 2020
Cited alongside, same era.
Transformers in vision: A survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah · 2021
Cited alongside, same era.
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2021
Cited alongside, same era.
Uncertainty-guided probabilistic transformer for complex action recognition
Hongji Guo, Hanjing Wang, and Qiang Ji · 2022
Closest in time.
Graph transformer network with temporal kernel attention for skeleton-based action recognition
Yanan Liu, Hao Zhang, Dan Xu, and Kangjian He · 2022
Closest in time.
Shrinking temporal attention in transformers for video action recognition
Bonan Li, Pengfei Xiong, Congying Han, and Tiande Guo · 2022
Closest in time.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Closest in time.
An efficient spatio-temporal pyramid transformer for action detection
Yuetian Weng, Zizheng Pan, Mingfei Han, Xiaojun Chang, and Bohan Zhuang · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preksha Pareek and Ankit Thakkar · 2021
Cited alongside, same era.
Evaluating transformers for lightweight action recognition
Raivo Koot, Markus Hennerbichler, and Haiping Lu · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lu, and Cordelia Schmid · 2021
Cited alongside, same era.
Co-training transformer with videos and images improves action recognition
Bowen Zhang, Jiahui Yu, Christopher Fifty, Wei Han, Andrew M Dai, Ruoming Pang, and Fei Sha · 2021
Cited alongside, same era.
Spatial temporal transformer network for skeleton-based action recognition
Chiara Plizzari, Marco Cannici, and Matteo Matteucci · 2021
Cited alongside, same era.
Videolightformer: Lightweight action recognition using transformers
Raivo Koot and Haiping Lu · 2021
Cited alongside, same era.
Vittorio Mazzia, Simone Angarano, Francesco Salvetti, Federico Angelini, and Marcello Chiaberge · 2022
Closest in time.
Deformable video transformer
Jue Wang and Lorenzo Torresani · 2022
Closest in time.
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2022
Closest in time.
Recurring the transformer for video action recognition
Jiewen Yang, Xingbo Dong, Liujun Liu, Chao Zhang, Jiajun Shen, and Dahai Yu · 2022
Closest in time.
Direcformer: A directed attention in transformer approach to robust action recognition
Thanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo, Son Lam Phung, Xin Li, and Khoa Luu · 2022
Closest in time.
Graph transformer network with temporal kernel attention for skeleton-based action recognition
Yanan Liu, Hao Zhang, Dan Xu, and Kangjian He · 2022
Closest in time.
Mm-vit: Multi-modal video transformer for compressed video action recognition
Jiawei Chen and Chiu Man Ho · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi · 2022
Closest in time.
Audio-adaptive activity recognition across video domains
Yunhua Zhang, Hazel Doughty, Ling Shao, and Cees GM Snoek · 2022
Closest in time.
Everything at once-multi-modal fusion transformer for video retrieval
Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogerio S Feris, David Harwath, James Glass, and Hilde Kuehne · 2022
Closest in time.
Cross-modal contrastive distillation for instructional activity anticipation
Zhengyuan Yang, Jingen Liu, Jing Huang, Xiaodong He, Tao Mei, Chenliang Xu, and Jiebo Luo · 2022
Closest in time.
Language supervised training for skeleton-based action recognition
Wangmeng Xiang, Chao Li, Yuxuan Zhou, Biao Wang, and Lei Zhang · 2022
Closest in time.
M&m mix: A multimodal multiview transformer ensemble
Xuehan Xiong, Anurag Arnab, Arsha Nagrani, and Cordelia Schmid · 2022
Closest in time.
Multimodal transformer for nursing activity recognition
Momal Ijaz, Renato Diaz, and Chen Chen · 2022
Closest in time.
Entity-aware and motion-aware transformers for language-driven action localization in videos
Shuo Yang and Xinxiao Wu · 2022
Closest in time.
Unsupervised domain adaptation for video transformers in action recognition
Victor G Turrisi da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, and Elisa Ricci · 2022
Closest in time.
Transformers in action: Weakly supervised action segmentation
John Ridley, Huseyin Coskun, David Joseph Tan, Nassir Navab, and Federico Tombari · 2022
Closest in time.
Efficient u-transformer with boundary-aware loss for action segmentation
Dazhao Du, Bing Su, Yu Li, Zhongang Qi, Lingyu Si, and Ying Shan · 2022
Closest in time.
Cross-enhancement transformer for action segmentation
Jiahui Wang, Zhenyou Wang, Shanna Zhuang, and Hui Wang · 2022
Closest in time.
Tuber: Tubelet transformer for video action detection
Jiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen, Bing Shuai, Mingze Xu, Chunhui Liu, Kaustav Kundu, Yuanjun Xiong, Davide Modolo, et al · 2022
Closest in time.
Stargazer: A transformer-based driver action detection system for intelligent transportation
Junwei Liang, He Zhu, Enwei Zhang, and Jun Zhang · 2022
Closest in time.
Unified recurrence modeling for video action anticipation
Tsung-Ming Tai, Giuseppe Fiameni, Cheng-Kuang Lee, Simon See, and Oswald Lanz · 2022
Closest in time.
Future transformer for long-term action anticipation
Dayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha, and Minsu Cho · 2022
Closest in time.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Closest in time.
Masked feature prediction for self-supervised visual pre-training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer · 2022
Closest in time.
Nsnet: Non-saliency suppression sampler for efficient video recognition
Boyang Xia, Wenhao Wu, Haoran Wang, Rui Su, Dongliang He, Haosen Yang, Xiaoran Fan, and Wanli Ouyang · 2022
Closest in time.
Temporal saliency query network for efficient video recognition
Boyang Xia, Zhihao Wang, Wenhao Wu, Haoran Wang, and Jungong Han · 2022
Closest in time.
Revisiting skeleton-based action recognition
Haodong Duan, Yue Zhao, Kai Chen, Dahua Lin, and Bo Dai · 2022
Closest in time.
Gatehub: Gated history unit with background suppression for online action detection
Junwen Chen, Gaurav Mittal, Ye Yu, Yu Kong, and Mei Chen · 2022
Closest in time.
Bridge-prompt: Towards ordinal action understanding in instructional videos
Muheng Li, Lei Chen, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie Zhou, and Jiwen Lu · 2022
Closest in time.
Maximization and restoration: Action segmentation through dilation passing and temporal reconstruction
Junyong Park, Daekyum Kim, Sejoon Huh, and Sungho Jo · 2022
Closest in time.
Epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2022
Closest in time.
Omnivore: A single model for many visual modalities
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Closest in time.
Cross-modal representation learning for zero-shot action recognition
Chung-Ching Lin, Kevin Lin, Lijuan Wang, Zicheng Liu, and Linjie Li · 2022
Closest in time.