Fetching the paper…
Reading the bibliography…
Action recognition models have achieved impressive results by incorporating scene-level annotations, such as objects, their relations, 3D structure, and more.
Multitask learning
Rich Caruana · 1998
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E. Hinton, Nitish Srivastava, A. Krizhevsky, Ilya Sutskever, and R. Salakhutdinov · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Hierarchical recurrent neural network for skeleton based action recognition
Yong Du, Wei Wang, and Liang Wang · 2015
Earlier work this paper cites.
Smpl: a skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
Learning deep object detectors from 3d models
Xingchao Peng, Baochen Sun, Karim Ali, and Kate Saenko · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
Germán Ros, Laura Sellart, Joanna Materzynska, David Vázquez, and Antonio M. López · 2016
Earlier work this paper cites.
Ntu rgb+d: A large scale dataset for 3d human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Earlier work this paper cites.
Joint 2d-3d-semantic data for indoor scene understanding
Iro Armeni, Sasha Sax, Amir Roshan Zamir, and Silvio Savarese · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Microsoft kinect v2 vision system in a manufacturing application
L. Caruso, R. Russo, and S. Savino · 2017
Earlier work this paper cites.
Procedural generation of videos to train deep action recognition networks
CR De Souza, A Gaidon, Y Cabon, and AM Lopez Pena · 2017
Earlier work this paper cites.
Procedural generation of videos to train deep action recognition networks
César Roberto de Souza, Adrien Gaidon, Yohann Cabon, and Antonio Manuel López Peña · 2017
Earlier work this paper cites.
Actionvlad: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Kumar Gupta, Josef Sivic, and Bryan C. Russell · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Action tubelet detector for spatio-temporal action localization
Vicky S. Kalogeiton, Philippe Weinzaepfel, Vittorio Ferrari, and Cordelia Schmid · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Enhanced skeleton visualization for view invariant human action recognition
Mengyuan Liu, Hong Liu, and Chen Chen · 2017
Earlier work this paper cites.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Earlier work this paper cites.
Object level visual reasoning in videos
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, and Greg Mori · 2018
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al · 2018
Earlier work this paper cites.
Potion: Pose motion representation for action recognition
Vasileios Choutas, Philippe Weinzaepfel, Jérôme Revaud, and Cordelia Schmid · 2018
Earlier work this paper cites.
AVA: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik · 2018
Earlier work this paper cites.
Mapping images to scene graphs with permutation-invariant structured prediction
Roei Herzig, Moshiko Raboh, Gal Chechik, Jonathan Berant, and Amir Globerson · 2018
Earlier work this paper cites.
Image generation from scene graphs
Justin Johnson, Agrim Gupta, and Li Fei-Fei · 2018
Earlier work this paper cites.
Learning 3d human dynamics from video, 2018
Angjoo Kanazawa, Jason Y. Zhang, Panna Felsen, and Jitendra Malik · 2018
Earlier work this paper cites.
Compositional learning for human object interaction
Keizo Kato, Yin Li, and Abhinav Gupta · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Earlier work this paper cites.
Referring relationships
Ranjay Krishna, Ines Chami, Michael S. Bernstein, and Li Fei-Fei · 2018
Earlier work this paper cites.
Resound: Towards action recognition without representation bias
Yingwei Li, Yi Li, and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding, 2018
Ji Lin, Chuang Gan, and Song Han · 2018
Earlier work this paper cites.
Actor-centric relation network
Chen Sun, Abhinav Shrivastava, Carl Vondrick, Kevin Murphy, Rahul Sukthankar, and Cordelia Schmid · 2018
Earlier work this paper cites.
Videos as space-time region graphs
Xiaolong Wang and Abhinav Gupta · 2018
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents
F. Xia, Amir Roshan Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, and Silvio Savarese · 2018
Earlier work this paper cites.
Relational deep reinforcement learning
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, et al · 2018
Earlier work this paper cites.
Taskonomy: Disentangling task transfer learning
Amir Roshan Zamir, Alexander Sax, Bokui (William) Shen, Leonidas J. Guibas, Jitendra Malik, and Silvio Savarese · 2018
Earlier work this paper cites.
Learning semantic segmentation from synthetic data: A geometrically guided input-output adaptation approach
Yuhua Chen, Wen Li, Xiaoran Chen, and Luc Van Gool · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2019
Earlier work this paper cites.
Nddr-cnn: Layerwise feature fusing in multi-task cnns by neural discriminative dimensionality reduction
Yuan Gao, Qi She, Jiayi Ma, Mingbo Zhao, W. Liu, and Alan Loddon Yuille · 2019
Earlier work this paper cites.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Cited alongside, same era.
Spatio-temporal action graph networks
Roei Herzig, Elad Levi, Huijuan Xu, Hang Gao, Eli Brosh, Xiaolong Wang, Amir Globerson, and Trevor Darrell · 2019
Cited alongside, same era.
Action genome: Actions as composition of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2019
Cited alongside, same era.
A large-scale varying-view rgb-d action dataset for arbitrary-view human action recognition, 2019
Yanli Ji, Feixiang Xu, Yang Yang, Fumin Shen, Heng Tao Shen, and Wei-Shi Zheng · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
4dcontrast: Contrastive learning with dynamic correspondences for 3d scene understanding
Yujin Chen, Matthias Nießner, and Angela Dai · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Threedworld: A platform for interactive multi-modal physical simulation
Chuang Gan, Jeremy Schwartz, Seth Alter, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, Kuno Kim, Elias Wang, Damian Mrowca, Michael Lingelbach, Aidan Curtis, Kevin T. Feigelis, Daniel Bear, Dan Gutfreund, David Cox, James J. DiCarlo, Josh H. McDermott, Joshua B. Tenenbaum, and Daniel L. K. Yamins · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peeking into the future: Predicting future person activities and locations in videos
Junwei Liang, Lu Jiang, Juan Carlos Niebles, Alexander G. Hauptmann, and Li Fei-Fei · 2019
Cited alongside, same era.
Bmn: Boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Cited alongside, same era.
End-to-end multi-task learning with attention
Shikun Liu, Edward Johns, and Andrew J. Davison · 2019
Cited alongside, same era.
Structured domain randomization: Bridging the reality gap by context-aware synthetic data
Aayush Prakash, Shaad Boochoon, Mark Brophy, David Acuna, Eric Cameracci, Gavriel State, Omer Shapira, and Stan Birchfield · 2019
Cited alongside, same era.
Generalized intersection over union
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Cited alongside, same era.
Habitat: A platform for embodied ai research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
Video visual relation detection via multi-modal feature fusion
Xu Sun, Tongwei Ren, Yuan Zi, and Gangshan Wu · 2019
Cited alongside, same era.
Later among the works it cites.
Unit: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
Conflict-averse gradient descent for multi-task learning
Bo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone, and Qiang Liu · 2021
Later among the works it cites.
Towards impartial multi-task learning
Liyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang · 2021
Later among the works it cites.
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2021
Later among the works it cites.
A scaling law for synthetic-to-real transfer: How much is your pre-training effective?, 2021
Hiroaki Mikami, Kenji Fukumizu, Shogo Murai, Shuji Suzuki, Yuta Kikuchi, Taiji Suzuki, Shin ichi Maeda, and Kohei Hayashi · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers, 2021
Mandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Joao F. Henriques · 2021
Later among the works it cites.
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer · 2021
Later among the works it cites.
Vip-deeplab: Learning visual perception with depth-aware video panoptic segmentation
Siyuan Qiao, Yukun Zhu, Hartwig Adam, Alan Loddon Yuille, and Liang-Chieh Chen · 2021
Later among the works it cites.
Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding
Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M. Susskind · 2021
Later among the works it cites.
Task switching network for multi-task learning
Guolei Sun, Thomas Probst, Danda Pani Paudel, Nikola Popovic, Menelaos Kanakis, Jagruti R. Patel, Dengxin Dai, and Luc Van Gool · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Synthetic humans for action recognition from unseen viewpoints
Gül Varol, Ivan Laptev, Cordelia Schmid, and Andrew Zisserman · 2021
Later among the works it cites.
Synthetic humans for action recognition from unseen viewpoints
Gül Varol, Ivan Laptev, Cordelia Schmid, and Andrew Zisserman · 2021
Later among the works it cites.
Towards Long-Form Video Understanding
Chao-Yuan Wu and Philipp Krähenbühl · 2021
Later among the works it cites.
Temporal query networks for fine-grained video understanding
Chuhan Zhang, Ankush Gputa, and Andrew Zisserman · 2021
Later among the works it cites.
Attentional mixtures of soft prompt tuning for parameter-efficient multi-task knowledge sharing
Akari Asai, Mohammadreza Salehi, Matthew E. Peters, and Hannaneh Hajishirzi · 2022
Closest in time.
Bringing image scene structure to video via frame-clip consistency of object tokens
Elad Ben Avraham, Roei Herzig, Karttikeya Mangalam, Amir Bar, Anna Rohrbach, Leonid Karlinsky, Trevor Darrell, and Amir Globerson · 2022
Closest in time.
Muit: An end-to-end multitask learning transformer
Deblina Bhattacharjee, Tong Zhang, Sabine Süsstrunk, and Mathieu Salzmann · 2022
Closest in time.
Object-region video transformers
Roei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar, Gal Chechik, Anna Rohrbach, Trevor Darrell, and Amir Globerson · 2022
Closest in time.
Visual prompt tuning
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim · 2022
Closest in time.
Egocentric human-object interaction detection exploiting synthetic data, 2022
Rosario Leonardi, Francesco Ragusa, Antonino Furnari, and Giovanni Maria Farinella · 2022
Closest in time.
Uniformer: Unified transformer for efficient spatiotemporal representation learning, 2022
Kunchang Li, Yali Wang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao · 2022
Closest in time.
Mvitv2: Improved multiscale vision transformers for classification and detection
Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Closest in time.
Egocentric video-language pretraining
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Closest in time.
Task2sim: Towards effective pre-training and transfer from synthetic data
Samarth Mishra, Rameswar Panda, Cheng Perng Phoo, Chun-Fu Chen, Leonid Karlinsky, Kate Saenko, Venkatesh Saligrama, and Rogério Schmidt Feris · 2022
Closest in time.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training, 2022
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Closest in time.
SPoT: Better frozen model adaptation through soft prompt transfer
Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou’, and Daniel Cer · 2022
Closest in time.
Dualprompt: Complementary prompting for rehearsal-free continual learning
Zifeng Wang, Zizhao Zhang, Sayna Ebrahimi, Ruoxi Sun, Han Zhang, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, et al · 2022
Closest in time.
Learning to prompt for continual learning
Zifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang, Ruoxi Sun, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, and Tomas Pfister · 2022
Closest in time.
How transferable are video representations based on synthetic data?, 2022
Yo whan Kim, SouYoung Jin, Rameswar Panda, Hilde Kuehne, Leonid Karlinsky, Samarth Mishra, Venkatesh Saligrama, Kate Saenko, Aude Oliva, and Rogério Schmidt Feris · 2022
Closest in time.
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Closest in time.
Multi-task learning with multi-query transformer for dense prediction
Yangyang Xu, Xiangtai Li, Haobo Yuan, Yibo Yang, Jing Zhang, Yunhai Tong, Lefei Zhang, and Dacheng Tao · 2022
Closest in time.
Tuber: Tubelet transformer for video action detection
Jiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen, Shuai Bing, Mingze Xu, Chunhui Liu, Kaustav Kundu, Yuanjun Xiong, Davide Modolo, Ivan Marsic, Cees G. M. Snoek, and Joseph Tighe · 2022
Closest in time.
Conditional prompt learning for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Closest in time.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Closest in time.