Fetching the paper…
Reading the bibliography…
Action understanding has attracted long-term attention.
Wordnet: a lexical database for english
George A Miller · 1995
Earlier work this paper cites.
The berkeley framenet project
Collin F Baker, Charles J Fillmore, and John B Lowe · 1998
Earlier work this paper cites.
From treebank to propbank
Paul R Kingsbury and Martha Palmer · 2002
Earlier work this paper cites.
VerbNet: A broad-coverage, comprehensive verb lexicon
Karin Kipper Schuler · 2005
Earlier work this paper cites.
Recognizing realistic actions from videos “in the wild”
Jingen Liu, Jiebo Luo, and Mubarak Shah · 2009
Earlier work this paper cites.
Recognizing human actions in still images: a study of bag-of-features and part-based representations
V. Delaitre, I. Laptev, and J. Sivic · 2010
Earlier work this paper cites.
Modeling temporal structure of decomposable motion segments for activity classification
Juan Carlos Niebles, Chih-Wei Chen, and Li Fei-Fei · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Recognition using visual phrases
Mohammad Amin Sadeghi and Ali Farhadi · 2011
Earlier work this paper cites.
The action similarity labeling challenge
O. Kliper-Gross, T. Hassner, and L. Wolf · 2012
Earlier work this paper cites.
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu · 2013
Earlier work this paper cites.
From actemes to action: A strongly-supervised representation for detailed action understanding
Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis · 2013
Earlier work this paper cites.
2d human pose estimation: New benchmark and state of the art analysis
Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele · 2014
Earlier work this paper cites.
Tuhoi: Trento universal human object interaction dataset
Dieu-Thu Le, Jasper Uijlings, and Raffaella Bernardi · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Yu Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng · 2015
Earlier work this paper cites.
Utd-mhad: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor
Chen Chen, Roozbeh Jafari, and Nasser Kehtarnavaz · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Contextual action recognition with r* cnn
Georgia Gkioxari, Ross Girshick, and Jitendra Malik · 2015
Earlier work this paper cites.
Saurabh Gupta and Jitendra Malik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Unsupervised visual sense disambiguation for verbs using multimodal embeddings
Spandana Gella, Mirella Lapata, and Frank Keller · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning models for actions and person-object interactions with transfer to question answering
Arun Mallya and Svetlana Lazebnik · 2016
Earlier work this paper cites.
Pytextrank, a python implementation of textrank for phrase extraction and summarization of text documents
Paco Nathan · 2016
Earlier work this paper cites.
The KIT motion-language dataset
Matthias Plappert, Christian Mandery, and Tamim Asfour · 2016
Earlier work this paper cites.
Ntu rgb+ d: A large scale dataset for 3d human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Andrew Zisserman Joao Carreira · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
A simple yet effective baseline for 3d human pose estimation
Julieta Martinez, Rayat Hossain, Javier Romero, and James J Little · 2017
Earlier work this paper cites.
Poincaré embeddings for learning hierarchical representations
Maximillian Nickel and Douwe Kiela · 2017
Cited alongside, same era.
Coarse-to-fine volumetric prediction for single-image 3d human pose
Georgios Pavlakos, Xiaowei Zhou, Konstantinos G Derpanis, and Kostas Daniilidis · 2017
Cited alongside, same era.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Cited alongside, same era.
Compositional human pose regression
Xiao Sun, Jiaxiang Shang, Shuang Liang, and Yichen Wei · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng · 2018
Searching for actions on the hyperbole
Teng Long, Pascal Mettes, Heng Tao Shen, and Cees GM Snoek · 2020
Later among the works it cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Later among the works it cites.
Finegym: A hierarchical video dataset for fine-grained action understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin · 2020
Later among the works it cites.
Temporal pyramid network for action recognition
Ceyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai, and Bolei Zhou · 2020
Later among the works it cites.
Elaborative rehearsal for zero-shot action recognition
Shizhe Chen and Dong Huang · 2021
Later among the works it cites.
Haa500: Human-centric atomic action dataset with curated videos
Jihoon Chung, Cheng hsin Wuu, Hsuan ru Yang, Yu-Wing Tai, and Chi-Keung Tang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Potion: Pose motion representation for action recognition
Vasileios Choutas, Philippe Weinzaepfel, Jérôme Revaud, and Cordelia Schmid · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Automatic emotion and attention analysis of young children at home: a researchkit autism feasibility study
Helen L Egger, Geraldine Dawson, Jordan Hashemi, Kimberly LH Carpenter, Steven Espinosa, Kathleen Campbell, Samuel Brotkin, Jana Schaich-Borg, Qiang Qiu, Mariano Tepper, et al · 2018
Cited alongside, same era.
Pairwise body-part attention for recognizing human-object interactions
Hao Shu Fang, Jinkun Cao, Yu Wing Tai, and Cewu Lu · 2018
Cited alongside, same era.
Hyperbolic entailment cones for learning hierarchical embeddings
Octavian Ganea, Gary Bécigneul, and Thomas Hofmann · 2018
Cited alongside, same era.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al · 2018
Cited alongside, same era.
Later among the works it cites.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Later among the works it cites.
Partial success in closing the gap between human and machine vision
Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A Wichmann, and Wieland Brendel · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Exemplar fine-tuning for 3d human model fitting towards in-the-wild 3d human pose estimation
Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi · 2021
Later among the works it cites.
Action-conditioned 3D human motion synthesis with transformer VAE
Mathis Petrovich, Michael J. Black, and Gül Varol · 2021
Later among the works it cites.
BABEL: Bodies, action and behavior with english labels
Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J. Black · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Home action genome: Cooperative compositional action understanding
Nishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai, Kazuki Kozuka, Shun Ishizaka, Ehsan Adeli, and Juan Carlos Niebles · 2021
Later among the works it cites.
Monocular, one-stage, regression of multiple 3d people
Yu Sun, Qian Bao, Wu Liu, Yili Fu, Black Michael J., and Tao Mei · 2021
Later among the works it cites.
Learning the predictability of the future
Dídac Surís, Ruoshi Liu, and Carl Vondrick · 2021
Later among the works it cites.
QPIC: Query-based pairwise human-object interaction detection with image-wide contextual information
Masato Tamura, Hiroki Ohashi, and Tomoaki Yoshinaga · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
Positive unlabeled contrastive learning
Anish Acharya, Sujay Sanghavi, Li Jing, Bhargav Bhushanam, Dhruv Choudhary, Michael Rabbat, and Inderjit Dhillon · 2022
Later among the works it cites.
Hake: A knowledge engine foundation for human activity understanding
Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li, Zuoyu Qiu, Liang Xu, Yue Xu, Hao-Shu Fang, and Cewu Lu · 2022
Later among the works it cites.
Frozen clip models are efficient video learners
Ziyi Lin, Shijie Geng, Renrui Zhang, Peng Gao, Gerard de Melo, Xiaogang Wang, Jifeng Dai, Yu Qiao, and Hongsheng Li · 2022
Later among the works it cites.
Posegpt: Quantization-based 3d human motion generation and forecasting
Thomas Lucas*, Fabien Baradel*, Philippe Weinzaepfel, and Grégory Rogez · 2022
Later among the works it cites.
Relvit: Concept-guided vision transformer for visual relational reasoning
Xiaojian Ma, Weili Nie, Zhiding Yu, Huaizu Jiang, Chaowei Xiao, Yuke Zhu, Song-Chun Zhu, and Anima Anandkumar · 2022
Later among the works it cites.
TEMOS: Generating diverse human motions from textual descriptions
Mathis Petrovich, Michael J. Black, and Gül Varol · 2022
Later among the works it cites.
Rethinking the openness of clip
Shuhuai Ren, Lei Li, Xuancheng Ren, Guangxiang Zhao, and Xu Sun · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Later among the works it cites.
Motionclip: Exposing human motion generation to clip space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
Haa4d: Few-shot human atomic action recognition via 3d spatio-temporal skeletal alignment, 2022
Mu-Ruei Tseng, Abhishek Gupta, Chi-Keung Tang, and Yu-Wing Tai · 2022
Later among the works it cites.
HumanNeRF: Free-viewpoint rendering of moving people from monocular video
Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron, and Ira Kemelmacher-Shlizerman · 2022
Later among the works it cites.
Mining cross-person cues for body-part interactiveness learning in hoi detection
Xiaoqian Wu, Yong-Lu Li, Xinpeng Liu, Junyi Zhang, Yuzhe Wu, and Cewu Lu · 2022
Later among the works it cites.
Motiondiffuse: Text-driven human motion generation with diffusion model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu · 2022
Later among the works it cites.
Hyperbolic image-text representations
Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shanmukha Ramakrishna Vedantam · 2023
Closest in time.
Llama-adapter v2: Parameter-efficient visual instruction model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Bridging the gap between human motion and action semantics via kinematic phrases
Xinpeng Liu, Yong-Lu Li, Ailing Zeng, Zizheng Zhou, Yang You, and Cewu Lu · 2023
Closest in time.
Symbol-llm: Leverage language models for symbolic system in visual human activity reasoning
Xiaoqian Wu, Yong-Lu Li, Jianhua Sun, and Cewu Lu · 2023
Closest in time.