Hake: Human activity knowledge engine
Original
Yong-Lu Li, Liang Xu, Xinpeng Liu, Xijie Huang, Yue Xu, Mingyang Chen, Ze Ma, Shiyi Wang, Hao-Shu Fang, and Cewu Lu · 1904
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Original
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 1908
Earlier work this paper cites.
A stochastic grammar of images
Song-Chun Zhu and David Mumford · 2007
Earlier work this paper cites.
Sift flow: Dense correspondence across scenes and its applications
Ce Liu, Jenny Yuen, and Antonio Torralba · 2010
Earlier work this paper cites.
Phrasal recognition
Ali Farhadi and Mohammad Amin Sadeghi · 2013
Earlier work this paper cites.
HICO: A benchmark for recognizing human-object interactions in images
Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
SM Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, Geoffrey E Hinton, et al · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Original
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Learning models for actions and person-object interactions with transfer to question answering
Arun Mallya and Svetlana Lazebnik · 2016
Earlier work this paper cites.
Attentional pooling for action recognition
Original
Rohit Girdhar and Deva Ramanan · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Shapeworld-a new test methodology for multimodal language understanding
Original
Alexander Kuhnle and Ann Copestake · 2017
Earlier work this paper cites.
Transitive invariance for self-supervised visual representation learning
Xiaolong Wang, Kaiming He, and Abhinav Gupta · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Earlier work this paper cites.
Systematic generalization: what is required and can it be learned?
Original
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, and Aaron Courville · 2018
Earlier work this paper cites.
Measuring abstract reasoning in neural networks
David Barrett, Felix Hill, Adam Santoro, Ari Morcos, and Timothy Lillicrap · 2018
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Earlier work this paper cites.
Pairwise body-part attention for recognizing human-object interactions
Hao-Shu Fang, Jinkun Cao, Yu-Wing Tai, and Cewu Lu · 2018
Earlier work this paper cites.