Fetching the paper…
Reading the bibliography…
Reasoning in the real world is not divorced from situations.
Situations, actions, and causal laws
J. McCarthy · 1963
Earlier work this paper cites.
Reasoning about action using a possible models approach
M. S. Winslett · 1988
Earlier work this paper cites.
Situated cognition and the culture of learning
J. S. Brown, A. Collins, and P. Duguid · 1989
Earlier work this paper cites.
Situated cognition: Stepping out of representational flatland
W. J. Clancey · 1991
Earlier work this paper cites.
The frame problem in the situation calculus: A simple solution (sometimes) and a completeness result for goal regression
R. Reiter · 1991
Earlier work this paper cites.
Situated learning: Legitimate peripheral participation
M. Bloch · 1994
Earlier work this paper cites.
Reasoning about action and change
H. Prendinger and G. Schurz · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The situation calculus: A case for modal logic
G. Lakemeyer · 2010
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Earlier work this paper cites.
Actions˜ transformations
X. Wang, A. Farhadi, and A. Gupta · 2016
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
RMPE: Regional multi-person pose estimation
H.-S. Fang, S. Xie, Y.-W. Tai, and C. Lu · 2017
Cited alongside, same era.
Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
C. Gan, Y. Li, H. Li, C. Sun, and B. Gong · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Cited alongside, same era.
Tvqa+: Spatio-temporal grounding for video question answering
J. Lei, L. Yu, T. L. Berg, and M. Bansal · 2019
Later among the works it cites.
Visualbert: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Later among the works it cites.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
R. Girdhar and D. Ramanan · 2020
Later among the works it cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
J. Ji, R. Krishna, L. Fei-Fei, and J. C. Niebles · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Johnson, B. Hariharan, L. Van Der Maaten, J. Hoffman, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Cited alongside, same era.
Marioqa: Answering questions by watching gameplay videos
J. Mun, P. Hongsuck Seo, I. Jung, and B. Han · 2017
Cited alongside, same era.
What actions are needed for understanding human actions in videos?
G. A. Sigurdsson, O. Russakovsky, and A. Gupta · 2017
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Tvqa: Localized, compositional video question answering
J. Lei, L. Yu, M. Bansal, and T. L. Berg · 2018
Cited alongside, same era.
Hierarchical conditional relation networks for video question answering
T. M. Le, V. Le, S. Venkatesh, and T. Tran · 2020
Later among the works it cites.
Gector–grammatical error correction: Tag, not rewrite
K. Omelianchuk, V. Atrasevych, A. Chernodub, and O. Skurzhanskyi · 2020
Later among the works it cites.
Unbiased scene graph generation from biased training
K. Tang, Y. Niu, J. Huang, J. Shi, and H. Zhang · 2020
Later among the works it cites.
Clevrer: Collision events for video representation and reasoning
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum · 2020
Later among the works it cites.
Grounding physical concepts of objects and events through dynamic visual reasoning
Z. Chen, J. Mao, J. Wu, K.-Y. K. Wong, J. B. Tenenbaum, and C. Gan · 2021
Later among the works it cites.
Dynamic visual reasoning by learning differentiable physics models from video and language
M. Ding, Z. Chen, T. Du, P. Luo, J. B. Tenenbaum, and C. Gan · 2021
Later among the works it cites.
Agqa: A benchmark for compositional spatio-temporal reasoning
M. Grunde-McLaughlin, R. Krishna, and M. Agrawala · 2021
Later among the works it cites.
Ptr: A benchmark for part-based conceptual, relational, and physical reasoning
Y. Hong, L. Yi, J. Tenenbaum, A. Torralba, and C. Gan · 2021
Later among the works it cites.
Movinets: Mobile video networks for efficient video recognition
D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid · 2021
Later among the works it cites.