Fetching the paper…
Reading the bibliography…
In this work, we propose a unified framework, called Visual Reasoning with Differ-entiable Physics (VRDP), that can jointly learn visual concepts and infer physics models of objects and their interactions from videos and language.
A stability property of implicit runge-kutta methods
J. C. Butcher · 1975
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Real time physics: class notes
M. Müller, J. Stam, D. James, and N. Thürey · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Modeling and solving constraints
E. Catto · 2009
Earlier work this paper cites.
Bullet physics engine
E. Coumans · 2010
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
P. W. Battaglia, J. B. Hamrick, and J. B. Tenenbaum · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Imagining the unseen: Stability-based cuboid arrangements for scene understanding
T. Shao, A. Monszpart, Y. Zheng, B. Koo, W. Xu, K. Zhou, and N. J. Mitra · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Earlier work this paper cites.
Galileo: Perceiving physical object properties by integrating a physics engine with deep learning
J. Wu, I. Yildirim, J. J. Lim, W. T. Freeman, and J. B. Tenenbaum · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Earlier work this paper cites.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
P. W. Battaglia, R. Pascanu, M. Lai, D. Rezende, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Learning physical intuition of block towers by example
A. Lerer, S. Gross, and R. Fergus · 2016
Earlier work this paper cites.
“what happens if…” learning to predict the effect of forces in images
R. Mottaghi, M. Rastegari, A. Gupta, and A. Farhadi · 2016
Earlier work this paper cites.
MovieQA: Understanding Stories in Movies through Question-Answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Earlier work this paper cites.
What value do explicit high level concepts have in vision to language problems?
Q. Wu, C. Shen, L. Liu, A. Dick, and A. van den Hengel · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
A compositional object-based approach to learning physical dynamics
M. B. Chang, T. Ullman, A. Torralba, and J. B. Tenenbaum · 2017
Earlier work this paper cites.
VQS: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
C. Gan, Y. Li, H. Li, C. Sun, and B. Gong · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
Inferring and executing programs for visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, J. Hoffman, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
T. N. Kipf and M. Welling · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Visual interaction networks: Learning a physics simulator from video
N. Watters, D. Zoran, T. Weber, P. Battaglia, R. Pascanu, and A. Tacchetti · 2017
Cited alongside, same era.
Learning to see physics via visual de-animation
J. Wu, E. Lu, P. Kohli, B. Freeman, and J. Tenenbaum · 2017
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
J. Mao, C. Gan, P. Kohli, J. B. Tenenbaum, and J. Wu · 2019
Later among the works it cites.
Modeling expectation violation in intuitive physics with coarse probabilistic object representations
K. Smith, L. Mei, S. Yao, J. Wu, E. Spelke, J. Tenenbaum, and T. Ullman · 2019
Later among the works it cites.
Differentiable physics and stable modes for tool-use and manipulation planning
M. A. Toussaint, K. R. Allen, K. A. Smith, and J. B. Tenenbaum · 2019
Later among the works it cites.
Compositional video prediction
Y. Ye, M. Singh, A. Gupta, and S. Tulsiani · 2019
Later among the works it cites.
Social-iq: A question answering benchmark for artificial social intelligence
A. Zadeh, M. Chan, P. P. Liang, E. Tong, and L.-P. Morency · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Cited alongside, same era.
Video question answering via attribute-augmented attention network learning
Y. Ye, Z. Zhao, Y. Li, L. Chen, J. Xiao, and Y. Zhuang · 2017
Cited alongside, same era.
End-to-end differentiable physics for learning and control
F. de Avila Belbute-Peres, K. Smith, K. Allen, J. Tenenbaum, and J. Z. Kolter · 2018
Cited alongside, same era.
Explainable neural computation via stack neural module networks
R. Hu, J. Andreas, T. Darrell, and K. Saenko · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning
D. A. Hudson and C. D. Manning · 2018
Cited alongside, same era.
Tvqa: Localized, compositional video question answering
J. Lei, L. Yu, M. Bansal, and T. L. Berg · 2018
Cited alongside, same era.
S. Amizadeh, H. Palangi, O. Polozov, Y. Huang, and K. Koishida · 2020
Later among the works it cites.
Cophy: Counterfactual learning of physical dynamics
F. Baradel, N. Neverova, J. Mille, G. Mori, and C. Wolf · 2020
Later among the works it cites.
Learning physical graph representations from visual scenes
D. M. Bear, C. Fan, D. Mrowca, Y. Li, S. Alter, A. Nayebi, J. Schwartz, L. Fei-Fei, J. Wu, J. B. Tenenbaum, et al · 2020
Later among the works it cites.
D. Ding, F. Hill, A. Santoro, and M. Botvinick · 2020
Later among the works it cites.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
R. Girdhar and D. Ramanan · 2020
Later among the works it cites.
Interpretable visual reasoning via probabilistic formulation under natural supervision
X. Han, S. Wang, C. Su, W. Zhang, Q. Huang, and Q. Tian · 2020
Later among the works it cites.
Neuralsim: Augmenting differentiable simulators with neural networks
E. Heiden, D. Millard, E. Coumans, Y. Sheng, and G. S. Sukhatme · 2020
Later among the works it cites.
Difftaichi: Differentiable programming for physical simulation
Y. Hu, L. Anderson, T.-M. Li, Q. Sun, N. Carr, J. Ragan-Kelley, and F. Durand · 2020
Later among the works it cites.
Location-aware graph convolutional networks for video question answering
D. Huang, P. Chen, R. Zeng, Q. Du, M. Tan, and C. Gan · 2020
Later among the works it cites.
Contrastive learning of structured world models
T. Kipf, E. van der Pol, and M. Welling · 2020
Later among the works it cites.
Hierarchical conditional relation networks for video question answering
T. M. Le, V. Le, S. Venkatesh, and T. Tran · 2020
Later among the works it cites.
Visual grounding of learned physical models
Y. Li, T. Lin, K. Yi, D. Bear, D. Yamins, J. Wu, J. Tenenbaum, and A. Torralba · 2020
Later among the works it cites.
Differentiable physics simulation
J. Liang and M. C. Lin · 2020
Later among the works it cites.
Object-centric learning with slot attention
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf · 2020
Later among the works it cites.
Entity abstraction in visual model-based reinforcement learning
R. Veerapaneni, J. D. Co-Reyes, M. Chang, M. Janner, C. Finn, J. Wu, J. Tenenbaum, and S. Levine · 2020
Later among the works it cites.
Clevrer: Collision events for video representation and reasoning
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum · 2020
Later among the works it cites.
Grounding physical concepts of objects and events through dynamic visual reasoning
Z. Chen, J. Mao, J. Wu, K.-Y. K. Wong, J. B. Tenenbaum, and C. Gan · 2021
Closest in time.
Threedworld: A platform for interactive multi-modal physical simulation
C. Gan, J. Schwartz, S. Alter, M. Schrimpf, J. Traer, J. De Freitas, J. Kubilius, A. Bhandwaldar, N. Haber, M. Sano, et al · 2021
Closest in time.
C. Gan, S. Zhou, J. Schwartz, S. Alter, A. Bhandwaldar, D. Gutfreund, D. L. Yamins, J. J. DiCarlo, J. McDermott, A. Torralba, et al · 2021
Closest in time.
Plasticinelab: A soft-body manipulation benchmark with differentiable physics
Z. Huang, Y. Hu, T. Du, S. Zhou, H. Su, J. B. Tenenbaum, and C. Gan · 2021
Closest in time.
gradsim: Differentiable simulation for system identification and visuomotor control
K. M. Jatavallabhula, M. Macklin, F. Golemo, V. Voleti, L. Petrini, M. Weiss, B. Considine, J. Parent-Levesque, K. Xie, K. Erleben, et al · 2021
Closest in time.
Learning long-term visual dynamics with region proposal interaction networks
H. Qi, X. Wang, D. Pathak, Y. Ma, and J. Malik · 2021
Closest in time.
Hopper: Multi-hop transformer for spatiotemporal reasoning
H. Zhou, A. Kadav, F. Lai, A. Niculescu-Mizil, M. R. Min, M. Kapadia, and H. P. Graf · 2021
Closest in time.