Fetching the paper…
Reading the bibliography…
In-depth scene descriptions and question answering tasks have greatly increased the scope of today's definition of scene understanding.
Intuitive physics
McCloskey, M.: · 1983
Earlier work this paper cites.
Neuroanimator: Fast neural network emulation and control of physics-based models
Grzeszczuk, R., Terzopoulos, D., Hinton, G.: · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction. Volume 1
Sutton, R.S., Barto, A.G.: · 1998
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
Battaglia, P.W., Hamrick, J.B., Tenenbaum, J.B.: · 2013
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
Ranzato, M., Szlam, A., Bruna, J., Mathieu, M., Collobert, R., Chopra, S.: · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Malinowski, M., Fritz, M.: · 2014
Earlier work this paper cites.
Towards a visual turing challenge
Malinowski, M., Fritz, M.: · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., Manning, C.D.: · 2014
Earlier work this paper cites.
Galileo: Perceiving physical object properties by integrating a physics engine with deep learning
Wu, J., Yildirim, I., Lim, J.J., Freeman, B., Tenenbaum, J.: · 2015
Earlier work this paper cites.
Visual Turing test for computer vision systems
Geman, D., Geman, S., Hallonquist, N., Younes, L.: · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Ren, M., Kiros, R., Zemel, R.: · 2015
Earlier work this paper cites.
Visual madlibs: Fill in the blank description generation and question answering
Yu, L., Park, E., Berg, A.C., Berg, T.L.: · 2015
Earlier work this paper cites.
VQA: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., Parikh, D.: · 2015
Earlier work this paper cites.
Benchmarking in manipulation research: The YCB object and model set and benchmarking protocols
Çalli, B., Walsman, A., Singh, A., Srinivasa, S., Abbeel, P., Dollar, A.M.: · 2015
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems (2015) Software available from tensorflow.org
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., Zheng, X.: · 2015
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
Mathieu, M., Couprie, C., LeCun, Y.: · 2016
Cited alongside, same era.
Learning physical intuition of block towers by example
Lerer, A., Gross, S., Fergus, R.: · 2016
Cited alongside, same era.
“What happens if…” Learning to predict the effect of forces in images
Mottaghi, R., Rastegari, M., Gupta, A., Farhadi, A.: · 2016
Cited alongside, same era.
Interaction networks for learning about objects, relations and physics
Battaglia, P., Pascanu, R., Lai, M., Rezende, D.J., et al.: · 2016
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: · 2016
Later among the works it cites.
Visual interaction networks
Watters, N., Tacchetti, A., Weber, T., Pascanu, R., Battaglia, P., Zoran, D.: · 2017
Later among the works it cites.
Visual stability prediction for robotic manipulation
Li, W., Leonardis, A., Fritz, M.: · 2017
Later among the works it cites.
Ask your neurons: A deep learning approach to visual question answering
Malinowski, M., Rohrbach, M., Fritz, M.: · 2017
Later among the works it cites.
Learning to reason: End-to-end module networks for visual question answering
Hu, R., Andreas, J., Rohrbach, M., Darrell, T., Saenko, K.: · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D.G., Malinowski, M., Pascanu, R., Battaglia, P., Lillicrap, T.: · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
To fall or not to fall: A visual approach to physical stability prediction
Li, W., Azimi, S., Leonardis, A., Fritz, M.: · 2016
Cited alongside, same era.
Newtonian scene understanding: Unfolding the dynamics of objects in static images
Mottaghi, R., Bagherinezhad, H., Rastegari, M., Farhadi, A.: · 2016
Cited alongside, same era.
Visual7W: Grounded question answering in images
Zhu, Y., Groth, O., Bernstein, M., Fei-Fei, L.: · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Yang, Z., He, X., Gao, J., Deng, L., Smola, A.: · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M.: · 2016
Cited alongside, same era.
Beattie, C., Leibo, J.Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., Petersen, S.: · 2016
Cited alongside, same era.
Later among the works it cites.
Taking visual motion prediction to new heightfields
Ehrhardt, S., Monszpart, A., Mitra, N.J., Vedaldi, A.: · 2017
Later among the works it cites.
Long-term image boundary extrapolation
Bhattacharyya, A., Malinowski, M., Schiele, B., Fritz, M.: · 2018
Closest in time.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Agrawal, A., Batra, D., Parikh, D., Kembhavi, A.: · 2018
Closest in time.
Dvqa: Understanding data visualizations via question answering
Kafle, K., Cohen, S., Price, B., Kanan, C.: · 2018
Closest in time.
Airsim: High-fidelity visual and physical simulation for autonomous vehicles
Shah, S., Dey, D., Lovett, C., Kapoor, A.: · 2018
Closest in time.
Building generalizable agents with a realistic and rich 3d environment
Wu, Y., Wu, Y., Gkioxari, G., Tian, Y.: · 2018
Closest in time.
Bullet 3
Coumans, E.: · 2018
Closest in time.