Fetching the paper…
Reading the bibliography…
We propose associating language utterances to 3D visual abstractions of the scene they describe.
The Ecological Approach to Visual Perception
James J. Gibson · 1979
Earlier work this paper cites.
Symbol grounding and meaning: A comparison of high-dimensional and embodied theories of meaning
A. Glenberg and D. Robertson · 2000
Earlier work this paper cites.
Grounding language in action
Arthur Glenberg and Michael P. Kaschak · 2002
Earlier work this paper cites.
Embodied meaning in a neural theory of language
Jerome Feldman and Srinivas Narayanan · 2004
Earlier work this paper cites.
Mental simulation in spatial language processing
Benjamin Bergen · 2005
Earlier work this paper cites.
From Molecule to Metaphor: A Neural Theory of Language
Jerome A. Feldman · 2006
Earlier work this paper cites.
Experimental methods for simulation semantics
B Bergen · 2007
Earlier work this paper cites.
Modulation of the ffa and ppa by language related to faces and places
Lisa Aziz-Zadeh, Christian Fiebach, Srini Narayanan, Jerome Feldman, Ellen Dodge, and Richard Ivry · 2008
Earlier work this paper cites.
Neural dissociations between action verb understanding and motor imagery
Roel M Willems, Ivan Toni, Peter Hagoort, and Daniel Casasanto · 2009
Earlier work this paper cites.
Modulation of bold response in motion-sensitive lateral temporal cortex by real and fictive motion sentences
Ayse Pinar Saygin, Stephen Mccullough, Morana Alac, and Karen Emmorey · 2009
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele · 2012
Earlier work this paper cites.
Perception as an inference problem
Bruno A. Olshausen · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
Marcus Rohrbach, Qiu Wei, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele · 2013
Earlier work this paper cites.
Semantic parsing via paraphrasing
Jonathan Berant and Percy Liang · 2014
Earlier work this paper cites.
Control-limited differential dynamic programming
Yuval Tassa, Nicolas Mansard, and Emanuel Todorov · 2014
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Embodiment, simulation and meaning
Benjamin Bergen · 2015
Cited alongside, same era.
Bullet physics simulation
Erwin Coumans · 2015
Cited alongside, same era.
Language models for image captioning: The quirks and what works
Jacob Devlin, Hao Cheng, Hao Fang, Saurabh Gupta, Li Deng, Xiaodong He, Geoffrey Zweig, and Margaret Mitchell · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Cited alongside, same era.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al · 2015
Cited alongside, same era.
Teaching machines to read and comprehend
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick · 2016
Later among the works it cites.
Text understanding with the attention sum reader network
Rudolf Kadlec, Martin Schmid, Ondrej Bajgar, and Jan Kleindienst · 2016
Later among the works it cites.
Learning dexterous manipulation policies from experience and imitation
Vikash Kumar, Abhishek Gupta, Emanuel Todorov, and Sergey Levine · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Cited alongside, same era.
Deep convolutional inverse graphics network
Tejas D. Kulkarni, Will Whitney, Pushmeet Kohli, and Joshua B. Tenenbaum · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Cited alongside, same era.
Learning common sense through visual abstraction
Ramakrishna Vedantam, Xiao Lin, Tanmay Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Structured training for neural network transition-based parsing
David Weiss, Chris Alberti, Michael Collins, and Slav Petrov · 2015
Cited alongside, same era.
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick · 2017
Later among the works it cites.
Generating descriptions with grounded and co-referenced people
Anna Rohrbach, Marcus Rohrbach, Siyu Tang, Seong Joon Oh, and Bernt Schiele · 2017
Later among the works it cites.
Vision-as-inverse-graphics: Obtaining a rich 3d explanation of a scene from a single image
Lukasz Romaszko, Christopher K. I. Williams, Pol Moreno, and Pushmeet Kohli · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Later among the works it cites.
Hsiao-Yu Fish Tung, Adam W. Harley, William Seto, and Katerina Fragkiadaki · 2017
Later among the works it cites.
Using syntax to ground referring expressions in natural images
Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency · 2018
Later among the works it cites.
Probabilistic neural programmed networks for scene generation
Zhiwei Deng, Jiacheng Chen, YIFANG FU, and Greg Mori · 2018
Later among the works it cites.
Reward learning using natural language
Fish Tung and Katerina Fragkiadaki · 2018
Later among the works it cites.
The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu · 2019
Closest in time.
Intphys: A benchmark for visual intuitive physics reasoning
Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux · 2019
Closest in time.
Cycle-consistency for robust visual question answering
Meet Shah, Xinlei Chen, Marcus Rohrbach, and Devi Parikh · 2019
Closest in time.
Learning spatial common sense with geometry-aware recurrent networks
Hsiao-Yu Fish Tung, Ricson Cheng, and Katerina Fragkiadaki · 2019
Closest in time.