Fetching the paper…
Reading the bibliography…
We propose the Neuro-Symbolic Concept Learner (NS-CL), a model that learns visual concepts, words, and semantic parsing of sentences without explicit supervision on any of them; instead, our model learns by simply looking at images and reading paired questions and answers.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
English verb classes and alternations , volume 1
Beth Levin · 1993
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A probabilistic computational model of cross-situational word learning
Afsaneh Fazly, Afra Alishahi, and Suzanne Stevenson · 2010
Earlier work this paper cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Yoav Artzi and Luke Zettlemoyer · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M Malinowski and M Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Abc-cnn: An attention based convolutional neural network for visual question answering
Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia · 2015
Earlier work this paper cites.
Learning language through pictures
Grzegorz Chrupała, Akos Kádár, and Afra Alishahi · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
Deep convolutional inverse graphics network
Tejas D Kulkarni, William F Whitney, Pushmeet Kohli, and Josh Tenenbaum · 2015
Cited alongside, same era.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning · 2015
Cited alongside, same era.
Weakly-supervised disentangling with recurrent transformations for 3d view synthesis
Jimei Yang, Scott E Reed, Ming-Hsuan Yang, and Honglak Lee · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Learning to compose neural networks for question answering
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Zero-shot task generalization with multi-task deep reinforcement learning
Junhyuk Oh, Satinder Singh, Honglak Lee, and Pushmeet Kohli · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David GT Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Later among the works it cites.
Learning disentangled representations with semi-supervised deep generative models
N Siddharth, T. B. Paige, J.W. Meent, A. Desmaison, N. Goodman, P. Kohli, F. Wood, and P. Torr · 2017
Later among the works it cites.
Neural scene de-rendering
Jiajun Wu, Joshua B Tenenbaum, and Pushmeet Kohli · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Later among the works it cites.
Object level visual reasoning in videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li Dong and Mirella Lapata · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten · 2016
Cited alongside, same era.
DenseCap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Cited alongside, same era.
Training and evaluating multimodal word embeddings with large-scale web annotated images
Junhua Mao, Jiajing Xu, Kevin Jing, and Alan L Yuille · 2016
Cited alongside, same era.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Huijuan Xu and Kate Saenko · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Cited alongside, same era.
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, and Greg Mori · 2018
Later among the works it cites.
Word learning and the acquisition of syntactic–semantic overhypotheses
Jon Gauthier, Roger Levy, and Joshua B Tenenbaum · 2018
Later among the works it cites.
Scan: learning abstract hierarchical compositional visual concepts
Irina Higgins, Nicolas Sonnerat, Loic Matthey, Arka Pal, Christopher P Burgess, Matthew Botvinick, Demis Hassabis, and Alexander Lerchner · 2018
Later among the works it cites.
Explainable neural computation via stack neural module networks
Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2018
Later among the works it cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Later among the works it cites.
From skills to symbols: Learning symbolic representations for abstract high-level planning
George Konidaris, Leslie Pack Kaelbling, and Tomas Lozano-Perez · 2018
Later among the works it cites.
Transparency by design: Closing the gap between performance and interpretability in visual reasoning
David Mascharka, Philip Tran, Ryan Soklaski, and Arjun Majumdar · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Later among the works it cites.
Learning visually-grounded semantics from contrastive adversarial samples
Haoyue Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang, and Jian Sun · 2018
Later among the works it cites.
DDRprog: A clevr differentiable dynamic reasoning programmer
Joseph Suarez, Justin Johnson, and Fei-Fei Li · 2018
Later among the works it cites.
Neural-Symbolic VQA: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Joshua B Tenenbaum · 2018
Later among the works it cites.