Fetching the paper…
Reading the bibliography…
We present a new AI task -- Embodied Question Answering (EmbodiedQA) -- where an agent is spawned at a random location in a 3D environment and asked a question ("What color is the car?").
T. Winograd, “Understanding natural language,”
1972
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
1992
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. J. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in
1999
Earlier work this paper cites.
L. Smith and M. Gasser, “The development of embodied cognition: six lessons from babies.,”
2005
Earlier work this paper cites.
MIT press, 2005
S. Thrun, W. Burgard, and D. Fox, · 2005
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in
2015
Earlier work this paper cites.
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra, “Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?,” in
2016
Earlier work this paper cites.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,”
2016
Earlier work this paper cites.
A. Graves, “Adaptive computation time for recurrent neural networks,”
2016
Earlier work this paper cites.
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh, “Yin and Yang: Balancing and Answering Binary Visual Questions,” in
2016
Earlier work this paper cites.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal compact bilinear pooling for visual question answering and visual grounding,” in
2016
Earlier work this paper cites.
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “MovieQA: Understanding Stories in Movies through Question-Answering,”
2016
Cited alongside, same era.
S. I. Wang, P. Liang, and C. D. Manning, “Learning language games through interaction,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in
2017
Cited alongside, same era.
D. Misra, J. Langford, and Y. Artzi, “Mapping instructions and visual observations to actions with reinforcement learning,” in
2017
Closest in time.
S. Brahmbhatt and J. Hays, “DeepNav: Learning to Navigate Large Cities,” in
2017
Closest in time.
J. Oh, S. Singh, H. Lee, and P. Kohli, “Zero-shot task generalization with multi-task deep reinforcement learning,” in
2017
Closest in time.
M. Pfeiffer, M. Schaeuble, J. Nieto, R. Siegwart, and C. Cadena, “From perception to decision: A data-driven approach to end-to-end motion planning for autonomous ground robots,” in
2017
Closest in time.
M. Denil, S. G. Colmenarejo, S. Cabi, D. Saxton, and N. de Freitas, “Programmable agents,”
2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser, “Semantic scene completion from a single depth image,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “TGIF-QA: toward spatio-temporal reasoning in visual question answering,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi, “Target-driven visual navigation in indoor scenes using deep reinforcement learning,” in
2017
Cited alongside, same era.
Y. Zhu, D. Gordon, E. Kolve, D. Fox, L. Fei-Fei, A. Gupta, R. Mottaghi, and A. Farhadi, “Visual Semantic Planning using Deep Successor Representations,” in
2017
Cited alongside, same era.
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik, “Cognitive mapping and planning for visual navigation,” in
2017
Cited alongside, same era.
2017
Closest in time.
J. Andreas, D. Klein, and S. Levine, “Modular multitask reinforcement learning with policy sketches,” in
2017
Closest in time.
D. K. Misra, J. Langford, and Y. Artzi, “Mapping instructions and visual observations to actions with reinforcement learning,” in
2017
Closest in time.
I. Armeni, A. Sax, A. R. Zamir, and S. Savarese, “Joint 2D-3D-Semantic Data for Indoor Scene Understanding,”
2017
Closest in time.
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor, “A deep hierarchical approach to lifelong learning in minecraft,” in
2017
Closest in time.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick, “CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning,” in
2017
Closest in time.
Anonymous, “Building generalizable agents with a realistic and rich 3d environments,” in
2018
Closest in time.