Fetching the paper…
Reading the bibliography…
We introduce Interactive Question Answering (IQA), the task of answering questions that require an autonomous agent to interact with a dynamic visual environment.
Strips: A new approach to the application of theorem proving to problem solving
R. E. Fikes and N. J. Nilsson · 1971
Earlier work this paper cites.
Real-time obstacle avoidance for fast mobile robots
J. Borenstein and Y. Koren · 1989
Earlier work this paper cites.
The vector field histogram-fast obstacle avoidance for mobile robots
J. Borenstein and Y. Koren · 1991
Earlier work this paper cites.
On-line map building and navigation for autonomous mobile robots
G. Oriolo, M. Vendittelli, and G. Ulivi · 1995
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
R. E. Parr and S. J. Russell · 1997
Earlier work this paper cites.
Reactive navigation in outdoor environments using potential fields
H. Haddad, M. Khatib, S. Lacroix, and R. Chatila · 1998
Earlier work this paper cites.
Symbolic navigation with a generic map
D. Kim and R. Nevatia · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. P. Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
T. G. Dietterich · 2000
Earlier work this paper cites.
Torcs, the open racing car simulator
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner · 2000
Earlier work this paper cites.
Real-time simultaneous localisation and mapping with a single camera
A. J. Davison · 2003
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
N. Kohl and P. Stone · 2004
Earlier work this paper cites.
High speed obstacle avoidance using monocular vision and reinforcement learning
J. Michels, A. Saxena, and A. Y. Ng · 2005
Earlier work this paper cites.
Vision-based 3-d trajectory tracking for unknown environments
P. Saeedi, P. D. Lawrence, and D. G. Lowe · 2006
Earlier work this paper cites.
Autonomous vision-based exploration and mapping using hybrid maps and rao-blackwellised particle filters
R. Sim and J. J. Little · 2006
Earlier work this paper cites.
3-d object map building using dense object models with sift-based recognition features
M. Tomono · 2006
Earlier work this paper cites.
A guide to vision-based map building
D. Wooden · 2006
Earlier work this paper cites.
Trajectory optimization using reinforcement learning for map exploration
T. Kollar and N. Roy · 2008
Earlier work this paper cites.
Integrating symbolic and geometric planning for mobile manipulation
C. Dornhege, M. Gissler, M. Teschner, and B. Nebel · 2009
Earlier work this paper cites.
Learning appearance in virtual scenarios for pedestrian detection
J. Marin, D. Vázquez, D. Gerónimo, and A. M. López · 2010
Earlier work this paper cites.
Hierarchical task and motion planning in the now
L. P. Kaelbling and T. Lozano-Pérez · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Using classical planners for tasks with continuous operators in robotics
S. Srivastava, L. Riano, S. Russell, and P. Abbeel · 2013
Earlier work this paper cites.
Lsd-slam: Large-scale direct monocular slam
J. Engel, T. Schöps, and D. Cremers · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Cited alongside, same era.
Combined task and motion planning through an extensible planner-independent interface layer
S. Srivastava, E. Fang, L. Riano, R. Chitnis, S. Russell, and P. Abbeel · 2014
Cited alongside, same era.
Deepdriving: Learning affordance for direct perception in autonomous driving
C. Chen, A. Seff, A. Kornhauser, and J. Xiao · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
Orb-slam: a versatile and accurate monocular slam system
Communicating hierarchical neural controllers for learning zero-shot task generalization
J. Oh, S. Singh, H. Lee, and P. Kohli · 2016
Later among the works it cites.
Playing for data: Ground truth from computer games
S. R. Richter, V. Vineet, S. Roth, and V. Koltun · 2016
Later among the works it cites.
The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez · 2016
Later among the works it cites.
Cad2rl: Real single-image flight without a single real image
F. Sadeghi and S. Levine · 2016
Later among the works it cites.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos · 2015
Cited alongside, same era.
Semantic pose using deep networks trained on synthetic rgb-d
J. Papon and M. Schoeler · 2015
Cited alongside, same era.
Galileo: Perceiving physical object properties by integrating a physics engine with deep learning
J. Wu, I. Yildirim, J. J. Lim, B. Freeman, and J. Tenenbaum · 2015
Cited alongside, same era.
Play and learn: Using video games to train computer vision models
J. L. Alireza Shafaei and M. Schmidt · 2016
Cited alongside, same era.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Understanding realworld indoor sceneswith synthetic data
A. Handa, V. PǍtrǍucean, V. Badrinarayanan, R. Cipolla, et al · 2016
Cited alongside, same era.
Fast, robust, continuous monocular egomotion computation
A. Jaegle, S. Phillips, and K. Daniilidis · 2016
Cited alongside, same era.
Q. Wu, D. Teney, P. Wang, C. Shen, A. R. Dick, and A. van den Hengel · 2016
Later among the works it cites.
Vqa: Visual question answering
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Parikh, and D. Batra · 2017
Closest in time.
Gated-attention architectures for task-oriented language grounding
D. S. Chaplot, K. M. Sathyendra, R. K. Pasumarthi, D. Rajagopal, and R. Salakhutdinov · 2017
Closest in time.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Closest in time.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Closest in time.
Understanding grounded language learning agents
F. Hill, K. M. Hermann, P. Blunsom, and S. Clark · 2017
Closest in time.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim · 2017
Closest in time.
Inferring and executing programs for visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, F. fei Li, C. L. Zitnick, and R. B. Girshick · 2017
Closest in time.
Figureqa: An annotated figure dataset for visual reasoning
S. E. Kahou, A. Atkinson, V. Michalski, Á. Kádár, A. Trischler, and Y. Bengio · 2017
Closest in time.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi · 2017
Closest in time.
Deepstory: Video story qa by deep embedded memory networks
K.-M. Kim, M.-O. Heo, S.-H. Choi, and B.-T. Zhang · 2017
Closest in time.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Closest in time.
A deep hierarchical approach to lifelong learning in minecraft
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor · 2017
Closest in time.
Fvqa: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. van den Hengel, and A. R. Dick · 2017
Closest in time.
Visual semantic planning using deep successor representations
Y. Zhu, D. Gordon, E. Kolve, D. Fox, L. Fei-Fei, A. Gupta, R. Mottaghi, and A. Farhadi · 2017
Closest in time.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi · 2017
Closest in time.
Yolov3: An incremental improvement
J. Redmon and A. Farhadi · 2018
Closest in time.