Fetching the paper…
Reading the bibliography…
To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments (Matterport 3D).
Twenty-two colors of maximum contrast
Kenneth L Kelly · 1965
Earlier work this paper cites.
Lazy theta*: Any-angle path planning and path length analysis in 3d
Alex Nash, Sven Koenig, and Craig Tovey · 2010
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Voxnet: A 3d convolutional neural network for real-time object recognition
Daniel Maturana and Sebastian Scherer · 2015
Earlier work this paper cites.
Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture
David Eigen and Rob Fergus · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Vizdoom: A doom-based AI research platform for visual reinforcement learning
Michal Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaskowski · 2016
Earlier work this paper cites.
Volumetric and multi-view cnns for object classification on 3d data
Charles R Qi, Hao Su, Matthias Nießner, Angela Dai, Mengyuan Yan, and Leonidas J Guibas · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Cognitive mapping and planning for visual navigation
Saurabh Gupta, James Davidson, Sergey Levine, Rahul Sukthankar, and Jitendra Malik · 2017
Earlier work this paper cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi · 2017
Earlier work this paper cites.
Visual Semantic Planning using Deep Successor Representations
Yuke Zhu, Daniel Gordon, Eric Kolve, Dieter Fox, Li Fei-Fei, Abhinav Gupta, Roozbeh Mottaghi, and Ali Farhadi · 2017
Earlier work this paper cites.
Gated-attention architectures for task-oriented language grounding
Devendra Singh Chaplot, Kanthashree Mysore Sathyendra, Rama Kumar Pasumarthi, Dheeraj Rajagopal, and Ruslan Salakhutdinov · 2017
Earlier work this paper cites.
Grounded language learning in a simulated 3d world
Karl Moritz Hermann, Felix Hill, Simon Green, Fumin Wang, Ryan Faulkner, Hubert Soyer, David Szepesvari, Wojtek Czarnecki, Max Jaderberg, Denis Teplyashin, et al · 2017
Cited alongside, same era.
Semantic scene completion from a single depth image
Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser · 2017
Cited alongside, same era.
Visual odometry and mapping for autonomous flight using an rgb-d camera
Albert S Huang, Abraham Bachrach, Peter Henry, Michael Krainin, Daniel Maturana, Dieter Fox, and Nicholas Roy · 2017
Cited alongside, same era.
Multirobot symbolic planning under temporal uncertainty
Shiqi Zhang, Yuqian Jiang, Guni Sharon, and Peter Stone · 2017
Cited alongside, same era.
Matterport3D: Learning from RGB-D data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2017
Cited alongside, same era.
O-cnn: Octree-based convolutional neural networks for 3d shape analysis
Peng-Shuai Wang, Yang Liu, Yu-Xiao Guo, Chun-Yu Sun, and Xin Tong · 2017
Later among the works it cites.
Learning representations and generative models for 3d point clouds
Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Embodied Question Answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Later among the works it cites.
Building generalizable agents with a realistic and rich 3d environments
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CAD2RL: Real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2017
Cited alongside, same era.
Learning to fly by crashing
Abhinav Gupta Dhiraj Gandhi, Lerrel Pinto · 2017
Cited alongside, same era.
MINOS: Multimodal indoor simulator for navigation in complex environments
Manolis Savva, Angel X. Chang, Alexey Dosovitskiy, Thomas Funkhouser, and Vladlen Koltun · 2017
Cited alongside, same era.
Home: a household multimodal environment
Simon Brodeur, Ethan Perez, Ankesh Anand, Florian Golemo, Luca Celotti, Florian Strub, Jean Rouat, Hugo Larochelle, and Aaron C. Courville · 2017
Cited alongside, same era.
AI2-THOR: An Interactive 3D Environment for Visual AI
Eric Kolve, Roozbeh Mottaghi, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Cited alongside, same era.
Semantic scene completion from a single depth image
Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser · 2017
Cited alongside, same era.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Cited alongside, same era.
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel · 2018
Later among the works it cites.
IQA: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi · 2018
Later among the works it cites.
Neural Modular Control for Embodied Question Answering
Abhishek Das, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Later among the works it cites.
Unity: A general platform for intelligent agents, 2018
Arthur Juliani, Vincent-Pierre Berges, Esh Vckay, Yuan Gao, Hunter Henry, Marwan Mattar, and Danny Lange · 2018
Later among the works it cites.
Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching
Andy Zeng, Shuran Song, Kuan-Ting Yu, Elliott Donlon, Francois Robert Hogan, Maria Bauza, Daolin Ma, Orion Taylor, Melody Liu, Eudald Romo, Nima Fazeli, Ferran Alet, Nikhil Chavan Dafle, Rachel Holladay, Isabella Morona, Prem Qu Nair, Druck Green, Ian Taylor, Weber Liu, Thomas Funkhouser, and Alberto Rodriguez · 2018
Later among the works it cites.
Attentional shapecontextnet for point cloud recognition
Saining Xie, Sainan Liu, Zeyu Chen, and Zhuowen Tu · 2018
Later among the works it cites.
Splatnet: Sparse lattice networks for point cloud processing
Hang Su, Varun Jampani, Deqing Sun, Subhransu Maji, Evangelos Kalogerakis, Ming-Hsuan Yang, and Jan Kautz · 2018
Later among the works it cites.
Blindfold Baselines for Embodied QA
Ankesh Anand, Eugene Belilovsky, Kyle Kastner, Hugo Larochelle, and Aaron Courville · 2018
Later among the works it cites.
Shifting the baseline: Single modality performance on visual navigation & qa
Jesse Thomason, Daniel Gordan, and Yonatan Bisk · 2018
Later among the works it cites.
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.