Fetching the paper…
Reading the bibliography…
Remarkable progress has been made in recent years in the fields of vision, language, and robotics.
“6-DOF GraspNet: Variational Grasp Generation for Object Manipulation”, 2019
Arsalan Mousavian, Clemens Eppner and Dieter Fox · 1905
Earlier work this paper cites.
“Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, 2019
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
“S4G: Amodal Single-view Single-Shot SE(3) Grasp Detection in Cluttered Scenes”, 2019
Yuzhe Qin et al · 1910
Earlier work this paper cites.
“Language Conditioned Imitation Learning over Unstructured Data”
Corey Lynch and Pierre Sermanet · 2005
Earlier work this paper cites.
“Path planning for mobile robot navigation using voronoi diagram and fast marching”
Santiago Garrido, Luis Moreno, Mohamed Abderrahim and Fernando Martin · 2006
Earlier work this paper cites.
“Supersizing Self-supervision: Learning to Grasp from 50K Tries and 700 Robot Hours”, 2015
Lerrel Pinto and Abhinav Gupta · 2015
Earlier work this paper cites.
Jeffrey Mahler et al · 2017
Earlier work this paper cites.
“Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics”
Jeffrey Mahler et al · 2017
Earlier work this paper cites.
“Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection”
Sergey Levine et al · 2018
Earlier work this paper cites.
“Robot Learning in Homes: Improving Generalization and Reducing Dataset Bias”
Abhinav Gupta, Adithyavairavan Murali, Dhiraj Gandhi and Lerrel Pinto · 2018
Earlier work this paper cites.
Jeffrey Mahler et al · 2018
Earlier work this paper cites.
“QT-Opt: Scalable deep reinforcement learning for vision-based robotic manipulation”
Dmitry Kalashnikov et al · 2018
Earlier work this paper cites.
“Ok-vqa: A visual question answering benchmark requiring external knowledge”
Kenneth Marino, Mohammad Rastegari, Ali Farhadi and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“Visual Representations for Semantic Target Driven Navigation”
Arsalan Mousavian et al · 2019
Earlier work this paper cites.
“Graspnet-1billion: a large-scale benchmark for general object grasping”
Hao-Shu Fang, Chenxi Wang, Minghao Gou and Cewu Lu · 2020
Earlier work this paper cites.
“Learning latent plans from play”
Corey Lynch et al · 2020
Earlier work this paper cites.
“Nerf: Representing scenes as neural radiance fields for view synthesis”
Ben Mildenhall et al · 2020
Earlier work this paper cites.
“Learning Transferable Visual Models From Natural Language Supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes”
Martin Sundermeyer, Arsalan Mousavian, Rudolph Triebel and Dieter Fox · 2021
Earlier work this paper cites.
“Film: Following instructions in language with modular methods”
So Min et al · 2021
Earlier work this paper cites.
“Acronym: A large-scale grasp dataset based on simulation”
Clemens Eppner, Arsalan Mousavian and Dieter Fox · 2021
Earlier work this paper cites.
“Do as I can, not as I say: Grounding language in robotic affordances”
Michael Ahn et al · 2022
Earlier work this paper cites.
“Rt-1: Robotics transformer for real-world control at scale”
Anthony Brohan et al · 2022
Earlier work this paper cites.
“Detecting twenty-thousand classes using image-level supervision”
Xingyi Zhou et al · 2022
Earlier work this paper cites.
“Simple Open-Vocabulary Object Detection with Vision Transformers”
Matthias Minderer et al · 2022
Cited alongside, same era.
“Flamingo: a Visual Language Model for Few-Shot Learning”, 2022
Jean-Baptiste Alayrac et al · 2022
Cited alongside, same era.
“Ifor: Iterative flow minimization for robotic object rearrangement”
Ankit Goyal et al · 2022
Cited alongside, same era.
“Language-Grounded Indoor 3D Semantic Segmentation in the Wild”, 2022
David Rozenberszki, Or Litany and Angela Dai · 2022
Cited alongside, same era.
“A persistent spatial semantic representation for high-level natural language instruction execution”
Valts Blukis et al · 2022
Cited alongside, same era.
“Segment Anything”
Alexander Kirillov et al · 2023
Later among the works it cites.
“Lerf: Language embedded radiance fields”
Justin Kerr et al · 2023
Later among the works it cites.
Matthew Chang et al · 2023
Later among the works it cites.
“Audio Visual Language Maps for Robot Navigation”
Chenguang Huang, Oier Mees, Andy Zeng and Wolfram Burgard · 2023
Later among the works it cites.
“Learning Dexterous Manipulation from Exemplar Object Trajectories and Pre-Grasps”, 2023
Sudeep Dasari, Abhinav Gupta and Vikash Kumar · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matt Deitke et al · 2022
Cited alongside, same era.
“Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models”
Huy Ha and Shuran Song · 2022
Cited alongside, same era.
“Open-vocabulary Queryable Scene Representations for Real World Planning”
Boyuan Chen et al · 2022
Cited alongside, same era.
“Behavior Transformers: Cloning k k modes with one stone”
Nur Shafiullah, Zichen Cui, Ariuntuya Altanzaya and Lerrel Pinto · 2022
Cited alongside, same era.
“From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data”, 2022
Zichen Cui, Yibin Wang, Nur Muhammad Shafiullah and Lerrel Pinto · 2022
Cited alongside, same era.
“CLIPort: What and where pathways for robotic manipulation”
Mohit Shridhar, Lucas Manuelli and Dieter Fox · 2022
Cited alongside, same era.
“The design of Stretch: A compact, lightweight mobile manipulator for indoor human environments”
Charles Kemp, Aaron Edsinger, Henry Clever and Blaine Matulevich · 2022
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
“Objaverse: A universe of annotated 3d objects”
Matt Deitke et al · 2023
Later among the works it cites.
“Visual language maps for robot navigation”
Chenguang Huang, Oier Mees, Andy Zeng and Wolfram Burgard · 2023
Later among the works it cites.
“Conceptfusion: Open-set multimodal 3d mapping”
Krishna Jatavallabhula et al · 2023
Later among the works it cites.
“ViNT: A Foundation Model for Visual Navigation”
Dhruv Shah et al · 2023
Later among the works it cites.
“Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning”
Qiao Gu et al · 2023
Later among the works it cites.
“Progprompt: Generating situated robot task plans using large language models”
I. Singh et al · 2023
Later among the works it cites.
“Perceiver-Actor: A multi-task transformer for robotic manipulation”
Mohit Shridhar, Lucas Manuelli and Dieter Fox · 2023
Later among the works it cites.
“Spatial-Language Attention Policies for Efficient Robot Learning”
Priyam Parashar, Jay Vakil, Sam Powers and Chris Paxton · 2023
Later among the works it cites.
“Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation”
Theophile Gervet, Zhou Xian, Nikolaos Gkanatsios and Katerina Fragkiadaki · 2023
Later among the works it cites.
“Learning Hybrid Actor-Critic Maps for 6D Non-Prehensile Manipulation”
Wenxuan Zhou et al · 2023
Later among the works it cites.
“VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models”
Wenlong Huang et al · 2023
Later among the works it cites.
“Code as Policies: Language model programs for embodied control”
Jacky Liang et al · 2023
Later among the works it cites.
“Voyager: An Open-Ended Embodied Agent with Large Language Models”
Guanzhi Wang et al · 2023
Later among the works it cites.
“ProgPrompt: Generating Situated Robot Task Plans using Large Language Models”
Ishika Singh et al · 2023
Later among the works it cites.
“ASC: Adaptive Skill Coordination for Robotic Mobile Manipulation”
Naoki Yokoyama et al · 2023
Later among the works it cites.
“Open-World Object Manipulation using Pre-trained Vision-Language Models”, 2023
Austin Stone et al · 2023
Later among the works it cites.
“GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration”
Naoki Wake et al · 2023
Later among the works it cites.
“Sayplan: Grounding large language models using 3d scene graphs for scalable task planning”
Krishan Rana et al · 2023
Later among the works it cites.