Fetching the paper…
Reading the bibliography…
Generalization in robotic manipulation remains a critical challenge, particularly when scaling to new environments with limited demonstrations.
“ALVINN: An Autonomous Land Vehicle in a Neural Network”
Dean Pomerleau · 1988
Earlier work this paper cites.
“U-Net: Convolutional Networks for Biomedical Image Segmentation”
Olaf Ronneberger, Philipp Fischer and Thomas Brox · 2015
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“The ”Something Something” Video Database for Learning and Evaluating Visual Common Sense”
Raghav Goyal et al · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“FiLM: Visual Reasoning with a General Conditioning Layer”
Ethan Perez et al · 2018
Earlier work this paper cites.
“Bootstrap Your Own Latent - A New Approach to Self-Supervised Learning”
Jean-Bastien Grill et al · 2020
Earlier work this paper cites.
“Understanding Human Hands in Contact at Internet Scale”
Dandan Shan, Jiaqi Geng, Michelle Shu and David. Fouhey · 2020
Earlier work this paper cites.
“Emerging Properties in Self-Supervised Vision Transformers”
Mathilde Caron et al · 2021
Earlier work this paper cites.
“An Empirical Study of Training Self-Supervised Vision Transformers”
Xinlei Chen, Saining Xie and Kaiming He · 2021
Earlier work this paper cites.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”
Alexey Dosovitskiy et al · 2021
Earlier work this paper cites.
“What Matters in Learning from Offline Human Demonstrations for Robot Manipulation”
Ajay Mandlekar et al · 2021
Earlier work this paper cites.
“Learning Transferable Visual Models From Natural Language Supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“Denoising Diffusion Implicit Models”
Jiaming Song, Chenlin Meng and Stefano Ermon · 2021
Earlier work this paper cites.
“Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100”
Dima Damen et al · 2022
Earlier work this paper cites.
“Ego4D: Around the World in 3, 000 Hours of Egocentric Video”
Kristen Grauman et al · 2022
Earlier work this paper cites.
“Masked Autoencoders Are Scalable Vision Learners”
Kaiming He et al · 2022
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”
Edward. Hu et al · 2022
Earlier work this paper cites.
“Perceiver IO: A General Architecture for Structured Inputs & Outputs”
Andrew Jaegle et al · 2022
Earlier work this paper cites.
“R3M: A Universal Visual Representation for Robot Manipulation”
Suraj Nair et al · 2022
Earlier work this paper cites.
“The Surprising Effectiveness of Representation Learning for Visual Imitation”
Jyothish Pari, Nur(Mahi) Shafiullah, Sridhar Arunachalam and Lerrel Pinto · 2022
Cited alongside, same era.
“Real-World Robot Learning with Masked Visual Pre-training”
Ilija Radosavovic et al · 2022
Cited alongside, same era.
“High-Resolution Image Synthesis with Latent Diffusion Models”
Robin Rombach et al · 2022
Cited alongside, same era.
“Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation”
Mohit Shridhar, Lucas Manuelli and Dieter Fox · 2022
Cited alongside, same era.
“RT-1: Robotics Transformer for Real-World Control at Scale”
Anthony Brohan et al · 2023
Cited alongside, same era.
“What Makes Pre-Trained Visual Representations Successful for Robust Manipulation?”
“Open X-Embodiment: Robotic Learning Datasets and RT-X Models”
Open-Embodiment Collaboration et al · 2024
Closest in time.
URL: https://en.dh-robotics.com/product/ag
“Dahuan AG Series Gripper”, 2024 · 2024
Closest in time.
“Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics”
Norman Di and Edward Johns · 2024
Closest in time.
“RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot”
Haoshu Fang et al · 2024
Closest in time.
URL: https://www.flexiv.com/product/rizon
“Flexiv Rizon Robot”, 2024 · 2024
Closest in time.
URL: https://www.forcedimension.com/products/sigma
“Force Dimension - sigma.7”, 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaylee Burns et al · 2023
Cited alongside, same era.
“Diffusion Policy: Visuomotor Policy Learning via Action Diffusion”
Cheng Chi et al · 2023
Cited alongside, same era.
“AnyGrasp: Robust and Efficient Grasp Perception in Spatial and Temporal Domains”
Haoshu Fang et al · 2023
Cited alongside, same era.
“Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation”
Théophile Gervet, Zhou Xian, Nikolaos Gkanatsios and Katerina Fragkiadaki · 2023
Cited alongside, same era.
“Segment Anything”
Alexander Kirillov et al · 2023
Cited alongside, same era.
“LIV: Language-Image Representations and Rewards for Robotic Control”
Yecheng Ma et al · 2023
Cited alongside, same era.
“Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?”
Arjun Majumdar et al · 2023
Cited alongside, same era.
URL: https://www.intelrealsense.com/depth-camera-d415
“Intel RealSense Depth Camera D415”, 2024 · 2024
Closest in time.
URL: https://www.intelrealsense.com/depth-camera-d435
“Intel RealSense Depth Camera D435”, 2024 · 2024
Closest in time.
“Droid: A large-scale in-the-wild robot manipulation dataset”
Alexander Khazatsky et al · 2024
Closest in time.
“OpenVLA: An Open-Source Vision-Language-Action Model”
Moo Kim et al · 2024
Closest in time.
“SpawnNet: Learning Generalizable Visuomotor Skills from Pre-trained Network”
Xingyu Lin et al · 2024
Closest in time.
“ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data”
Zeyi Liu et al · 2024
Closest in time.
“Octo: An Open-Source Generalist Robot Policy”
Octo Model Team et al · 2024
Closest in time.
“DINOv2: Learning Robust Visual Features without Supervision”
Maxime Oquab et al · 2024
Closest in time.
“Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies”
Jianing Qian, Anastasios Panagopoulos and Dinesh Jayaraman · 2024
Closest in time.
“Theia: Distilling Diverse Vision Foundation Models for Robot Learning”
Jinghuan Shang et al · 2024
Closest in time.
“DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation”
Chen Wang et al · 2024
Closest in time.
“RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective”
Chenxi Wang, Hongjie Fang, Hao-Shu Fang and Cewu Lu · 2024
Closest in time.
“3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations”
Yanjie Ze et al · 2024
Closest in time.
“SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation”
Junjie Zhang et al · 2024
Closest in time.