Fetching the paper…
Reading the bibliography…
In order to *generalize* to various tasks in the wild, robotic agents will need a suitable representation (i.e., vision network) that enables the robot to predict optimal actions given high dimensional vision inputs.
Smoothing and differentiation of data by simplified least squares procedures
Abraham Savitzky and Marcel JE Golay · 1964
Earlier work this paper cites.
The ecological approach to visual perception
JJ Gibson · 1979
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal et al · 1997
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Object tracking using sift features and mean shift
Huiyu Zhou, Yuan Yuan, and Chunmei Shi · 2009
Earlier work this paper cites.
From 3d scene geometry to human workspace
Abhinav Gupta, Scott Satkin, Alexei A Efros, and Martial Hebert · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso · 2013
Earlier work this paper cites.
Scene parsing by integrating function, geometry and appearance models
Yibiao Zhao and Song-Chun Zhu · 2013
Earlier work this paper cites.
D Eigen and R Fergus · 2014
Earlier work this paper cites.
Action-reaction: Forecasting the dynamics of human interaction
De-An Huang and Kris M Kitani · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A hierarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese · 2014
Earlier work this paper cites.
Reasoning about object affordances in a knowledge base representation
Yuke Zhu, Alireza Fathi, and Li Fei-Fei · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Anticipating human activities using object affordances for reactive robotic response
Hema S Koppula and Ashutosh Saxena · 2015
Earlier work this paper cites.
Affordance detection of tool parts from geometric features
Austin Myers, Ching L Teo, Cornelia Fermüller, and Yiannis Aloimonos · 2015
Earlier work this paper cites.
Marr revisited: 2d-3d alignment via surface normal prediction
Aayush Bansal, Bryan Russell, and Abhinav Gupta · 2016
Earlier work this paper cites.
End to end learning for self-driving cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al · 2016
Earlier work this paper cites.
Tutorial on variational autoencoders
Carl Doersch · 2016
Earlier work this paper cites.
Recurrent neural networks for driver activity anticipation via sensory-fusion architecture
Ashesh Jain, Avi Singh, Hema S Koppula, Shane Soh, and Ashutosh Saxena · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta · 2016
Earlier work this paper cites.
Learning action maps of large environments via first-person vision
Nicholas Rhinehart and Kris M Kitani · 2016
Earlier work this paper cites.
A multi-scale cnn for affordance segmentation in rgb images
Anirban Roy and Sinisa Todorovic · 2016
Earlier work this paper cites.
Predicting motivations of actions by leveraging text
Carl Vondrick, Deniz Oktay, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Inferring forces and learning human utilities from videos
Yixin Zhu, Chenfanfu Jiang, Yibiao Zhao, Demetri Terzopoulos, and Song-Chun Zhu · 2016
Earlier work this paper cites.
Next-active-object prediction from egocentric videos
Antonino Furnari, Sebastiano Battiato, Kristen Grauman, and Giovanni Maria Farinella · 2017
Earlier work this paper cites.
Red: Reinforced encoder-decoder networks for action anticipation
Jiyang Gao, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
End-to-end recovery of human shape and pose
Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik · 2017
Earlier work this paper cites.
Weakly supervised affordance detection
Johann Sawatzky, Abhilash Srikantha, and Juergen Gall · 2017
Cited alongside, same era.
Decomposing motion and content for natural video sequence prediction
Ruben Villegas, Jimei Yang, Seunghoon Hong, Xunyu Lin, and Honglak Lee · 2017
Cited alongside, same era.
When will you do what?-anticipating temporal occurrences of activities
Yazan Abu Farha, Alexander Richard, and Juergen Gall · 2018
Cited alongside, same era.
Robot learning in homes: Improving generalization and reducing dataset bias
Abhinav Gupta, Adithyavairavan Murali, Dhiraj Prakashchand Gandhi, and Lerrel Pinto · 2018
Cited alongside, same era.
Visual affordance and function understanding: a survey. arxiv
M Hassanin, S Khan, and M Tahtali · 2018
Cited alongside, same era.
Rrl: Resnet as representation for reinforcement learning
Rutav M Shah and Vikash Kumar · 2021
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Malla, Deeksha Manjunath, Igor Mordatch, Ofir Nachum, Carolina Parada, Jodilyn Peralta, Emily Perez, Karl Pertsch, Jornell Quiambao, Kanishka Rao, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Kevin Sayed, Jaspiar Singh, Sumedh Sontakke, Austin Stone, Clayton Tan, Huong Tran, Vincent Vanhoucke, Steve Vega, Quan Vuong, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich · 2022
Later among the works it cites.
Epic-kitchens visor benchmark: Video segmentations and object relations
Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen · 2022
Later among the works it cites.
Human hands as probes for interactive object understanding
Mohit Goyal, Sahil Modi, Rishabh Goyal, and Saurabh Gupta · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Robonet: Large-scale multi-robot learning
Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn · 2019
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2022
Later among the works it cites.
Dexvip: Learning dexterous grasping with human hand pose priors from video
Priyanka Mandikal and Kristen Grauman · 2022
Later among the works it cites.
Intention-conditioned long-term human egocentric action forecasting@ ego4d challenge 2022
Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Later among the works it cites.
Real-world robot learning with masked visual pre-training
Ilija Radosavovic, Tete Xiao, Stephen James, Pieter Abbeel, Jitendra Malik, and Trevor Darrell · 2022
Later among the works it cites.
Videodex: Learning dexterity from internet videos
Kenneth Shaw, Shikhar Bahl, and Deepak Pathak · 2022
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2022
Later among the works it cites.
Masked visual pre-training for motor control
Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krähenbühl, and Ishan Misra · 2022
Later among the works it cites.
Affordances from human videos as a versatile representation for robotics
Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, and Deepak Pathak · 2023
Later among the works it cites.
Towards generalizable zero-shot manipulation via translating human interaction plans
Homanga Bharadhwaj, Abhinav Gupta, Vikash Kumar, and Shubham Tulsiani · 2023
Later among the works it cites.
What makes pre-trained visual representations successful for robust manipulation?
Kaylee Burns, Zach Witzel, Jubayer Ibn Hamid, Tianhe Yu, Chelsea Finn, and Karol Hausman · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song · 2023
Later among the works it cites.
An unbiased look at datasets for visuo-motor pre-training
Sudeep Dasari, Mohan Kumar Srirama, Unnat Jain, and Abhinav Gupta · 2023
Later among the works it cites.
The expressive power of tuning only the normalization layers
Angeliki Giannou, Shashank Rajput, and Dimitris Papailiopoulos · 2023
Later among the works it cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei · 2023
Later among the works it cites.
Deft: Dexterous fine-tuning for real-world hand policies
Aditya Kannan, Kenneth Shaw, Shikhar Bahl, Pragna Mannam, and Deepak Pathak · 2023
Later among the works it cites.
Where are we in the search for an artificial visual cortex for embodied intelligence?
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Pieter Abbeel, Jitendra Malik, et al · 2023
Later among the works it cites.
Leap hand:low-cost, efficient, and anthropomorphic hand for robot learning
Kenneth Shaw, Ananye Agarwal, and Deepak Pathak · 2023
Later among the works it cites.
Manipulate by seeing: Creating manipulation controllers from pre-trained representations
Jianren Wang, Sudeep Dasari, Mohan Kumar Srirama, Shubham Tulsiani, and Abhinav Gupta · 2023
Later among the works it cites.
Affordance diffusion: Synthesizing hand-object interactions
Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello, Stan Birchfield, Jiaming Song, Shubham Tulsiani, and Sifei Liu · 2023
Later among the works it cites.
Tuning layernorm in attention: Towards efficient multi-modal llm finetuning
Bingchen Zhao, Haoqin Tu, Chen Wei, Jieru Mei, and Cihang Xie · 2023
Later among the works it cites.
Hacman: Learning hybrid actor-critic maps for 6d non-prehensile manipulation
Wenxuan Zhou, Bowen Jiang, Fan Yang, Chris Paxton, and David Held · 2023
Later among the works it cites.
Look ma, no hands! agent-environment factorization of egocentric videos
Matthew Chang, Aditya Prakash, and Saurabh Gupta · 2024
Closest in time.
Yuanchen Ju, Kaizhe Hu, Guowei Zhang, Gu Zhang, Mingrun Jiang, and Huazhe Xu · 2024
Closest in time.