Fetching the paper…
Reading the bibliography…
To build general robotic agents that can operate in many environments, it is often imperative for the robot to collect experience in the real world.
Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography
M. A. Fischler and R. C. Bolles · 1981
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Dynamic movement primitives-a framework for motor control in humans and humanoid robotics
S. Schaal · 2006
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu · 2013
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Earlier work this paper cites.
Dynamical movement primitives: Learning attractor models for motor behaviors
A. J. Ijspeert, J. Nakanishi, H. Hoffmann, P. Pastor, and S. Schaal · 2013
Earlier work this paper cites.
Dynamic movement primitives for human-robot interaction: Comparison with human behavioral observation
M. Prada, A. Remazeilles, A. Koene, and S. Endo · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Earlier work this paper cites.
Smpl: A skinned multi-person linear model
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black · 2015
Earlier work this paper cites.
Mujoco haptix: A virtual reality system for hand manipulation
V. Kumar and E. Todorov · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
The curious robot: Learning visual representations via physical interactions
L. Pinto, D. Gandhi, Y. Han, Y.-L. Park, and A. Gupta · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
End to end learning for self-driving cars, 2016
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, and K. Zieba · 2016
Earlier work this paper cites.
Pixelwise View Selection for Unstructured Multi-View Stereo
J. L. Schönberger, E. Zheng, M. Pollefeys, and J.-M. Frahm · 2016
Earlier work this paper cites.
The ”something something” video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollar, and R. Girshick · 2017
Earlier work this paper cites.
Embodied hands: Modeling and capturing hands and bodies together
J. Romero, D. Tzionas, and M. J. Black · 2017
Earlier work this paper cites.
End-to-end recovery of human shape and pose
A. Kanazawa, M. J. Black, D. W. Jacobs, and J. Malik · 2017
Earlier work this paper cites.
Learning robot activities from first-person human videos using convolutional future regression
J. Lee and M. S. Ryoo · 2017
Cited alongside, same era.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2018
Cited alongside, same era.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, and S. Levine · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
A. V. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Dexpilot: Vision-based teleoperation of dexterous robotic hand-arm system
A. Handa, K. Van Wyk, W. Yang, J. Liang, Y.-W. Chao, Q. Wan, S. Birchfield, N. Ratliff, and D. Fox · 2020
Later among the works it cites.
Hierarchical neural dynamic policies
S. Bahl, A. Gupta, and D. Pathak · 2021
Later among the works it cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al · 2021
Later among the works it cites.
Dexmv: Imitation learning for dexterous manipulation from human videos
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Online optical marker-based hand tracking with deep labels
S. Han, B. Liu, R. Wang, Y. Ye, C. D. Twigg, and K. Kin · 2018
Cited alongside, same era.
Freihand: A dataset for markerless capture of hand pose and shape from single rgb images
C. Zimmermann, D. Ceylan, J. Yang, B. Russell, M. Argus, and T. Brox · 2019
Cited alongside, same era.
Expressive body capture: 3D hands, face, and body from a single image
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black · 2019
Cited alongside, same era.
Vision-based teleoperation of shadow dexterous hand using end-to-end deep neural network
S. Li, X. Ma, H. Liang, M. Görner, P. Ruppel, B. Fang, F. Sun, and J. Zhang · 2019
Cited alongside, same era.
Third-person visual imitation learning via decoupled hierarchical controller
P. Sharma, D. Pathak, and A. Gupta · 2019
Cited alongside, same era.
Openpose: realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, and Y. Sheikh · 2019
Cited alongside, same era.
Frankmocap: A monocular 3d whole-body pose estimation system via regression and integration
Y. Rong, T. Shiratori, and H. Joo · 2021
Later among the works it cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg · 2021
Later among the works it cites.
Learning generalizable robotic reward functions from” in-the-wild” human videos
A. S. Chen, S. Nair, and C. Finn · 2021
Later among the works it cites.
The surprising effectiveness of representation learning for visual imitation
J. Pari, N. Muhammad, S. P. Arunachalam, L. Pinto, et al · 2021
Later among the works it cites.
Rb2: Robotic manipulation benchmarking with a twist
S. Dasari, J. Wang, J. Hong, S. Bahl, Y. Lin, A. S. Wang, A. Thankaraj, K. S. Chahal, B. Calli, S. Gupta, et al · 2021
Later among the works it cites.
Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam
C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. Montiel, and J. D. Tardós · 2021
Later among the works it cites.
Adabins: Depth estimation using adaptive bins
S. F. Bhat, I. Alhashim, and P. Wonka · 2021
Later among the works it cites.
d3rlpy: An offline deep reinforcement library
M. I. Takuma Seno · 2021
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Closest in time.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Closest in time.
Dexterous imitation made easy: A learning-based framework for efficient dexterous manipulation, 2022
S. P. Arunachalam, S. Silwal, B. Evans, and L. Pinto · 2022
Closest in time.
Y. Qin, H. Su, and X. Wang · 2022
Closest in time.
Dexvip: Learning dexterous grasping with human hand pose priors from video
P. Mandikal and K. Grauman · 2022
Closest in time.
Human-to-robot imitation in the wild
S. Bahl, A. Gupta, and D. Pathak · 2022
Closest in time.
Robotic telekinesis: Learning a robotic hand imitator by watching humans on youtube, 2022
A. Sivakumar, K. Shaw, and D. Pathak · 2022
Closest in time.
Detecting twenty-thousand classes using image-level supervision
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Closest in time.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Closest in time.