Fetching the paper…
Reading the bibliography…
We study how visual representations pre-trained on diverse human video data can enable data-efficient learning of downstream robotic manipulation tasks.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and J. A. Bagnell · 2011
Earlier work this paper cites.
Learning state representations with robotic priors
R. Jonschkowski and O. Brock · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
X. Wang and A. K. Gupta · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Unsupervised perceptual rewards for imitation learning
P. Sermanet, K. Xu, and S. Levine · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, and S. Levine · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, J. Schulman, E. Todorov, and S. Levine · 2018
Earlier work this paper cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Y. Liu, A. Gupta, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
T. Yu, C. Finn, S. Dasari, A. Xie, T. Zhang, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Multiple interactions made easy (mime): Large scale demonstrations data for imitation
P. Sharma, L. Mohan, L. Pinto, and A. K. Gupta · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity
A. Mandlekar, J. Booher, M. Spero, A. Tung, A. Gupta, Y. Zhu, A. Garg, S. Savarese, and L. Fei-Fei · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Earlier work this paper cites.
Third-person visual imitation learning via decoupled hierarchical controller
P. Sharma, D. Pathak, and A. Gupta · 2019
Earlier work this paper cites.
Perceptual values from observation
A. D. Edwards and C. L. Isbell · 2019
Earlier work this paper cites.
Improving robot success detection using static object data
R. Scalise, J. Thomason, Y. Bisk, and S. Srinivasa · 2019
Earlier work this paper cites.
Online object representations with contrastive learning, 2019
S. Pirk, M. Khansari, Y. Bai, C. Lynch, and P. Sermanet · 2019
Cited alongside, same era.
Learning correspondence from the cycle-consistency of time
X. Wang, A. Jabri, and A. A. Efros · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Cited alongside, same era.
Towards image-based cancer cell lines authentication using deep neural networks
D. Mzurikwao, M. Khan, O. Samuel, J. Cinatl, M. Wass, M. Michaelis, G. Marcelli, and C. S. Ang · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Ego4D: Around the World in 3,000 Hours of Egocentric Video, 2021
K. Grauman et al · 2021
Later among the works it cites.
Learning generalizable robotic reward functions from ”in-the-wild” human videos
A. S. Chen, S. Nair, and C. Finn · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2021
Later among the works it cites.
The surprising effectiveness of representation learning for visual imitation
J. Pari, N. M. M. Shafiullah, S. P. Arunachalam, and L. Pinto · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT-XML: Large scale automated ICD coding using BERT pretraining
Z. Zhang, J. Liu, and N. Razavian · 2020
Cited alongside, same era.
Bert representations for video question answering
Z. Yang, N. Garcia, C. Chu, M. Otani, Y. Nakashima, and H. Takemura · 2020
Cited alongside, same era.
Visual imitation made easy
S. Young, D. Gandhi, S. Tulsiani, A. Gupta, P. Abbeel, and L. Pinto · 2020
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown et al · 2020
Cited alongside, same era.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. B. Girshick · 2020
Cited alongside, same era.
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine · 2021
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Later among the works it cites.
Simple but effective: Clip embeddings for embodied ai
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi · 2021
Later among the works it cites.
Rrl: Resnet as representation for reinforcement learning
R. Shah and V. Kumar · 2021
Later among the works it cites.
Learning by watching: Physical imitation of manipulation skills from human videos, 2021
H. Xiong, Q. Li, Y.-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg · 2021
Later among the works it cites.
Model-based inverse reinforcement learning from visual demonstrations, 2021
N. Das, S. Bechtle, T. Davchev, D. Jayaraman, A. Rai, and F. Meier · 2021
Later among the works it cites.
Xirl: Cross-embodiment inverse reinforcement learning, 2021
K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi · 2021
Later among the works it cites.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation
S. Nair, E. Mitchell, K. Chen, B. Ichter, S. Savarese, and C. Finn · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, and F. M. L. Z. C. Feichtenhofer · 2021
Later among the works it cites.
Vision models are more robust and fair when pretrained on uncurated images without supervision
P. Goyal, Q. Duval, I. Seessel, M. Caron, I. Misra, L. Sagun, A. Joulin, and P. Bojanowski · 2022
Closest in time.
The unsurprising effectiveness of pre-trained vision models for control
S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. K. Gupta · 2022
Closest in time.
Dynamics-aware metric embedding: Metric learning in a latent space for visual planning
M. Hong, K. Lee, M. Kang, W. Jung, and S. Oh · 2022
Closest in time.
Reinforcement learning with action-free pre-training from videos
Y. Seo, K. Lee, S. James, and P. Abbeel · 2022
Closest in time.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Closest in time.
Can Foundation Models Perform Zero-Shot Task Specification For Robot Manipulation?
Y. Cui, S. Niekum, A. Gupta, V. Kumar, and A. Rajeswaran · 2022
Closest in time.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Closest in time.
Human hands as probes for interactive object understanding
M. Goyal, S. Modi, R. Goyal, and S. Gupta · 2022
Closest in time.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2022
Closest in time.