Fetching the paper…
Reading the bibliography…
Large-scale visuomotor policy learning is a promising approach toward developing generalizable manipulation systems.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Sim2Real viewpoint invariant visual servoing by recurrent control
F. Sadeghi, A. Toshev, E. Jang, and S. Levine · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks
A. Chattopadhay, A. Sarkar, P. Howlader, and V. N. Balasubramanian · 2018
Earlier work this paper cites.
Accelerating 3D deep learning with PyTorch3D
N. Ravi, J. Reizenstein, D. Novotny, T. Gordon, W.-Y. Lo, J. Johnson, and G. Gkioxari · 2020
Earlier work this paper cites.
robosuite: A modular simulation framework and benchmark for robot learning
Y. Zhu, J. Wong, A. Mandlekar, R. Martín-Martín, A. Joshi, S. Nasiriany, and Y. Zhu · 2020
Earlier work this paper cites.
Unsupervised learning of visual 3d keypoints for control
B. Chen, P. Abbeel, and D. Pathak · 2021
Earlier work this paper cites.
pixelNeRF: Neural radiance fields from one or few images
A. Yu, V. Ye, M. Tancik, and A. Kanazawa · 2021
Earlier work this paper cites.
GRF: Learning a general radiance field for 3D representation and rendering
A. Trevithick and B. Yang · 2021
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Earlier work this paper cites.
NeRF: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2021
Earlier work this paper cites.
PyTorch library for cam methods
J. Gildenblat and contributors · 2021
Earlier work this paper cites.
Reinforcement learning with neural radiance fields
D. Driess, I. Schubert, P. Florence, Y. Li, and M. Toussaint · 2022
Earlier work this paper cites.
Vision-based manipulators need to also see from their hands
K. Hsu, M. J. Kim, R. Rafailov, J. Wu, and C. Finn · 2022
Earlier work this paper cites.
CACTI: A framework for scalable multi-task multi-scene visual imitation learning
Z. Mandi, H. Bharadhwaj, V. Moens, S. Song, A. Rajeswaran, and V. Kumar · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Multi-view masked world models for visual robotic manipulation
Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel · 2023
Cited alongside, same era.
ExAug: Robot-conditioned navigation policies via geometric experience augmentation
N. Hirose, D. Shah, A. Sridhar, and S. Levine · 2023
Cited alongside, same era.
NeRF in the palm of your hand: Corrective augmentation for robotics via novel-view synthesis
A. Zhou, M. J. Kim, L. Wang, P. Florence, and C. Finn · 2023
Cited alongside, same era.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Cited alongside, same era.
Act3D: Infinite resolution action detection transformer for robotic manipulation
Decomposing the generalization gap in imitation learning for visual robotic manipulation
A. Xie, L. Lee, T. Xiao, and C. Finn · 2024
Closest in time.
THE COLOSSEUM: A benchmark for evaluating generalization for robotic manipulation
W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox · 2024
Closest in time.
Efficient data collection for robotic manipulation via compositional generalization
J. Gao, A. Xie, T. Xiao, C. Finn, and D. Sadigh · 2024
Closest in time.
DROID: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Closest in time.
ZeroNVS: Zero-shot 360-degree view synthesis from a single real image
K. Sargent, Z. Li, T. Shah, C. Herrmann, H.-X. Yu, Y. Zhang, E. R. Chan, D. Lagun, L. Fei-Fei, D. Sun, and J. Wu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki · 2023
Cited alongside, same era.
Learning generalizable manipulation policies with object-centric 3D representations
Y. Zhu, Z. Jiang, P. Stone, and Y. Zhu · 2023
Cited alongside, same era.
RVT: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Cited alongside, same era.
Zero-1-to-3: Zero-shot one image to 3D object
R. Liu, R. Wu, B. V. Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Cited alongside, same era.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Cited alongside, same era.
Dall-E-Bot: Introducing web-scale diffusion models to robotics
I. Kapelyukh, V. Vosylius, and E. Johns · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Cited alongside, same era.
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Closest in time.
3D diffuser actor: Policy diffusion with 3D scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Closest in time.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song · 2024
Closest in time.
ReconFusion: 3D reconstruction with diffusion priors
R. Wu, B. Mildenhall, P. Henzler, K. Park, R. Gao, D. Watson, P. P. Srinivasan, D. Verbin, J. T. Barron, B. Poole, and A. Holynski · 2024
Closest in time.
CAT3D: Create anything in 3D with multi-view diffusion models
R. Gao*, A. Holynski*, P. Henzler, A. Brussee, R. Martin-Brualla, P. P. Srinivasan, J. T. Barron, and B. Poole* · 2024
Closest in time.
Mirage: Cross-embodiment zero-shot policy transfer with cross-painting
L. Y. Chen, K. Hari, K. Dharmarajan, C. Xu, Q. Vuong, and K. Goldberg · 2024
Closest in time.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel · 2024
Closest in time.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine · 2024
Closest in time.
Can pre-trained text-to-image models generate visual goals for reinforcement learning?
J. Gao, K. Hu, G. Xu, and H. Xu · 2024
Closest in time.
Video language planning
Y. Du, M. Yang, P. Florence, F. Xia, A. Wahid, B. Ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, et al · 2024
Closest in time.