Fetching the paper…
Reading the bibliography…
It is a long-standing problem in robotics to develop agents capable of executing diverse manipulation tasks from visual observations in unstructured real-world environments.
Rrt*-connect: Faster, asymptotically optimal motion planning
S. Klemm, J. Oberländer, A. Hermann, A. Roennau, T. Schamm, J. M. Zollner, and R. Dillmann · 2015
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
6-dof graspnet: Variational grasp generation for object manipulation
A. Mousavian, C. Eppner, and D. Fox · 2019
Earlier work this paper cites.
Occupancy networks: Learning 3d reconstruction in function space
L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger · 2019
Earlier work this paper cites.
Occupancy flow: 4d reconstruction by learning particle dynamics
M. Niemeyer, L. Mescheder, M. Oechsle, and A. Geiger · 2019
Earlier work this paper cites.
Deepsdf: Learning continuous signed distance functions for shape representation
J. J. Park, P. Florence, J. Straub, R. Newcombe, and S. Lovegrove · 2019
Earlier work this paper cites.
Scene representation networks: Continuous 3d-structure-aware neural scene representations
V. Sitzmann, M. Zollhöfer, and G. Wetzstein · 2019
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh · 2019
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Earlier work this paper cites.
Multi-task reinforcement learning with soft modularization
R. Yang, H. Xu, Y. Wu, and X. Wang · 2020
Earlier work this paper cites.
Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations
S. Song, A. Zeng, J. Lee, and T. Funkhouser · 2020
Earlier work this paper cites.
6-dof grasping for target-driven object manipulation in clutter
A. Murali, A. Mousavian, C. Eppner, C. Paxton, and D. Fox · 2020
Earlier work this paper cites.
Generalization in reinforcement learning by soft data augmentation
N. Hansen and X. Wang · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Earlier work this paper cites.
Learning continuous image representation with local implicit image function
Y. Chen, S. Liu, and X. Wang · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2021
Cited alongside, same era.
Synergies between affordance and geometry: 6-dof grasp detection via implicit representations
Z. Jiang, Y. Zhu, M. Svetlik, K. Fang, and Y. Zhu · 2021
Cited alongside, same era.
Codenerf: Disentangled neural radiance fields for object categories
W. Jang and L. Agapito · 2021
Cited alongside, same era.
Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
J. Reizenstein, R. Shapovalov, P. Henzler, L. Sbordone, P. Labatut, and D. Novotny · 2021
Cited alongside, same era.
Sharf: Shape-conditioned radiance fields from a single view
K. Rematas, R. Martin-Brualla, and V. Ferrari · 2021
Cited alongside, same era.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
N. Hansen, Z. Yuan, Y. Ze, T. Mu, A. Rajeswaran, H. Su, H. Xu, and X. Wang · 2022
Later among the works it cites.
3d neural scene representations for visuomotor control
Y. Li, S. Li, V. Sitzmann, P. Agrawal, and A. Torralba · 2022
Later among the works it cites.
Transformers as meta-learners for implicit neural representations
Y. Chen and X. Wang · 2022
Later among the works it cites.
Neural feature fusion fields: 3d distillation of self-supervised 2d image representations
V. Tschernezki, I. Laina, D. Larlus, and A. Vedaldi · 2022
Later among the works it cites.
Decomposing nerf for editing via feature field distillation
S. Kobayashi, E. Matsumoto, and V. Sitzmann · 2022
Later among the works it cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grf: Learning a general radiance field for 3d representation and rendering
A. Trevithick and B. Yang · 2021
Cited alongside, same era.
Ibrnet: Learning multi-view image-based rendering
Q. Wang, Z. Wang, K. Genova, P. P. Srinivasan, H. Zhou, J. T. Barron, R. Martin-Brualla, N. Snavely, and T. Funkhouser · 2021
Cited alongside, same era.
pixelnerf: Neural radiance fields from one or few images
A. Yu, V. Ye, M. Tancik, and A. Kanazawa · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Cited alongside, same era.
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Later among the works it cites.
Efficient neural radiance fields for interactive free-viewpoint video
H. Lin, S. Peng, Z. Xu, Y. Yan, Q. Shuai, H. Bao, and X. Zhou · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Closest in time.
Imitating task and motion planning with visuomotor transformers
M. Dalal, A. Mandlekar, C. Garrett, A. Handa, R. Salakhutdinov, and D. Fox · 2023
Closest in time.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2023
Closest in time.
Visual reinforcement learning with self-supervised 3d representations
Y. Ze, N. Hansen, Y. Chen, M. Jain, and X. Wang · 2023
Closest in time.
Snerl: Semantic-aware neural radiance fields for reinforcement learning
D. Shim, S. Lee, and H. J. Kim · 2023
Closest in time.
Mira: Mental imagery for robotic affordances
Y.-C. Lin, P. Florence, A. Zeng, J. T. Barron, Y. Du, W.-C. Ma, A. Simeonov, A. R. Garcia, and P. Isola · 2023
Closest in time.
Vision transformer for nerf-based view synthesis from a single input image
K.-E. Lin, Y.-C. Lin, W.-S. Lai, T.-Y. Lin, Y.-C. Shih, and R. Ramamoorthi · 2023
Closest in time.
Featurenerf: Learning generalizable nerfs by distilling foundation models
J. Ye, N. Wang, and X. Wang · 2023
Closest in time.
Lerf: Language embedded radiance fields
J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik · 2023
Closest in time.
Conceptfusion: Open-set multimodal 3d mapping
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, et al · 2023
Closest in time.
Neural implicit vision-language feature fields
K. Blomqvist, F. Milano, J. J. Chung, L. Ott, and R. Siegwart · 2023
Closest in time.