Fetching the paper…
Reading the bibliography…
Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization.
RTFM: Generalising to New Environment Dynamics via Reading
V. Zhong, T. Rocktäschel, and E. Grefenstette · 1910
Earlier work this paper cites.
Ray tracing volume densities
J. T. Kajiya and B. P. Von Herzen · 1984
Earlier work this paper cites.
Simultaneous map building and localization for an autonomous mobile robot
J. J. Leonard and H. F. Durrant-Whyte · 1991
Earlier work this paper cites.
Optical models for direct volume rendering
N. Max · 1995
Earlier work this paper cites.
Localization of Autonomous Guided Vehicles
H. Durrant-Whyte, D. Rye, and E. Nebot · 1996
Earlier work this paper cites.
Multiple View Geometry in Computer Vision
R. Hartley and A. Zisserman · 2004
Earlier work this paper cites.
Simultaneous localization and mapping: part i
H. Durrant-Whyte and T. Bailey · 2006
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2016
Earlier work this paper cites.
Universal Correspondence Network
C. B. Choy, J. Gwak, S. Savarese, and M. Chandraker · 2016
Earlier work this paper cites.
Structure-from-Motion Revisited
J. L. Schönberger and J.-M. Frahm · 2016
Earlier work this paper cites.
Pixelwise View Selection for Unstructured Multi-View Stereo
J. L. Schönberger, E. Zheng, M. Pollefeys, and J.-M. Frahm · 2016
Earlier work this paper cites.
Self-supervised visual descriptor learning for dense correspondence
T. Schmidt, R. Newcombe, and D. Fox · 2017
Earlier work this paper cites.
Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics
J. Mahler, J. Liang, S. Niyaz, M. Laskey, R. Doan, X. Liu, J. A. Ojea, and K. Goldberg · 2017
Earlier work this paper cites.
Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation
P. R. Florence, L. Manuelli, and R. Tedrake · 2018
Earlier work this paper cites.
PyBullet Planning
C. Garrett · 2018
Earlier work this paper cites.
Learning 6-dof grasping interaction via deep geometry-aware 3d representations
X. Yan, J. Hsu, M. Khansari, Y. Bai, A. Pathak, A. Gupta, J. Davidson, and H. Lee · 2018
Earlier work this paper cites.
Learning to understand goal specifications by modelling reward
D. Bahdanau, F. Hill, J. Leike, E. Hughes, A. Hosseini, P. Kohli, and E. Grefenstette · 2018
Earlier work this paper cites.
Learning spatial common sense with geometry-aware recurrent networks
H.-Y. F. Tung, R. Cheng, and K. Fragkiadaki · 2019
Earlier work this paper cites.
Learning from unlabelled videos using contrastive predictive neural 3d mapping
A. W. Harley, S. K. Lakshmikanth, F. Li, X. Zhou, H.-Y. F. Tung, and K. Fragkiadaki · 2019
Earlier work this paper cites.
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2020
Earlier work this paper cites.
Language-Conditioned Imitation Learning for Robot Manipulation Tasks
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor · 2020
Cited alongside, same era.
Self-Supervised Correspondence in Visuomotor Policy Learning
P. Florence, L. Manuelli, and R. Tedrake · 2020
Cited alongside, same era.
3D-OES: Viewpoint-invariant object-factorized environment simulators
H.-Y. F. Tung, Z. Xian, M. Prabhudesai, S. Lal, and K. Fragkiadaki · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
OpenCLIP, July 2021
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt · 2021
Cited alongside, same era.
Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
LM-nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osinski, B. Ichter, and S. Levine · 2022
Later among the works it cites.
RT-1: Robotics Transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Later among the works it cites.
Evo-NeRF: Evolving NeRF for Sequential Robot Grasping of Transparent Objects
J. Kerr, L. Fu, H. Huang, Y. Avigal, M. Tancik, J. Ichnowski, A. Kanazawa, and K. Goldberg · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
pixelNeRF: Neural radiance fields from one or few images
A. Yu, V. Ye, M. Tancik, and A. Kanazawa · 2021
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Cited alongside, same era.
Score-Based Generative Modeling through Stochastic Differential Equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2021
Cited alongside, same era.
Barf: Bundle-adjusting neural radiance fields
C.-H. Lin, W.-C. Ma, A. Torralba, and S. Lucey · 2021
Cited alongside, same era.
NeRF − − -- : Neural Radiance Fields Without Known Camera Parameters
Z. Wang, S. Wu, W. Xie, M. Chen, and V. A. Prisacariu · 2021
Cited alongside, same era.
A. Simeonov, Y. Du, L. Yen-Chen, , A. Rodriguez, L. P. Kaelbling, T. L. Perez, and P. Agrawal · 2022
Later among the works it cites.
Nerfstudio: A Modular Framework for Neural Radiance Field Development
M. Tancik, E. Weber, E. Ng, R. Li, B. Yi, J. Kerr, T. Wang, A. Kristoffersen, J. Austin, K. Salahi, A. Ahuja, D. McAllister, and A. Kanazawa · 2023
Closest in time.
When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou · 2023
Closest in time.
Language-driven representation learning for robotics
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang · 2023
Closest in time.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, et al · 2023
Closest in time.
PaLM-E: An Embodied Multimodal Language Model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Closest in time.
DINOv2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, R. Howes, P.-Y. Huang, H. Xu, V. Sharma, S.-W. Li, W. Galuba, M. Rabbat, M. Assran, N. Ballas, G. Synnaeve, I. Misra, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2023
Closest in time.
SE(3)-DiffusionFields: Learning smooth cost functions for joint grasp and motion optimization through diffusion
J. Urain, N. Funk, J. Peters, and G. Chalvatzaki · 2023
Closest in time.
Neural volumetric memory for visual locomotion control
R. Yang, G. Yang, and X. Wang · 2023
Closest in time.
GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF
Q. Dai, Y. Zhu, Y. Geng, C. Ruan, J. Zhang, and H. Wang · 2023
Closest in time.
OpenScene: 3D Scene Understanding with Open Vocabularies
S. Peng, K. Genova, C. M. Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser · 2023
Closest in time.
Clip-fields: Weakly supervised semantic fields for robotic memory
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam · 2023
Closest in time.
USA-Net: Unified Semantic and Affordance Representations for Robot Memory
B. Bolte, A. S. Wang, J. Yang, M. Mukadam, M. Kalakrishnan, and C. Paxton · 2023
Closest in time.
Visual Language Maps for Robot Navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Closest in time.
Conceptfusion: Open-set multimodal 3d mapping
K. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, J. Tenenbaum, C. de Melo, M. Krishna, L. Paull, F. Shkurti, and A. Torralba · 2023
Closest in time.
Lerf: Language embedded radiance fields
J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik · 2023
Closest in time.
Feature-Realistic Neural Fusion for Real-Time, Open Set Scene Understanding
K. Mazur, E. Sucar, and A. Davison · 2023
Closest in time.
Compressing volumetric radiance fields to 1 mb
L. Li, Z. Shen, Z. Wang, L. Shen, and L. Bo · 2023
Closest in time.