Fetching the paper…
Reading the bibliography…
We introduce Dream2Real, a robotics framework which integrates vision-language models (VLMs) trained on 2D data into a 3D object rearrangement pipeline.
B. Curless and M. Levoy, “A volumetric method for building complex models from range images,” in Proceedings of the 23rd annual conference on Computer graphics and interactive techniques , 1996, pp. 303–312
1996
Earlier work this paper cites.
J. Kuffner and S. LaValle, “RRT-connect: An efficient approach to single-query path planning,” in Proceedings 2000 ICRA. Millennium Conference. IEEE International Conference on Robotics and Automation. Symposia Proceedings (Cat. No.00CH37065) , vol. 2, 2000, pp. 995–1001 vol.2
2000
Earlier work this paper cites.
M. J. Schuster, D. Jain, M. Tenorth, and M. Beetz, “Learning organizational principles in human environments,” in International Conference on Robotics and Automation , 2012, pp. 3867–3874
2012
Earlier work this paper cites.
N. Abdo, C. Stachniss, L. Spinello, and W. Burgard, “Robot, organize my shelves! Tidying up objects by predicting user preferences,” in International Conference on Robotics and Automation , 2015
2015
Earlier work this paper cites.
2018
Earlier work this paper cites.
V. Satish, J. Mahler, and K. Goldberg, “On-policy dataset synthesis for learning robot grasping policies using fully convolutional deep networks,” IEEE Robotics and Automation Letters , 2019
2019
Earlier work this paper cites.
D. Batra, A. X. Chang, S. Chernova, A. J. Davison, J. Deng, V. Koltun, S. Levine, J. Malik, I. Mordatch, R. Mottaghi, M. Savva, and H. Su, “Rearrangement: A challenge for embodied AI,” arXiv , 2020
2020
Earlier work this paper cites.
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, and J. Lee, “Transporter networks: Rearranging the visual world for robotic manipulation,” Conference on Robot Learning (CoRL) , 2020
2020
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “CLIPort: What and where pathways for robotic manipulation,” in Conference on Robot Learning (CoRL) , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning, ICML , 2021
2021
Earlier work this paper cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
Earlier work this paper cites.
C. Paxton, C. Xie, T. Hermans, and D. Fox, “Predicting stable configurations for semantic placement of novel objects,” in Conference on Robot Learning, 8-11 November 2021, London, UK , ser. Proceedings of Machine Learning Research, vol. 164. PMLR, 2021, pp. 806–815
2021
Earlier work this paper cites.
I. Kapelyukh and E. Johns, “My house, my rules: Learning tidying preferences with graph neural networks,” in Conference on Robot Learning (CoRL) , 2021
2021
Earlier work this paper cites.
A. Yu, V. Ye, M. Tancik, and A. Kanazawa, “pixelNeRF: Neural radiance fields from one or few images,” in CVPR , 2021
2021
Earlier work this paper cites.
G. Sarch, Z. Fang, A. W. Harley, P. Schydlo, M. J. Tarr, S. Gupta, and K. Fragkiadaki, “TIDEE: Tidying up novel rooms using visuo-semantic commonsense priors,” in European Conference on Computer Vision , 2022
2022
Earlier work this paper cites.
Y. Lin, A. S. Wang, E. Undersander, and A. Rai, “Efficient and interpretable robot manipulation with graph neural networks,” IEEE Robotics and Automation Letters , vol. 7, pp. 2740–2747, 2022
2022
Earlier work this paper cites.
V. Jain, Y. Lin, E. Undersander, Y. Bisk, and A. Rai, “Transformers are adaptable task planners,” in Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
W. Liu, C. Paxton, T. Hermans, and D. Fox, “StructFormer: Learning spatial structure for language-guided semantic rearrangement of novel objects,” International Conference on Robotics and Automation , 2022
2022
Earlier work this paper cites.
Y. Kant, A. Ramachandran, S. Yenamandra, I. Gilitschenski, D. Batra, A. Szot, and H. Agrawal, “Housekeep: Tidying virtual households using commonsense reasoning,” arXiv , 2022
2022
Earlier work this paper cites.
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng, “Do as I can, not as I say: Grounding language in robotic affordances,” arXiv , 2022
2022
Earlier work this paper cites.
Z. Mandi, H. Bharadhwaj, V. Moens, S. Song, A. Rajeswaran, and V. Kumar, “CACTI: A framework for scalable multi-task multi-scene visual imitation learning,” arXiv , 2022
2022
Earlier work this paper cites.
Y. Cui, S. Niekum, A. Gupta, V. Kumar, and A. Rajeswaran, “Can foundation models perform zero-shot task specification for robot manipulation?” in Proceedings of The 4th Annual Learning for Dynamics and Control Conference , ser. Proceedings of Machine Learning Research, R. Firoozi, N. Mehr, E. Yel, R. Antonova, J. Bohg, M. Schwager, and M. Kochenderfer, Eds., vol. 168. PMLR, 23–24 Jun 2022, pp. 893–905
2022
Cited alongside, same era.
L. Yen-Chen, P. Florence, A. Zeng, J. T. Barron, Y. Du, W.-C. Ma, A. Simeonov, A. R. Garcia, and P. Isola, “MIRA: Mental imagery for robotic affordances,” in Conference on Robot Learning (CoRL) , 2022
2022
Cited alongside, same era.
T. Müller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” ACM Transactions on Graphics (ToG) , vol. 41, no. 4, pp. 1–15, 2022
2022
Cited alongside, same era.
C. Wang, M. Chai, M. He, D. Chen, and J. Liao, “Clip-nerf: Text-and-image driven manipulation of neural radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3835–3844
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, D. M, J. Peralta, B. Ichter, K. Hausman, and F. Xia, “Scaling robot learning with semantically imagined experience,” arXiv , 2023
2023
Closest in time.
Z. Chen, S. Kiami, A. Gupta, and V. Kumar, “GenAug: Retargeting behaviors to unseen situations via generative augmentation,” arXiv , 2023
2023
Closest in time.
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei, “Voxposer: Composable 3d value maps for robotic manipulation with language models,” in Conference on Robot Learning (CoRL) , 2023
2023
Closest in time.
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola, “Distilled feature fields enable few-shot language-guided manipulation,” arXiv preprint:2308.07931 , 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
A. Mirzaei, Y. Kant, J. Kelly, and I. Gilitschenski, “Laterf: Label and text driven object radiance fields,” in European Conference on Computer Vision . Springer, 2022, pp. 20–36
2022
Cited alongside, same era.
H. Ha and S. Song, “Semantic abstraction: Open-world 3D scene understanding from 2D vision-language models,” in Proceedings of the 2022 Conference on Robot Learning , 2022
2022
Cited alongside, same era.
H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” in European Conference on Computer Vision . Springer, 2022, pp. 640–658
2022
Cited alongside, same era.
K. Wada, S. James, and A. J. Davison, “ReorientBot: Learning object reorientation for specific-posed placement,” in IEEE International Conference on Robotics and Automation (ICRA) , 2022
2022
Cited alongside, same era.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich, “RT-2: Vision-language-action models transfer web knowledge to robotic control,” in arXiv , 2023
2023
Cited alongside, same era.
I. Kapelyukh, V. Vosylius, and E. Johns, “DALL-E-Bot: Introducing web-scale diffusion models to robotics,” IEEE Robotics and Automation Letters (RA-L) , 2023
2023
Cited alongside, same era.
K. Ramachandruni, M. Zuo, and S. Chernova, “Consor: A context-aware semantic object rearrangement framework for partially arranged scenes,” in 2023 IEEE International Conference on Intelligent Robots and Systems , 2023
2023
Cited alongside, same era.
Y. Zeng, M. Wu, L. Yang, J. Zhang, H. Ding, H. Cheng, and H. Dong, “Distilling functional rearrangement priors from large models,” 2023
2023
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
S. Sharma, A. Rashid, C. M. Kim, J. Kerr, L. Y. Chen, A. Kanazawa, and K. Goldberg, “Language embedded radiance fields for zero-shot task-oriented grasping,” in 7th Annual Conference on Robot Learning , 2023
2023
Closest in time.
D. Driess, Z. Huang, Y. Li, R. Tedrake, and M. Toussaint, “Learning multi-object dynamics with compositional neural radiance fields,” in Proceedings of The 6th Conference on Robot Learning , ser. Proceedings of Machine Learning Research, vol. 205. PMLR, 14–18 Dec 2023, pp. 1755–1768
2023
Closest in time.
X. Kong, S. Liu, M. Taher, and A. J. Davison, “vMAP: Vectorised object mapping for neural field slam,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 952–961
2023
Closest in time.
K. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, J. Tenenbaum, C. de Melo, M. Krishna, L. Paull, F. Shkurti, and A. Torralba, “Conceptfusion: Open-set multimodal 3d mapping,” Robotics: Science and Systems , 2023
2023
Closest in time.
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam, “Clip-fields: Weakly supervised semantic fields for robotic memory,” in Robotics: Science and Systems , 2023
2023
Closest in time.
2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” 2023
2023
Closest in time.
OpenAI, “GPT-4 technical report,” 2023
2023
Closest in time.
J. Urain, N. Funk, J. Peters, and G. Chalvatzaki, “SE(3)-DiffusionFields: Learning smooth cost functions for joint grasp and motion optimization through diffusion,” IEEE International Conference on Robotics and Automation (ICRA) , 2023
2023
Closest in time.
L. Melas-Kyriazi, C. Rupprecht, I. Laina, and A. Vedaldi, “Realfusion: 360° reconstruction of any object from a single image,” in Arxiv , 2023
2023
Closest in time.
R. Liu, R. Wu, B. V. Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick, “Zero-1-to-3: Zero-shot one image to 3D object,” in ICCV , 2023
2023
Closest in time.
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou, “When and why vision-language models behave like bags-of-words, and what to do about it?” in International Conference on Learning Representations , 2023
2023
Closest in time.
G. Zhai, X. Cai, D. Huang, Y. Di, F. Manhardt, F. Tombari, N. Navab, and B. Busam, “SG-Bot: Object rearrangement via coarse-to-fine robotic imagination on scene graphs,” in ICRA , 2024
2024
Closest in time.
M. Taher, I. Alzugaray, and A. J. Davison, “Fit-NGP: Fitting object models to neural graphics primitives,” in ICRA , 2024
2024
Closest in time.