Fetching the paper…
Reading the bibliography…
We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
A flexible and scalable slam system with full 3d motion estimation
S. Kohlbrecher, J. Meyer, O. von Stryk, and U. Klingauf · 2011
Earlier work this paper cites.
Scene representation networks: Continuous 3d-structure-aware neural scene representations
Vincent Sitzmann, Michael Zollhöfer, and Gordon Wetzstein · 2019
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner · 2020
Earlier work this paper cites.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge · 2020
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Earlier work this paper cites.
Seal: Self-supervised embodied active learning using exploration and 3d consistency
Devendra Singh Chaplot, Murtaza Dalal, Saurabh Gupta, Jitendra Malik, and Russ R Salakhutdinov · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2021
Earlier work this paper cites.
Film: Following instructions in language with modular methods
So Yeon Min, Devendra Singh Chaplot, Pradeep Ravikumar, Yonatan Bisk, and Ruslan Salakhutdinov · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
imap: Implicit mapping and positioning in real-time
Edgar Sucar, Shikun Liu, Joseph Ortiz, and Andrew J Davison · 2021
Earlier work this paper cites.
Nesf: Neural semantic fields for generalizable semantic segmentation of 3d scenes
Suhani Vora, Noha Radwan, Klaus Greff, Henning Meyer, Kyle Genova, Mehdi SM Sajjadi, Etienne Pot, Andrea Tagliasacchi, and Daniel Duckworth · 2021
Earlier work this paper cites.
Structured scene memory for vision-language navigation
Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen · 2021
Cited alongside, same era.
ilabel: Interactive neural scene labelling
Shuaifeng Zhi, Edgar Sucar, Andre Mouton, Iain Haughton, Tristan Laidlow, and Andrew J Davison · 2021
Cited alongside, same era.
Scanqa: 3d question answering for spatial scene understanding
Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoki Kawanabe · 2022
Cited alongside, same era.
A persistent spatial semantic representation for high-level natural language instruction execution
Valts Blukis, Chris Paxton, Dieter Fox, Animesh Garg, and Yoav Artzi · 2022
Cited alongside, same era.
” this is my unicorn, fluffy”: Personalizing frozen vision-language representations
Niv Cohen, Rinon Gal, Eli A Meirom, Gal Chechik, and Yuval Atzmon · 2022
Cited alongside, same era.
3d neural scene representations for visuomotor control
Yunzhu Li, Shuang Li, Vincent Sitzmann, Pulkit Agrawal, and Antonio Torralba · 2022
Closest in time.
Instant neural graphics primitives with a multiresolution hash encoding
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller · 2022
Closest in time.
isdf: Real-time neural signed distance fields for robot perception
Joseph Ortiz, Alexander Clegg, Jing Dong, Edgar Sucar, David Novotny, Michael Zollhoefer, and Mustafa Mukadam · 2022
Closest in time.
Language-grounded indoor 3d semantic segmentation in the wild
David Rozenberszki, Or Litany, and Angela Dai · 2022
Closest in time.
Cliport: What and where pathways for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Episodic memory question answering
Samyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai, Mukul Khanna, Dhruv Batra, and Devi Parikh · 2022
Cited alongside, same era.
Learning multi-object dynamics with compositional neural radiance fields
Danny Driess, Zhiao Huang, Yunzhu Li, Russ Tedrake, and Marc Toussaint · 2022
Cited alongside, same era.
Continuous scene representations for embodied ai
Samir Yitzhak Gadre, Kiana Ehsani, Shuran Song, and Roozbeh Mottaghi · 2022
Cited alongside, same era.
Navigating to objects in the real world
Theophile Gervet, Soumith Chintala, Dhruv Batra, Jitendra Malik, and Devendra Singh Chaplot · 2022
Cited alongside, same era.
Semantic abstraction: Open-world 3d scene understanding from 2d vision-language models
Huy Ha and Shuran Song · 2022
Cited alongside, same era.
Visual language maps for robot navigation
Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard · 2022
Cited alongside, same era.
Decomposing nerf for editing via feature field distillation
Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitzmann · 2022
Cited alongside, same era.
Neural descriptor fields: Se (3)-equivariant object representations for manipulation
Anthony Simeonov, Yilun Du, Andrea Tagliasacchi, Joshua B Tenenbaum, Alberto Rodriguez, Pulkit Agrawal, and Vincent Sitzmann · 2022
Closest in time.
Language grounding with 3d objects
Jesse Thomason, Mohit Shridhar, Yonatan Bisk, Chris Paxton, and Luke Zettlemoyer · 2022
Closest in time.
Neural feature fusion fields: 3d distillation of self-supervised 2d image representations
Vadim Tschernezki, Iro Laina, Diane Larlus, and Andrea Vedaldi · 2022
Closest in time.
Neural fields in visual computing and beyond
Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar · 2022
Closest in time.
Habitat-matterport 3d semantics dataset
Karmesh Yadav, Ram Ramrakhya, Santhosh Kumar Ramakrishnan, Theo Gervet, John Turner, Aaron Gokaslan, Noah Maestre, Angel Xuan Chang, Dhruv Batra, Manolis Savva, et al · 2022
Closest in time.
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krähenbühl, and Ishan Misra · 2022
Closest in time.