Fetching the paper…
Reading the bibliography…
Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success.
A solution for the best rotation to relate two sets of vectors
W. Kabsch · 1976
Earlier work this paper cites.
Least-squares estimation of transformation parameters between two point patterns
S. Umeyama · 1991
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
CIDER: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Learning visual features from large weakly supervised data
A. Joulin, L. Van Der Maaten, A. Jabri, and N. Vasilache · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Earlier work this paper cites.
Structure-from-motion revisited
J. L. Schonberger and J.-M. Frahm · 2016
Earlier work this paper cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Earlier work this paper cites.
Matterport3D: Learning from RGB-D data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
ScanNet: Richly-annotated 3D reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
S. Gidaris, P. Singh, and N. Komodakis · 2018
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. Van Der Maaten · 2018
Earlier work this paper cites.
ReferIt3D: Neural listeners for fine-grained 3D object identification in real-world scenes
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. Guibas · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
PointContrast: Unsupervised pre-training for 3D point cloud understanding
S. Xie, J. Gu, D. Guo, C. R. Qi, L. Guibas, and O. Litany · 2020
Earlier work this paper cites.
VATT: Transformers for multimodal self-supervised learning from raw video, audio and text
H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
Exploring simple siamese representation learning
X. Chen and K. He · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Exploring data-efficient 3D scene understanding with contrastive scene contexts
J. Hou, B. Graham, M. Nießner, and S. Xie · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
HyperSim: A photorealistic synthetic dataset for holistic indoor scene understanding
M. Roberts, J. Ramapuram, A. Ranjan, A. Kumar, M. A. Bautista, N. Paczan, R. Webb, and J. M. Susskind · 2021
Earlier work this paper cites.
VideoCLIP: Contrastive pre-training for zero-shot video-text understanding
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer · 2021
Earlier work this paper cites.
Self-supervised pretraining of 3D features on any point-cloud
Z. Zhang, R. Girdhar, A. Joulin, and I. Misra · 2021
Earlier work this paper cites.
ScanQA: 3D question answering for spatial scene understanding
D. Azuma, T. Miyanishi, S. Kurita, and M. Kawanabe · 2022
Cited alongside, same era.
BeiT: Bert pre-training of image transformers
H. Bao, L. Dong, S. Piao, and F. Wei · 2022
Cited alongside, same era.
Label-efficient semantic segmentation with diffusion models
D. Baranchuk, I. Rubachev, A. Voynov, V. Khrulkov, and A. Babenko · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Cited alongside, same era.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Y. LeCun · 2022
Cited alongside, same era.
Masked discrimination for self-supervised learning on point clouds
H. Liu, M. Cai, and Y. J. Lee · 2022
Cited alongside, same era.
BEV-guided multi-modality fusion for driving perception
Y. Man, L.-Y. Gui, and Y.-X. Wang · 2023
Later among the works it cites.
OpenScene: 3D scene understanding with open vocabularies
S. Peng, K. Genova, C. M. Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser · 2023
Later among the works it cites.
Monocular depth estimation using diffusion models
S. Saxena, A. Kar, M. Norouzi, and D. J. Fleet · 2023
Later among the works it cites.
DriveLM: Driving with graph visual question answering
C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, P. Luo, A. Geiger, and H. Li · 2023
Later among the works it cites.
Emergent correspondence from image diffusion
L. Tang, M. Jia, Q. Wang, C. P. Phoo, and B. Hariharan · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Masked autoencoders for point cloud self-supervised learning
Y. Pang, W. Wang, F. E. Tay, W. Liu, Y. Tian, and L. Yuan · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Laion-5B: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev · 2022
Cited alongside, same era.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Z. Tong, Y. Song, J. Wang, and L. Wang · 2022
Cited alongside, same era.
InternVideo: General video foundation models via generative and discriminative learning
Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wang, S. Xing, G. Chen, J. Pan, J. Yu, Y. Wang, L. Wang, and Y. Qiao · 2022
Cited alongside, same era.
REGTR: End-to-end point cloud correspondences with transformers
Z. J. Yew and G. H. Lee · 2022
Cited alongside, same era.
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Open-vocabulary panoptic segmentation with text-to-image diffusion models
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello · 2023
Later among the works it cites.
Swin3D: A pretrained transformer backbone for 3D indoor scene understanding
Y.-Q. Yang, Y.-X. Guo, J.-Y. Xiong, Y. Liu, H. Pan, P.-S. Wang, X. Tong, and B. Guo · 2023
Later among the works it cites.
ScanNet++: A high-fidelity dataset of 3D indoor scenes
C. Yeshwanth, Y.-C. Liu, M. Nießner, and A. Dai · 2023
Later among the works it cites.
What does stable diffusion know about the 3d scene?
G. Zhan, C. Zheng, W. Xie, and A. Zisserman · 2023
Later among the works it cites.
Unleashing text-to-image diffusion models for visual perception
W. Zhao, Y. Rao, Z. Liu, B. Liu, J. Zhou, and J. Lu · 2023
Later among the works it cites.
3D-VisTA: Pre-trained transformer for 3D vision and text alignment
Z. Zhu, X. Ma, Y. Chen, Z. Deng, S. Huang, and Q. Li · 2023
Later among the works it cites.
Probing the 3D awareness of visual foundation models
M. E. Banani, A. Raj, K.-K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani · 2024
Closest in time.
V-JEPA: Latent video prediction for visual representation learning
A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y. LeCun, M. Assran, and N. Ballas · 2024
Closest in time.
GraphDreamer: Compositional 3D scene synthesis from scene graphs
G. Gao, W. Liu, A. Chen, A. Geiger, and B. Schölkopf · 2024
Closest in time.
MultiPLY: A multisensory object-centric embodied large language model in 3D world
Y. Hong, Z. Zheng, P. Chen, Y. Wang, J. Li, and C. Gan · 2024
Closest in time.
OpenShape: Scaling up 3D shape representation towards open-world understanding
M. Liu, R. Shi, K. Kuang, Y. Zhu, X. Li, S. Han, H. Cai, F. Porikli, and H. Su · 2024
Closest in time.
Situational awareness matters in 3D vision language reasoning
Y. Man, L.-Y. Gui, and Y.-X. Wang · 2024
Closest in time.
EmerDiff: Emerging pixel-level semantic knowledge in diffusion models
K. Namekata, A. Sabour, S. Fidler, and S. W. Kim · 2024
Closest in time.
PIVOT: Iterative visual prompting elicits actionable knowledge for VLMs
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, Q. Vuong, T. Zhang, T.-W. E. Lee, K.-H. Lee, P. Xu, S. Kirmani, Y. Zhu, A. Zeng, K. Hausman, N. Heess, C. Finn, S. Levine, and B. Ichter · 2024
Closest in time.
DINOv2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2024
Closest in time.
Frozen transformers in language models are effective visual encoder layers
Z. Pang, Z. Xie, Y. Man, and Y.-X. Wang · 2024
Closest in time.
Kosmos-2: Grounding multimodal large language models to the world
Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, Q. Ye, and F. Wei · 2024
Closest in time.
AM-RADIO: Agglomerative model–reduce all domains into one
M. Ranzinger, G. Heinrich, J. Kautz, and P. Molchanov · 2024
Closest in time.
GLaMM: Pixel grounding large multimodal model
H. Rasheed, M. Maaz, S. Shaji, A. Shaker, S. Khan, H. Cholakkal, R. M. Anwer, E. Xing, M.-H. Yang, and F. S. Khan · 2024
Closest in time.
ControlRoom3D: Room generation using semantic proxy rooms
J. Schult, S. Tsai, L. Höllein, B. Wu, J. Wang, C.-Y. Ma, K. Li, X. Wang, F. Wimbauer, Z. He, P. Zhang, B. Leibe, P. Vajda, and J. Hou · 2024
Closest in time.
DriveVLM: The convergence of autonomous driving and large vision-language models
X. Tian, J. Gu, B. Li, Y. Liu, C. Hu, Y. Wang, K. Zhan, P. Jia, X. Lang, and H. Zhao · 2024
Closest in time.
Pixel aligned language models
J. Xu, X. Zhou, S. Yan, X. Gu, A. Arnab, C. Sun, X. Wang, and C. Schmid · 2024
Closest in time.
3D feature prediction for masked-autoencoder-based point cloud pretraining
S. Yan, Y. Yang, Y. Guo, H. Pan, P.-s. Wang, X. Tong, Y. Liu, and Q. Huang · 2024
Closest in time.
SceneCraft: Layout-guided 3D scene generation
X. Yang, Y. Man, J.-K. Chen, and Y.-X. Wang · 2024
Closest in time.
LLaMA-Adapter: Efficient fine-tuning of language models with zero-init attention
R. Zhang, J. Han, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, P. Gao, and Y. Qiao · 2024
Closest in time.
3D-VLA: A 3D vision-language-action generative world model
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan · 2024
Closest in time.