Fetching the paper…
Reading the bibliography…
Most modern deep learning-based multi-view 3D reconstruction techniques use RNNs or fusion modules to combine information from multiple images after independently encoding them.
Three-mode factor analysis
Joseph Levin · 1965
Earlier work this paper cites.
Some mathematical notes on three-mode factor analysis
Ledyard R Tucker · 1966
Earlier work this paper cites.
Foundations of the parafac procedure: Models and conditions for an” explanatory” multimodal factor analysis
Richard A Harshman et al · 1970
Earlier work this paper cites.
Distance field compression
Mark W Jones · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Matrix factorization techniques for recommender systems
Yehuda Koren, Robert Bell, and Chris Volinsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Real-time 3d reconstruction at scale using voxel hashing
Matthias Nießner, Michael Zollhöfer, Shahram Izadi, and Marc Stamminger · 2013
Earlier work this paper cites.
Octree-based fusion for realtime 3d reconstruction
Ming Zeng, Fukai Zhao, Jiaxiang Zheng, and Xinguo Liu · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Visual simultaneous localization and mapping: a survey
Jorge Fuentes-Pacheco, José Ruiz-Ascencio, and Juan Manuel Rendón-Mancha · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
3d-r2n2: A unified approach for single and multi-view 3d object reconstruction
Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Learning a multi-view stereo machine
Abhishek Kar, Christian Häne, and Jitendra Malik · 2017
Cited alongside, same era.
Tanks and temples: Benchmarking large-scale scene reconstruction
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun · 2017
Cited alongside, same era.
A survey of structure from motion
Onur Ozyesil, Vladislav Voroninski, Ronen Basri, and Amit Singer · 2017
Cited alongside, same era.
Octree generating networks: Efficient convolutional architectures for high-resolution 3d outputs
Maxim Tatarchenko, Alexey Dosovitskiy, and Thomas Brox · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
What do single-view 3d reconstruction networks learn?
Maxim Tatarchenko, Stephan R Richter, René Ranftl, Zhuwen Li, Vladlen Koltun, and Thomas Brox · 2019
Later among the works it cites.
Pix2vox: Context-aware 3d reconstruction from single and multi-view images
Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, and Shengping Zhang · 2019
Later among the works it cites.
Etc: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanón, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang · 2020
Later among the works it cites.
Tt-tsdf: Memory-efficient tsdf with low-rank tensor train decomposition
Alexey I Boyko, Mikhail P Matrosov, Ivan V Oseledets, Dzmitry Tsetserukou, and Gonzalo Ferrer · 2020
Later among the works it cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A papier-mâché approach to learning 3d surface generation
Thibault Groueix, Matthew Fisher, Vladimir G Kim, Bryan C Russell, and Mathieu Aubry · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Cited alongside, same era.
Matryoshka networks: Predicting 3d geometry via nested shape layers
Stephan R Richter and Stefan Roth · 2018
Cited alongside, same era.
Tensormachine: probabilistic boolean tensor decomposition
Tammo Rukat, Chris C Holmes, and Christopher Yau · 2018
Cited alongside, same era.
Pix3d: Dataset and methods for single-image 3d shape modeling
Xingyuan Sun, Jiajun Wu, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Joshua B Tenenbaum, and William T Freeman · 2018
Cited alongside, same era.
Pixel2mesh: Generating 3d mesh models from single rgb images
Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang · 2018
Cited alongside, same era.
Learning implicit fields for generative shape modeling
Zhiqin Chen and Hao Zhang · 2019
Cited alongside, same era.
Implicit functions in feature space for 3d shape reconstruction and completion
Julian Chibane, Thiemo Alldieck, and Gerard Pons-Moll · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Later among the works it cites.
Recent developments in boolean matrix factorization
Pauli Miettinen and Stefan Neumann · 2020
Later among the works it cites.
A divide et impera approach for 3d shape reconstruction from multiple views
Riccardo Spezialetti, David Joseph Tan, Alessio Tonioni, Keisuke Tateno, and Federico Tombari · 2020
Later among the works it cites.
Geometric all-way boolean tensor decomposition
Changlin Wan, Wennan Chang, Tong Zhao, Sha Cao, and Chi Zhang · 2020
Later among the works it cites.
Pix2vox++: multi-scale context-aware 3d object reconstruction from single and multiple images
Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, and Wenxiu Sun · 2020
Later among the works it cites.
Robust attentional aggregation of deep feature sets for multi-view 3d reconstruction
Bo Yang, Sen Wang, Andrew Markham, and Niki Trigoni · 2020
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Levit: a vision transformer in convnet’s clothing for faster inference
Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze · 2021
Closest in time.
Multi-view 3d reconstruction with transformers
Dan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou, Tianyang Shi, Septimiu Salcudean, Z Jane Wang, and Rabab Ward · 2021
Closest in time.