Fetching the paper…
Reading the bibliography…
Monocular depth estimation (MDE) is inherently ambiguous, as a given image may result from many different 3D scenes and vice versa.
Monocular Outdoor Semantic Mapping with a Multi-task Network
Yucai Bai, Lei Fan, Ziyu Pan, and Long Chen · 1901
Earlier work this paper cites.
SharpNet: Fast and Accurate Recovery of Occluding Contours in Monocular Depth Estimation
Michaël Ramamonjisoa and Vincent Lepetit · 1905
Earlier work this paper cites.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Mingxing Tan and Quoc V. Le · 1905
Earlier work this paper cites.
Allyson Ettinger · 1907
Earlier work this paper cites.
From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, and Il Hong Suh · 1907
Earlier work this paper cites.
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 1908
Earlier work this paper cites.
Chameleons use accommodation cues to judge distance
Lindesay Harkness · 1977
Earlier work this paper cites.
Barn owls (Tyto alba) use accommodation as a distance cue
Hermann Wagner and Frank Schaeffel · 1991
Earlier work this paper cites.
Pictorial Cues, Oculomotor Adjustments, Automatic Organizing Processes, and Observer Tendencies
Maurice Hershenson · 1998
Earlier work this paper cites.
Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition
Erik F. Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Language Models are Open Knowledge Graphs
Chenguang Wang, Xiao Liu, and Dawn Song · 2010
Earlier work this paper cites.
AdaBins: Depth Estimation using Adaptive Bins
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2011
Earlier work this paper cites.
Judging an unfamiliar object’s distance from its retinal image size
Rita Sousa, Eli Brenner, and Jeroen B. J. Smeets · 2011
Cited alongside, same era.
Depth Perception from Image Defocus in a Jumping Spider
Takashi Nagata, Mitsumasa Koyanagi, Hisao Tsukamoto, Shinjiro Saeki, Kunio Isono, Yoshinori Shichida, Fumio Tokunaga, Michiyo Kinoshita, Kentaro Arikawa, and Akihisa Terakita · 2012
Cited alongside, same era.
Efficient Estimation of Word Representations in Vector Space, Sept. 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Cited alongside, same era.
Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Cited alongside, same era.
Glove: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
A Structural Probe for Finding Syntax in Word Representations
John Hewitt and Christopher D Manning · 2019
Later among the works it cites.
On Measuring Social Biases in Sentence Encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger · 2019
Later among the works it cites.
Language Models as Knowledge Bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture
David Eigen and Rob Fergus · 2015
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Scene Parsing Through ADE20K Dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Cited alongside, same era.
Deep Ordinal Regression Network for Monocular Depth Estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2018
Cited alongside, same era.
Look Deeper into Depth: Monocular Depth Estimation with Semantic Booster and Attention-Driven Loss
Jianbo Jiao, Ying Cao, Yibing Song, and Rynson Lau · 2018
Cited alongside, same era.
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Monocular Depth Estimation Using Cues Inspired by Biological Vision Systems
Dylan Auty and Krystian Mikolajczyk · 2022
Later among the works it cites.
What do Toothbrushes do in the Kitchen? How Transformers Think our World is Structured
Alexander Henlein and Alexander Mehler · 2022
Later among the works it cites.
Zhenyu Li, Zehui Chen, Xianming Liu, and Junjun Jiang · 2022
Later among the works it cites.
BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation
Zhenyu Li, Xuyang Wang, Xianming Liu, and Junjun Jiang · 2022
Later among the works it cites.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Can Language Understand Depth?
Renrui Zhang, Ziyao Zeng, and Ziyu Guo · 2022
Later among the works it cites.