Fetching the paper…
Reading the bibliography…
Attention-based models such as transformers have shown outstanding performance on dense prediction tasks, such as semantic segmentation, owing to their capability of capturing long-range dependency in an image.
“ImageNet: A Large-Scale Hierarchical Image Database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Indoor Segmentation And Support Inference From RGBD Images,”
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus, · 2012
Earlier work this paper cites.
“Vision Meets Robotics: The KITTI Dataset,”
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun, · 2013
Earlier work this paper cites.
“Depthmap Prediction From A Single Image Using A Multi-scale Deep network,”
David Eigen, Christian Puhrsch, and Rob Fergus, · 2014
Earlier work this paper cites.
“Learning Depth From Single Monocular Images Using Deep Convolutional Neural Fields,”
Fayao Liu, Chunhua Shen, Guosheng Lin, and Ian Reid, · 2016
Earlier work this paper cites.
“Unsupervised CNN For Single View Depth Estimation: Geometry To The Rescue,”
Ravi Garg, Vijay Kumar B.G., Gustavo Carneiro, and Ian Reid, · 2016
Earlier work this paper cites.
“Deep Compositional Captioning: Describing Novel Object Categories Without Paired Training Data,”
Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, and Trevor Darrell, · 2016
Earlier work this paper cites.
“Unsupervised Video Summarization With Adversarial LSTM Networks,”
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic, · 2017
Earlier work this paper cites.
“Semi-Supervised Deep Learning For Monocular Depth Map Prediction,”
Yevhen Kuznietsov, Jorg Stuckler, and Bastian Leibe, · 2017
Cited alongside, same era.
“Feature pyramid networks for object detection,”
Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie, · 2017
Cited alongside, same era.
“Deep Ordinal Regression Network For Monocular Depth Estimation,”
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao, · 2018
Cited alongside, same era.
“Monocular Depth Estimation With Affinity, Vertical Pooling, And Label Enhancement,”
Yukang Gan, Xiangyu Xu, Wenxiu Sun, and Liang Lin, · 2018
Cited alongside, same era.
“Enforcing Geometric Constraints Of Virtual Normal For Depth Prediction,”
Wei Yin, Yifan Liu, Chunhua Shen, and Youliang Yan, · 2019
Cited alongside, same era.
“Decoupled Weight Decay Regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
“SegFormer: Simple and Efficient Design For Semantic Segmentation With Transformers,”
Enze Xie, · 2021
Later among the works it cites.
“Pyramid Vision Transformer: A Versatile Backbone For Dense Prediction Without Convolutions,”
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao, · 2021
Later among the works it cites.
“An Image Is Worth 16x16 Words: Transformers For Image Recognition At Scale,”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby, · 2021
Later among the works it cites.
“Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,”
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo, · 2021
Later among the works it cites.
“Twins: Revisiting The Design Of Spatial Attention In Vision Transformers ,”
Xiangxiang Chu, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Guiding Monocular Depth Estimation Using Depth-Attention Volume,”
Lam Huynh, Phong Nguyen-Ha, Jiri Matas, Esa Rahtu, and Janne Heikkilä, · 2020
Cited alongside, same era.
“End-to-End Object Detection with Transformers,”
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko, · 2020
Cited alongside, same era.
“From Big To small: Multi-scale Local Planar Guidance For Monocular Depth Estimation,”
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, and Il Hong Suh,
Cited in the paper.
“Structure-Aware Residual Pyramid Network For Monocular Depth Estimation,”
Xiaotian Chen, Xuejin Chen, and Zheng-Jun Zha,
Cited in the paper.
“Super-Convergence: Very Fast Training of Residual Networks Using Large Learning Rates,”
Leslie N. Smith and Nicholay Topin,
Cited in the paper.
“Vision Transformers For Dense Prediction,”
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun, · 2021
Later among the works it cites.
“AdaBins: Depth Estimation Using Adaptive Bins,”
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka, · 2021
Later among the works it cites.
“MPViT: Multi-Path Vision Transformer For Dense Prediction,”
Youngwan Lee, Jonghee Kim, Jeffrey Willette, and Sung Ju Hwang, · 2022
Closest in time.