Fetching the paper…
Reading the bibliography…
We present a unified formulation and model for three motion and 3D perception tasks: optical flow, rectified stereo matching and unrectified stereo depth estimation from posed images.
H. Xu and J. Zhang, “Aanet: Adaptive aggregation network for efficient stereo matching,” in CVPR , 2020, pp. 1959–1968
1968
Earlier work this paper cites.
B. K. Horn and B. G. Schunck, “Determining optical flow,” Artificial intelligence , vol. 17, no. 1-3, pp. 185–203, 1981
1981
Earlier work this paper cites.
Z. Yin and J. Shi, “Geonet: Unsupervised learning of dense depth, optical flow and camera pose,” in CVPR , 2018, pp. 1983–1992
1992
Earlier work this paper cites.
R. T. Collins, “A space-sweep approach to true multi-image matching,” in CVPR , 1996, pp. 358–363
1996
Earlier work this paper cites.
B. Triggs, P. F. McLauchlan, R. I. Hartley, and A. W. Fitzgibbon, “Bundle adjustment—a modern synthesis,” in International workshop on vision algorithms . Springer, 1999, pp. 298–372
1999
Earlier work this paper cites.
D. Scharstein and R. Szeliski, “A taxonomy and evaluation of dense two-frame stereo correspondence algorithms,” IJCV , vol. 47, no. 1, pp. 7–42, 2002
2002
Earlier work this paper cites.
R. Hartley and A. Zisserman, Multiple view geometry in computer vision . Cambridge university press, 2003
2003
Earlier work this paper cites.
T. Brox, A. Bruhn, N. Papenberg, and J. Weickert, “High accuracy optical flow estimation based on a theory for warping,” in ECCV . Springer, 2004, pp. 25–36
2004
Earlier work this paper cites.
A. Bruhn, J. Weickert, and C. Schnörr, “Lucas/kanade meets horn/schunck: Combining local and global optic flow methods,” IJCV , vol. 61, no. 3, pp. 211–231, 2005
2005
Earlier work this paper cites.
K.-J. Yoon and I. S. Kweon, “Adaptive support-weight approach for correspondence search,” TPAMI , vol. 28, no. 4, pp. 650–656, 2006
2006
Earlier work this paper cites.
G. Klein and D. Murray, “Parallel tracking and mapping for small ar workspaces,” in ISMAR . IEEE, 2007, pp. 225–234
2007
Earlier work this paper cites.
H. Hirschmuller, “Stereo processing by semiglobal matching and mutual information,” TPAMI , vol. 30, no. 2, pp. 328–341, 2007
2007
Earlier work this paper cites.
H. Hirschmuller and D. Scharstein, “Evaluation of cost functions for stereo matching,” in CVPR . IEEE, 2007, pp. 1–8
2007
Earlier work this paper cites.
D. Sun, S. Roth, and M. J. Black, “Secrets of optical flow estimation and their principles,” in CVPR . IEEE, 2010, pp. 2432–2439
2010
Earlier work this paper cites.
T. Brox and J. Malik, “Large displacement optical flow: descriptor matching in variational motion estimation,” TPAMI , vol. 33, no. 3, pp. 500–513, 2010
2010
Earlier work this paper cites.
S. Agarwal, Y. Furukawa, N. Snavely, I. Simon, B. Curless, S. M. Seitz, and R. Szeliski, “Building rome in a day,” Communications of the ACM , vol. 54, no. 10, pp. 105–112, 2011
2011
Earlier work this paper cites.
L. Xu, J. Jia, and Y. Matsushita, “Motion detail preserving optical flow estimation,” TPAMI , vol. 34, no. 9, pp. 1744–1757, 2011
2011
Earlier work this paper cites.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in CVPR . IEEE, 2012, pp. 3354–3361
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” NIPS , vol. 25, 2012
2012
Earlier work this paper cites.
D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black, “A naturalistic open source movie for optical flow evaluation,” in ECCV . Springer, 2012, pp. 611–625
2012
Earlier work this paper cites.
J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of rgb-d slam systems,” in 2012 IEEE/RSJ international conference on intelligent robots and systems . IEEE, 2012, pp. 573–580
2012
Earlier work this paper cites.
A. Hosni, C. Rhemann, M. Bleyer, C. Rother, and M. Gelautz, “Fast cost-volume filtering for visual correspondence and beyond,” TPAMI , vol. 35, no. 2, pp. 504–511, 2012
2012
Earlier work this paper cites.
J. Xiao, A. Owens, and A. Torralba, “Sun3d: A database of big spaces reconstructed using sfm and object labels,” in ICCV , 2013, pp. 1625–1632
2013
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” NIPS , vol. 27, 2014
2014
Earlier work this paper cites.
D. Scharstein, H. Hirschmüller, Y. Kitajima, G. Krathwohl, N. Nešić, X. Wang, and P. Westling, “High-resolution stereo datasets with subpixel-accurate ground truth,” in GCPR . Springer, 2014, pp. 31–42
2014
Earlier work this paper cites.
A. Dosovitskiy, P. Fischer, E. Ilg, P. Hausser, C. Hazirbas, V. Golkov, P. Van Der Smagt, D. Cremers, and T. Brox, “Flownet: Learning optical flow with convolutional networks,” in ICCV , 2015, pp. 2758–2766
2015
Earlier work this paper cites.
M. Menze and A. Geiger, “Object scene flow for autonomous vehicles,” in CVPR , 2015, pp. 3061–3070
2015
Earlier work this paper cites.
J. Revaud, P. Weinzaepfel, Z. Harchaoui, and C. Schmid, “Epicflow: Edge-preserving interpolation of correspondences for optical flow,” in CVPR , 2015, pp. 1164–1172
2015
Earlier work this paper cites.
J. Zbontar and Y. LeCun, “Stereo matching by training a convolutional neural network to compare image patches,” J. Mach. Learn. Res. , vol. 17, pp. 65:1–65:32, 2016
2016
Earlier work this paper cites.
N. Mayer, E. Ilg, P. Hausser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox, “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” in CVPR , 2016, pp. 4040–4048
2016
Earlier work this paper cites.
R. Garg, V. K. Bg, G. Carneiro, and I. Reid, “Unsupervised cnn for single view depth estimation: Geometry to the rescue,” in ECCV . Springer, 2016, pp. 740–756
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” in CVPR , 2016, pp. 4104–4113
2016
Earlier work this paper cites.
D. Kondermann, R. Nair, K. Honauer, K. Krispin, J. Andrulis, A. Brock, B. Gussefeld, M. Rahimimoghaddam, S. Hofmann, C. Brenner et al. , “The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving,” in CVPR Workshops , 2016, pp. 19–28
2016
Earlier work this paper cites.
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox, “Flownet 2.0: Evolution of optical flow estimation with deep networks,” in CVPR , 2017, pp. 2462–2470
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
T. Schops, J. L. Schonberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” in CVPR , 2017, pp. 3260–3269
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in CVPR , 2017, pp. 5828–5839
2017
Earlier work this paper cites.
B. Ummenhofer, H. Zhou, J. Uhrig, N. Mayer, E. Ilg, A. Dosovitskiy, and T. Brox, “Demon: Depth and motion network for learning monocular stereo,” in CVPR , 2017, pp. 5038–5047
2017
Earlier work this paper cites.
A. Ranjan and M. J. Black, “Optical flow estimation using a spatial pyramid network,” in CVPR , 2017, pp. 4161–4170
2017
Cited alongside, same era.
A. Kendall, H. Martirosyan, S. Dasgupta, P. Henry, R. Kennedy, A. Bachrach, and A. Bry, “End-to-end learning of geometry and context for deep stereo regression,” in ICCV , 2017, pp. 66–75
2017
Cited alongside, same era.
C. Godard, O. Mac Aodha, and G. J. Brostow, “Unsupervised monocular depth estimation with left-right consistency,” in CVPR , 2017, pp. 270–279
2017
Cited alongside, same era.
T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, “Unsupervised learning of depth and ego-motion from video,” in CVPR , 2017, pp. 1851–1858
2017
Cited alongside, same era.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in CVPR , 2017, pp. 2117–2125
X. Cheng, Y. Zhong, M. Harandi, Y. Dai, X. Chang, H. Li, T. Drummond, and Z. Ge, “Hierarchical neural architecture search for deep stereo matching,” NeurIPS , vol. 33, pp. 22 158–22 169, 2020
2020
Later among the works it cites.
W. Wang, D. Zhu, X. Wang, Y. Hu, Y. Qiu, C. Wang, Y. Hu, A. Kapoor, and S. Scherer, “Tartanair: A dataset to push the limits of visual slam,” 2020
2020
Later among the works it cites.
W. Bao, W. Wang, Y. Xu, Y. Guo, S. Hong, and X. Zhang, “Instereo2k: A large real dataset for stereo matching in indoor scenes,” Science China Information Sciences , vol. 63, no. 11, pp. 1–11, 2020
2020
Later among the works it cites.
L. Lipson, Z. Teed, and J. Deng, “Raft-stereo: Multilevel recurrent field transforms for stereo matching,” in 3DV . IEEE, 2021, pp. 218–227
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Yao, Z. Luo, S. Li, T. Fang, and L. Quan, “Mvsnet: Depth inference for unstructured multi-view stereo,” in ECCV , 2018, pp. 767–783
2018
Cited alongside, same era.
D. Sun, X. Yang, M.-Y. Liu, and J. Kautz, “Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,” in CVPR , 2018, pp. 8934–8943
2018
Cited alongside, same era.
J.-R. Chang and Y.-S. Chen, “Pyramid stereo matching network,” in CVPR , 2018, pp. 5410–5418
2018
Cited alongside, same era.
Y. Zou, Z. Luo, and J.-B. Huang, “Df-net: Unsupervised joint learning of depth and flow using cross-task consistency,” in ECCV , 2018, pp. 36–53
2018
Cited alongside, same era.
S. Meister, J. Hur, and S. Roth, “Unflow: Unsupervised learning of optical flow with a bidirectional census loss,” in AAAI , 2018
2018
Cited alongside, same era.
P.-H. Huang, K. Matzen, J. Kopf, N. Ahuja, and J.-B. Huang, “Deepmvs: Learning multi-view stereopsis,” in CVPR , 2018, pp. 2821–2830
2018
Cited alongside, same era.
2021
Later among the works it cites.
H. Xu, J. Yang, J. Cai, J. Zhang, and X. Tong, “High-resolution optical flow from 1d attention and correlation,” in ICCV , 2021, pp. 10 498–10 507
2021
Later among the works it cites.
D. Sun, D. Vlasic, C. Herrmann, V. Jampani, M. Krainin, H. Chang, R. Zabih, W. T. Freeman, and C. Liu, “Autoflow: Learning a better training set for optical flow,” in CVPR , 2021, pp. 10 093–10 102
2021
Later among the works it cites.
F. Zhang, O. J. Woodford, V. A. Prisacariu, and P. H. Torr, “Separable flow: Learning motion cost volumes for optical flow estimation,” in ICCV , 2021, pp. 10 807–10 817
2021
Later among the works it cites.
Z. Shen, Y. Dai, and Z. Rao, “Cfnet: Cascade and fused cost volume for robust stereo matching,” in CVPR , 2021, pp. 13 906–13 915
2021
Later among the works it cites.
M. Poggi, F. Tosi, K. Batsos, P. Mordohai, and S. Mattoccia, “On the synergies between machine learning and binocular stereo for depth estimation from images: a survey,” TPAMI , vol. 44, no. 9, pp. 5314–5334, 2021
2021
Later among the works it cites.
Z. Li, X. Liu, N. Drenkow, A. Ding, F. X. Creighton, R. H. Taylor, and M. Unberath, “Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,” in ICCV , 2021, pp. 6197–6206
2021
Later among the works it cites.
J. Watson, O. Mac Aodha, V. Prisacariu, G. Brostow, and M. Firman, “The temporal opportunist: Self-supervised multi-frame monocular depth,” in CVPR , 2021, pp. 1164–1174
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” ICCV , 2021
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2021
2021
Later among the works it cites.
S. d’Ascoli, H. Touvron, M. Leavitt, A. Morcos, G. Biroli, and L. Sagun, “Convit: Improving vision transformers with soft convolutional inductive biases,” ICML , 2021
2021
Later among the works it cites.
Y. Xu, Q. Zhang, J. Zhang, and D. Tao, “Vitae: Vision transformer advanced by exploring intrinsic inductive bias,” NeurIPS , 2021
2021
Later among the works it cites.
T. Ke, T. Do, K. Vuong, K. Sartipi, and S. I. Roumeliotis, “Deep multi-view depth estimation with predicted uncertainty,” in ICRA . IEEE, 2021, pp. 9235–9241
2021
Later among the works it cites.
G. Xu, J. Cheng, P. Guo, and X. Yang, “Attention concatenation volume for accurate and efficient stereo matching,” in CVPR , 2022, pp. 12 981–12 990
2022
Closest in time.
A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer et al. , “Perceiver io: A general architecture for structured inputs & outputs,” in ICLR , 2022
2022
Closest in time.
W. Yifan, C. Doersch, R. Arandjelović, J. Carreira, and A. Zisserman, “Input-level inductive biases for 3d reconstruction,” in CVPR , 2022, pp. 6176–6186
2022
Closest in time.
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, and D. Tao, “Gmflow: Learning optical flow via global matching,” in CVPR , 2022, pp. 8121–8130
2022
Closest in time.
S. Bai, Z. Geng, Y. Savani, and J. Z. Kolter, “Deep equilibrium optical flow estimation,” in CVPR , 2022, pp. 620–630
2022
Closest in time.
A. Luo, F. Yang, X. Li, and S. Liu, “Learning optical flow with kernel patch attention,” in CVPR , 2022, pp. 8906–8915
2022
Closest in time.
X. Sui, S. Li, X. Geng, Y. Wu, X. Xu, Y. Liu, R. Goh, and H. Zhu, “Craft: Cross-attentional flow transformer for robust optical flow,” in CVPR , 2022, pp. 17 602–17 611
2022
Closest in time.
Z. Huang, X. Shi, C. Zhang, Q. Wang, K. C. Cheung, H. Qin, J. Dai, and H. Li, “FlowFormer: A transformer architecture for optical flow,” ECCV , 2022
2022
Closest in time.
Z. Zheng, N. Nie, Z. Ling, P. Xiong, J. Liu, H. Wang, and J. Li, “Dip: Deep inverse patchmatch for high-resolution optical flow,” in CVPR , 2022, pp. 8925–8934
2022
Closest in time.
A. Luo, F. Yang, K. Luo, X. Li, H. Fan, and S. Liu, “Learning optical flow with adaptive graph reasoning,” in AAAI , 2022
2022
Closest in time.
J. Li, P. Wang, P. Xiong, T. Cai, Z. Yan, L. Yang, J. Liu, H. Fan, and S. Liu, “Practical stereo matching via cascaded recurrent network with adaptive correlation,” in CVPR , 2022, pp. 16 263–16 272
2022
Closest in time.
W. Guo, Z. Li, Y. Yang, Z. Wang, R. H. Taylor, M. Unberath, A. Yuille, and Y. Li, “Context-enhanced stereo transformer,” in ECCV . Springer, 2022, pp. 263–279
2022
Closest in time.
V. Guizilini, R. Ambruș, D. Chen, S. Zakharov, and A. Gaidon, “Multi-frame self-supervised depth with transformers,” in CVPR , 2022, pp. 160–170
2022
Closest in time.
M. Sayed, J. Gibson, J. Watson, V. Prisacariu, M. Firman, and C. Godard, “Simplerecon: 3d reconstruction without 3d convolutions,” in ECCV . Springer, 2022, pp. 1–19
2022
Closest in time.
Z. Ma, Z. Teed, and J. Deng, “Multiview stereo with cascaded epipolar raft,” in ECCV . Springer, 2022, pp. 734–750
2022
Closest in time.
V. Guizilini, I. Vasiljevic, J. Fang, R. Ambru, G. Shakhnarovich, M. R. Walter, and A. Gaidon, “Depth field networks for generalizable multi-view scene representation,” in ECCV . Springer, 2022, pp. 245–262
2022
Closest in time.
Y. Ding, W. Yuan, Q. Zhu, H. Zhang, X. Liu, Y. Wang, and X. Liu, “Transmvsnet: Global context-aware multi-view stereo network with transformers,” in CVPR , 2022, pp. 8585–8594
2022
Closest in time.
2022
Closest in time.
S. Zhao, L. Zhao, Z. Zhang, E. Zhou, and D. Metaxas, “Global matching with overlapping attention for optical flow estimation,” in CVPR , 2022, pp. 17 592–17 601
2022
Closest in time.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in CVPR , 2022, pp. 16 000–16 009
2022
Closest in time.
J. Tremblay, T. To, and S. Birchfield, “Falling things: A synthetic dataset for 3d object detection and pose estimation,” in CVPR Workshops , 2018, pp. 2038–2041
2041
Closest in time.