Fetching the paper…
Reading the bibliography…
We formulate monocular depth estimation using denoising diffusion models, inspired by their recent successes in high fidelity image generation.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Learning depth from single monocular images
Saxena, A., Chung, S., and Ng, A · 2005
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Make3d: Learning 3d scene structure from a single still image
Saxena, A., Sun, M., and Ng, A. Y · 2009
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Silberman, N., Hoiem, D., Kohli, P., and Fergus, R · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Geiger, A., Lenz, P., Stiller, C., and Urtasun, R · 2013
Earlier work this paper cites.
Eigen, D. and Fergus, R · 2014
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
Eigen, D., Puhrsch, C., and Fergus, R · 2014
Earlier work this paper cites.
Inpainting of missing values in the kinect sensor’s depth maps based on background estimates
Stommel, M., Beetz, M., and Xu, W · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
ShapeNet: An information-rich 3d model repository
Chang, A. X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., Xiao, J., Yi, L., and Yu, F · 2015
Earlier work this paper cites.
Scenenet: Understanding real world indoor scenes with synthetic data
Handa, A., Patraucean, V., Badrinarayanan, V., Stent, S., and Cipolla, R · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Cao, Y., Wu, Z., and Shen, C · 2016
Earlier work this paper cites.
Unsupervised cnn for single view depth estimation: Geometry to the rescue
Garg, R., Bg, V. K., Carneiro, G., and Reid, I · 2016
Earlier work this paper cites.
Deeper depth prediction with fully convolutional residual networks
Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N · 2016
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks
Lamb, A., Goyal, A., Zhang, Y., Zhang, S., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Learning representations for automatic colorization
Larsson, G., Maire, M., and Shakhnarovich, G · 2016
Earlier work this paper cites.
Scenenet rgb-d: 5m photorealistic images of synthetic indoor trajectories with ground truth
McCormac, J., Handa, A., Leutenegger, S., and Davison, A. J · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2016
Earlier work this paper cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y · 2016
Earlier work this paper cites.
Zhang, R., Isola, P., and Efros, A. A · 2016
Earlier work this paper cites.
ScanNet: Richly-annotated 3d reconstructions of indoor scenes
Dai, A., Chang, A. X., Savva, M., Halber, M., Funkhouser, T., and Nießner, M · 2017
Cited alongside, same era.
Unsupervised monocular depth estimation with left-right consistency
Godard, C., Mac Aodha, O., and Brostow, G. J · 2017
Cited alongside, same era.
Cross-domain self-supervised multi-task feature learning using synthetic imagery
Ren, Z. and Lee, Y. J · 2017
Cited alongside, same era.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A · 2017
Cited alongside, same era.
Deep ordinal regression network for monocular depth estimation
Fu, H., Gong, M., Wang, C., Batmanghelich, K., and Tao, D · 2018
Cited alongside, same era.
Vision transformers for dense prediction
Ranftl, R., Bochkovskiy, A., and Koltun, V · 2021
Later among the works it cites.
Pixelsynth: Generating a 3d-consistent experience from a single image
Rockwell, C., Fouhey, D. F., and Johnson, J · 2021
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Later among the works it cites.
Transformer-based dual relation graph for multi-label image recognition
Zhao, J., Yan, K., Zhao, Y., Guo, X., Huang, F., and Li, J · 2021
Later among the works it cites.
Attention attention everywhere: Monocular depth prediction with skip attention
Agarwal, A. and Arora, C · 2022
Later among the works it cites.
Instructpix2pix: Learning to follow image editing instructions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Monocular depth estimation with affinity, vertical pooling, and label enhancement
Gan, Y., Xu, X., Sun, W., and Lin, L · 2018
Cited alongside, same era.
From big to small: Multi-scale local planar guidance for monocular depth estimation
Lee, J. H., Han, M.-K., Ko, D. W., and Suh, I. H · 2019
Cited alongside, same era.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., and Koltun, V · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Cited alongside, same era.
Enforcing geometric constraints of virtual normal for depth prediction
Yin, W., Liu, Y., Shen, C., and Yan, Y · 2019
Cited alongside, same era.
Does computer vision matter for action?
Zhou, B., Krähenbühl, P., and Koltun, V · 2019
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Brooks, T., Holynski, A., and Efros, A. A · 2022
Later among the works it cites.
A generalist framework for panoptic segmentation of images and videos
Chen, T., Li, L., Saxena, S., Hinton, G., and Fleet, D. J · 2022
Later among the works it cites.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A · 2022
Later among the works it cites.
Prompt-to-prompt image editing with cross attention control
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Later among the works it cites.
Jing, L., Yu, R., Kretzschmar, H., Li, K., Qi, C. R., Zhao, H., Ayvaci, A., Chen, X., Cower, D., Li, Y., You, Y., Deng, H., Li, C., and Anguelov, D · 2022
Later among the works it cites.
Binsformer: Revisiting adaptive bins for monocular depth estimation
Li, Z., Wang, X., Liu, X., and Jiang, J · 2022
Later among the works it cites.
Point-e: A system for generating 3d point clouds from complex prompts
Nichol, A., Jun, H., Dhariwal, P., Mishkin, P., and Chen, M · 2022
Later among the works it cites.
Dreamfusion: Text-to-3d using 2d diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Progressive distillation for fast sampling of diffusion models
Salimans, T. and Ho, J · 2022
Later among the works it cites.
Step-unrolled denoising autoencoders for text generation
Savinov, N., Chung, J., Binkowski, M., Elsen, E., and van den Oord, A · 2022
Later among the works it cites.
Imagen editor and editbench: Advancing and evaluating text-guided image inpainting
Wang, S., Saharia, C., Montgomery, C., Pont-Tuset, J., Noy, S., Pellegrini, S., Onoe, Y., Laszlo, S., Fleet, D. J., Soricut, R., Baldridge, J., Norouzi, M., Anderson, P., and Chan, W · 2022
Later among the works it cites.
Novel view synthesis with diffusion models
Watson, D., Chan, W., Martin-Brualla, R., Ho, J., Tagliasacchi, A., and Norouzi, M · 2022
Later among the works it cites.
Revealing the dark secrets of masked image modeling
Xie, Z., Geng, Z., Hu, J., Zhang, Z., Hu, H., and Cao, Y · 2022
Later among the works it cites.
All in tokens: Unifying output space of visual tasks via soft token
Ning, J., Li, C., Zhang, Z., Geng, Z., Dai, Q., He, K., and Hu, H · 2023
Closest in time.