Fetching the paper…
Reading the bibliography…
We introduce UViM, a unified approach capable of modeling a wide range of computer vision tasks.
An algorithm for vector quantizer design
Y. Linde, A. Buzo, and R. Gray · 1980
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
D. Koller and N. Friedman · 2009
Earlier work this paper cites.
Structured learning and prediction in computer vision
S. Nowozin and C. H. Lampert · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
D. Eigen, C. Puhrsch, and R. Fergus · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Chained predictions using convolutional neural networks
G. Gkioxari, A. Toshev, and N. Jaitly · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Conditional image generation with PixelCNN decoders
A. van den Oord, N. Kalchbrenner, L. Espeholt, K. Kavukcuoglu, O. Vinyals, and A. Graves · 2016
Cited alongside, same era.
Pixel recurrent neural networks
A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Cited alongside, same era.
DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2017
Cited alongside, same era.
PixColor: Pixel recursive colorization
S. Guadarrama, R. Dahl, D. Bieber, M. Norouzi, J. Shlens, and K. Murphy · 2017
Cited alongside, same era.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Later among the works it cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Later among the works it cites.
Masked-attention mask transformer for universal image segmentation
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar · 2021
Later among the works it cites.
Per-pixel classification is not all you need for semantic segmentation
B. Cheng, A. Schwing, and A. Kirillov · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GANs trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros · 2017
Cited alongside, same era.
PixelCNN models with auxiliary variables for natural image modeling
A. Kolesnikov and C. H. Lampert · 2017
Cited alongside, same era.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Cited alongside, same era.
Probabilistic image colorization
A. Royer, A. Kolesnikov, and C. H. Lampert · 2017
Cited alongside, same era.
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma · 2017
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Later among the works it cites.
Perceiver IO: A general architecture for structured inputs & outputs
A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, et al · 2021
Later among the works it cites.
Colorization transformer
M. Kumar, D. Weissenborn, and N. Kalchbrenner · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Palette: Image-to-image diffusion models
C. Saharia, W. Chan, H. Chang, C. A. Lee, J. Ho, T. Salimans, D. J. Fleet, and M. Norouzi · 2021
Later among the works it cites.
How to train your ViT? data, augmentation, and regularization in vision transformers
A. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer · 2021
Later among the works it cites.
Nüwa: Visual synthesis pre-training for neural visual world creation
C. Wu, J. Liang, L. Ji, F. Yang, Y. Fang, D. Jiang, and N. Duan · 2021
Later among the works it cites.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2021
Later among the works it cites.
Pix2seq: A language modeling framework for object detection
T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. Hinton · 2022
Closest in time.
BinsFormer: Revisiting adaptive bins for monocular depth estimation
Z. Li, X. Wang, X. Liu, and J. Jiang · 2022
Closest in time.
Transframer: Arbitrary frame prediction with generative models
C. Nash, J. Carreira, J. Walker, I. Barr, A. Jaegle, M. Malinowski, and P. Battaglia · 2022
Closest in time.
Vector-quantized image modeling with improved VQGAN
J. Yu, X. Li, J. Y. Koh, H. Zhang, R. Pang, J. Qin, A. Ku, Y. Xu, J. Baldridge, and Y. Wu · 2022
Closest in time.
NeW CRFs: Neural window fully-connected CRFs for monocular depth estimation
W. Yuan, X. Gu, Z. Dai, S. Zhu, and P. Tan · 2022
Closest in time.