Fetching the paper…
Reading the bibliography…
Pre-training over mixtured multi-task, multi-domain, and multi-modal data remains an open challenge in vision perception pre-training.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
X. Chen, H. Fan, R. B. Girshick, and K. He · 2003
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. Hinton · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Semantic contours from inverse detectors
B. Hariharan, P. Arbelaez, L. D. Bourdev, S. Maji, and J. Malik · 2011
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Layer normalization
L. J. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Earlier work this paper cites.
Universal representations: The missing link between faces, text, planktons, and cat breeds
H. Bilen and A. Vedaldi · 2017
Earlier work this paper cites.
Deformable convolutional networks
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei · 2017
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
C. J. Maddison, A. Mnih, and Y. W. Teh · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
S.-A. Rebuffi, H. Bilen, and A. Vedaldi · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Earlier work this paper cites.
A survey on multi-task learning
Y. Zhang and Q. Yang · 2017
Earlier work this paper cites.
Scene parsing through ADE20K dataset
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba · 2017
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
H. Caesar, J. R. R. Uijlings, and V. Ferrari · 2018
Earlier work this paper cites.
Cascade R-CNN: delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
S. Gidaris, P. Singh, and N. Komodakis · 2018
Earlier work this paper cites.
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, T. Duerig, and V. Ferrari · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks
S.-A. Rebuffi, H. Bilen, and A. Vedaldi · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun · 2018
Earlier work this paper cites.
Taskonomy: Disentangling task transfer learning
A. R. Zamir, A. Sax, W. B. Shen, L. J. Guibas, J. Malik, and S. Savarese · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cissé, Y. N. Dauphin, and D. Lopez-Paz · 2018
Cited alongside, same era.
Class-balanced loss based on effective number of samples
Y. Cui, M. Jia, T. Lin, Y. Song, and S. J. Belongie · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
LVIS: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollár, and R. B. Girshick · 2019
Cited alongside, same era.
Rethinking imagenet pre-training
K. He, R. B. Girshick, and P. Dollár · 2019
Cited alongside, same era.
From big to small: Multi-scale local planar guidance for monocular depth estimation
J. H. Lee, M.-K. Han, D. W. Ko, and I. H. Suh · 2019
VATT: transformers for multimodal self-supervised learning from raw video, audio and text
H. Akbari, L. Yuan, R. Qian, W. Chuang, S. Chang, Y. Cui, and B. Gong · 2021
Later among the works it cites.
Beit: BERT pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2021
Later among the works it cites.
GAIA: A transfer learning system of object detection that fits your needs
X. Bu, J. Peng, J. Yan, T. Tan, and Z. Zhang · 2021
Later among the works it cites.
Multisiam: Self-supervised multi-instance siamese representation learning for autonomous driving
K. Chen, L. Hong, H. Xu, Z. Li, and D. Yeung · 2021
Later among the works it cites.
An empirical study of training self-supervised vision transformers
X. Chen*, S. Xie*, and K. He · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An analysis of pre-training on object detection
H. Li, B. Singh, M. Najibi, Z. Wu, and L. S. Davis · 2019
Cited alongside, same era.
DARTS: differentiable architecture search
H. Liu, K. Simonyan, and Y. Yang · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Cited alongside, same era.
Nettailor: Tuning the architecture, not just the weights
P. Morgado and N. Vasconcelos · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Cited alongside, same era.
Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
H. Hazimeh, Z. Zhao, A. Chowdhery, M. Sathiamoorthy, Y. Chen, R. Mazumder, L. Hong, and E. H. Chi · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2021
Later among the works it cites.
Unit: Multimodal multitask learning with a unified transformer
R. Hu and A. Singh · 2021
Later among the works it cites.
MURAL: multimodal, multitask retrieval across languages
A. Jain, M. Guo, K. Srinivasan, T. Chen, S. Kudugunta, C. Jia, Y. Yang, and J. Baldridge · 2021
Later among the works it cites.
Introducing pathways: A next-generation ai architecture, 2021
D. Jeff · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Combined scaling for zero-shot transfer learning
H. Pham, Z. Dai, G. Ghiasi, H. Liu, A. W. Yu, M. Luong, M. Tan, and Q. V. Le · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Imagenet-21k pretraining for the masses, 2021
T. Ridnik, E. Ben-Baruch, A. Noy, and L. Zelnik-Manor · 2021
Later among the works it cites.
Multi-dataset pretraining: A unified model for semantic segmentation
B. Shi, X. Zhang, H. Xu, W. Dai, J. Zou, H. Xiong, and Q. Tian · 2021
Later among the works it cites.
Monocular depth estimation using laplacian pyramid-based depth residuals
M. Song, S. Lim, and W. Kim · 2021
Later among the works it cites.
Multi-path neural networks for on-device multi-domain visual classification
Q. Wang, J. Ke, J. Greaves, G. Chu, G. Bender, L. Sbaiz, A. Go, A. Howard, M.-H. Yang, J. Gilbert, P. Milanfar, and F. Yang · 2021
Later among the works it cites.
Dense contrastive learning for self-supervised visual pre-training
X. Wang, R. Zhang, C. Shen, T. Kong, and L. Li · 2021
Later among the works it cites.
Detco: Unsupervised contrastive learning for object detection
E. Xie, J. Ding, W. Wang, X. Zhan, H. Xu, P. Sun, Z. Li, and P. Luo · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
L. Yuan, D. Chen, Y. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, C. Liu, M. Liu, Z. Liu, Y. Lu, Y. Shi, L. Wang, J. Wang, B. Xiao, Z. Xiao, J. Yang, M. Zeng, L. Zhou, and P. Zhang · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Closest in time.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
Closest in time.
Binsformer: Revisiting adaptive bins for monocular depth estimation
Z. Li, X. Wang, X. Liu, and J. Jiang · 2022
Closest in time.
A convnet for the 2020s
Z. Liu, H. Mao, C. Wu, C. Feichtenhofer, T. Darrell, and S. Xie · 2022
Closest in time.
Controllable dynamic multi-task architectures
D. S. Raychaudhuri, Y. Suh, S. Schulter, X. Yu, M. Faraki, A. K. Roy-Chowdhury, and M. Chandraker · 2022
Closest in time.
Newcrfs: Neural window fully-connected crfs for monocular depth estimation
W. Yuan, X. Gu, Z. Dai, S. Zhu, and P. Tan · 2022
Closest in time.
ibot: Image bert pre-training with online tokenizer
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong · 2022
Closest in time.