Fetching the paper…
Reading the bibliography…
We propose CLIP-Lite, an information efficient method for visual representation learning by feature alignment with textual annotations.
Learning representations by maximizing mutual information across views
Bachman, P., Hjelm, R. D., and Buchwalter, W. (2019) · 1906
Earlier work this paper cites.
Lookahead optimizer: k steps forward, 1 step back
Zhang, M. R., Lucas, J., Hinton, G., and Ba, J. (2019) · 1907
Earlier work this paper cites.
Contrastive representation distillation
Tian, Y., Krishnan, D., and Isola, P. (2019) · 1910
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T. (1964) · 1964
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time. iv
Donsker, M. D. and Varadhan, S. S. (1983) · 1983
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K. (2020b) · 2003
Earlier work this paper cites.
What makes for good views for contrastive learning?
Tian, Y., Sun, C., Poole, B., Krishnan, D., Schmid, C., and Isola, P. (2020) · 2005
Earlier work this paper cites.
On mutual information in contrastive learning for visual representations
Wu, M., Zhuang, C., Mosse, M., Yamins, D., and Goodman, N. (2020) · 2005
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. (2020) · 2006
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., et al. (2020) · 2006
Earlier work this paper cites.
Space-time correspondence as a contrastive random walk
Jabri, A., Owens, A., and Efros, A. A. (2020) · 2006
Earlier work this paper cites.
Towards debiasing sentence representations
Liang, P. P., Li, I. M., Zheng, E., Lim, Y. C., Salakhutdinov, R., and Morency, L.-P. (2020) · 2007
Earlier work this paper cites.
Learning visual representations using images with captions
Quattoni, A., Collins, M., and Darrell, T. (2007) · 2007
Earlier work this paper cites.
Learning video representations from textual web supervision
Stroud, J. C., Lu, Z., Sun, C., Deng, J., Sukthankar, R., Schmid, C., and Ross, D. A. (2020) · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Everingham, M., Van Gool, L., Williams, C. K., Winn, J., and Zisserman, A. (2010) · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A. (2010) · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. (2014) · 2014
Cited alongside, same era.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. (2015) · 2015
Cited alongside, same era.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L. (2015) · 2015
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017) · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.-W. (2017) · 2017
Later among the works it cites.
Mutual information neural estimation
Belghazi, M. I., Baratin, A., Rajeshwar, S., Ozair, S., Bengio, Y., Courville, A., and Hjelm, D. (2018) · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Predicting deep zero-shot convolutional neural networks using textual descriptions
Lei Ba, J., Swersky, K., Fidler, S., et al. (2015) · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T. (2015) · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. (2015) · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Tishby, N. and Zaslavsky, N. (2015) · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015) · 2015
Cited alongside, same era.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T. (2016) · 2016
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. (2018) · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O. (2018) · 2018
Later among the works it cites.
Mutual information maximization for simple and accurate part-of-speech induction
Stratos, K. (2018) · 2018
Later among the works it cites.
The inaturalist species classification and detection dataset
Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S. (2018) · 2018
Later among the works it cites.
Veličković, P., Fedus, W., Hamilton, W. L., Liò, P., Bengio, Y., and Hjelm, R. D. (2018) · 2018
Later among the works it cites.
Scaling and benchmarking self-supervised visual representation learning
Goyal, P., Mahajan, D., Gupta, A., and Misra, I. (2019) · 2019
Later among the works it cites.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. (2019) · 2019
Later among the works it cites.
Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations
Wang, T., Zhao, J., Yatskar, M., Chang, K.-W., and Ordonez, V. (2019) · 2019
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020) · 2020
Later among the works it cites.
Formal limitations on the measurement of mutual information
McAllester, D. and Stratos, K. (2020) · 2020
Later among the works it cites.
Learning visual representations with caption annotations
Sariyildiz, M. B., Perez, J., and Larlus, D. (2020) · 2020
Later among the works it cites.
Towards fairness in visual recognition: Effective strategies for bias mitigation
Wang, Z., Qinami, K., Karakozis, I. C., Genova, K., Nair, P., Hata, K., and Russakovsky, O. (2020) · 2020
Later among the works it cites.
Virtex: Learning visual representations from textual annotations
Desai, K. and Johnson, J. (2021) · 2021
Closest in time.
Natural adversarial examples
Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D. (2021) · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Closest in time.