Fetching the paper…
Reading the bibliography…
In this paper, we aim to optimize a contrastive loss with individualized temperatures in a principled and systematic manner for self-supervised learning.
On general minimax theorems
Sion, M · 1958
Earlier work this paper cites.
Subgradient methods for saddle-point problems
Nedić, A. and Ozdaglar, A · 2009
Earlier work this paper cites.
Variational analysis , volume 317
Rockafellar, R. T. and Wets, R. J.-B · 2009
Earlier work this paper cites.
Robust solutions of optimization problems affected by uncertain probabilities
Ben-Tal, A., Den Hertog, D., De Waegenaere, A., Melenberg, B., and Rennen, G · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Capturing long-tail distributions of object subcategories
Zhu, X., Anguelov, D., and Ramanan, D · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., et al · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Variance-based regularization with convex objectives
Namkoong, H. and Duchi, J. C · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Earlier work this paper cites.
Data-driven robust optimization
Bertsimas, D., Gupta, V., and Kallus, N · 2018
Earlier work this paper cites.
The inaturalist species classification and detection dataset
Horn, G. V., Aodha, O. M., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S. J · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Robust wasserstein profile inference and applications to machine learning
Blanchet, J., Kang, Y., and Murthy, K · 2019
Earlier work this paper cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Cao, K., Wei, C., Gaidon, A., Arechiga, N., and Ma, T · 2019
Earlier work this paper cites.
Class-balanced loss based on effective number of samples
Cui, Y., Jia, M., Lin, T.-Y., Song, Y., and Belongie, S · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
Distributionally robust optimization and generalization in kernel methods
Staib, M. and Jegelka, S · 2019
Earlier work this paper cites.
Pytorch image models
Wightman, R · 2019
Earlier work this paper cites.
Large scale incremental learning
Wu, Y., Chen, Y., Wang, L., Ye, Y., Liu, Z., Guo, Y., and Fu, Y · 2019
Earlier work this paper cites.
Non-asymptotic analysis of stochastic methods for non-smooth non-convex regularized problems
Xu, Y., Jin, R., and Yang, T · 2019
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
Debiased contrastive learning
Chuang, C.-Y., Robinson, J., Lin, Y.-C., Torralba, A., and Jegelka, S · 2020
Cited alongside, same era.
Learning to discriminate information for online action detection
Eun, H., Moon, J., Park, J., Jung, C., and Kim, C · 2020
Cited alongside, same era.
Does learning require memorization? a short tale about a long tail
Feldman, V · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., et al · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Cited alongside, same era.
Hard negative mixing for contrastive learning
Kalantidis, Y., Sariyildiz, M. B., Pion, N., Weinzaepfel, P., and Larlus, D · 2020
Contrastive self-supervised learning for sensor-based human activity recognition
Khaertdinov, B., Ghaleb, E., and Asteriadis, S · 2021
Later among the works it cites.
Stochastic optimization of areas under precision-recall curves with provable convergence
Qi, Q., Luo, Y., Xu, Z., Ji, S., and Yang, T · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Contrastive learning with hard negative samples
Robinson, J. D., Chuang, C.-Y., Sra, S., and Jegelka, S · 2021
Later among the works it cites.
Viewmaker networks: Learning views for unsupervised representation learning
Tamkin, A., Wu, M., and Goodman, N · 2021
Later among the works it cites.
Understanding the behaviour of contrastive loss
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large-scale methods for distributionally robust optimization
Levy, D., Carmon, Y., Duchi, J. C., and Sidford, A · 2020
Cited alongside, same era.
Prototypical contrastive learning of unsupervised representations
Li, J., Zhou, P., Xiong, C., and Hoi, S. C · 2020
Cited alongside, same era.
Attentional biased stochastic gradient for imbalanced classification
Qi, Q., Xu, Y., Jin, R., Yin, W., and Yang, T · 2020
Cited alongside, same era.
Byol works even without batch statistics
Richemond, P. H., Grill, J.-B., Altché, F., Tallec, C., Strub, F., Brock, A., Smith, S., De, S., Pascanu, R., Piot, B., et al · 2020
Cited alongside, same era.
What makes for good views for contrastive learning?
Tian, Y., Sun, C., Poole, B., Krishnan, D., Schmid, C., and Isola, P · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Cited alongside, same era.
Wang, F. and Liu, H · 2021
Later among the works it cites.
Contrastive learning with stronger augmentations
Wang, X. and Qi, G.-J · 2021
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
Yao, L., Huang, R., Hou, L., Lu, G., Niu, M., Xu, H., Liang, X., Li, Z., Jiang, X., and Xu, C · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Later among the works it cites.
Temperature as uncertainty in contrastive learning
Zhang, O., Wu, M., Bayrooti, J., and Goodman, N · 2021
Later among the works it cites.
Coarse-to-fine vision-language pre-training with fusion in the backbone
Dou, Z.-Y., Kamath, A., Gan, Z., Zhang, P., Wang, J., Li, L., Liu, Z., Liu, C., LeCun, Y., Peng, N., et al · 2022
Later among the works it cites.
Cyclip: Cyclic contrastive language-image pretraining
Goel, S., Bansal, H., Bhatia, S., Rossi, R. A., Vinay, V., and Grover, A · 2022
Later among the works it cites.
A stochastic subgradient method for distributionally robust non-convex and non-smooth learning
Gürbüzbalaban, M., Ruszczyński, A., and Zhu, L · 2022
Later among the works it cites.
Contrastive masked autoencoders are stronger vision learners
Huang, Z., Jin, X., Lu, C., Hou, Q., Cheng, M.-M., Fu, D., Shen, X., and Feng, J · 2022
Later among the works it cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Liang, W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J · 2022
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Mu, N., Kirillov, A., Wagner, D., and Xie, S · 2022
Later among the works it cites.
Stochastic constrained dro with a complexity independent of sample size
Qi, Q., Lyu, J., Bai, E. W., Yang, T., et al · 2022
Later among the works it cites.
Tomasev, N., Bica, I., McWilliams, B., Buesing, L., Pascanu, R., Blundell, C., and Mitrovic, J · 2022
Later among the works it cites.
Finite-sum compositional stochastic optimization: Theory and applications
Wang, B. and Yang, T · 2022
Later among the works it cites.
Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix
Wang, T., Jiang, W., Lu, Z., Zheng, F., Cheng, R., Yin, C., and Luo, P · 2022
Later among the works it cites.
Progcl: Rethinking hard negative mining in graph contrastive learning
Xia, J., Wu, L., Wang, G., Chen, J., and Li, S. Z · 2022
Later among the works it cites.
Delving into inter-image invariance for unsupervised visual representations
Xie, J., Zhan, X., Liu, Z., Ong, Y.-S., and Loy, C. C · 2022
Later among the works it cites.
Algorithmic foundation of deep x-risk optimization
Yang, T · 2022
Later among the works it cites.
Provable stochastic optimization for global contrastive learning: Small batch does not harm performance
Yuan, Z., Wu, Y., Qiu, Z.-H., Du, X., Zhang, L., Zhou, D., and Yang, T · 2022
Later among the works it cites.
Dual temperature helps contrastive learning without many negative samples: Towards understanding and simplifying moco
Zhang, C., Zhang, K., Pham, T. X., Niu, A., Qiao, Z., Yoo, C. D., and Kweon, I. S · 2022
Later among the works it cites.
How well do unsupervised learning algorithms model human real-time and life-long learning?
Zhuang, C., Xiang, V., Bai, Y., Jia, X., Turk-Browne, N., Norman, K., DiCarlo, J. J., and Yamins, D. L · 2022
Later among the works it cites.
Zhu, L., Gürbüzbalaban, M., and Ruszczyński, A · 2023
Closest in time.