Fetching the paper…
Reading the bibliography…
The advancements in generative modeling, particularly the advent of diffusion models, have sparked a fundamental question: how can these models be effectively used for discriminative tasks? In this work, we find that generative models can be great test-time adapters for discriminative models.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Vision as bayesian inference: analysis by synthesis?
A. Yuille and D. Kersten · 2006
Earlier work this paper cites.
Brain states: Top-down influences in sensory processing
C. D. Gilbert and M. Sigman · 2007
Earlier work this paper cites.
To recognize shapes, first learn to generate images
G. E. Hinton · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
WILDS: A benchmark of in-the-wild distribution shifts
P. W. Koh, S. Sagawa, H. Marklund, S. M. Xie, M. Zhang, A. Balsubramani, W. Hu, M. Yasunaga, R. L. Phillips, S. Beery, J. Leskovec, A. Kundaje, E. Pierson, S. Levine, C. Finn, and P. Liang · 2012
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Perception as an inference problem
B. A. Olshausen · 2013
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
L. Bossard, M. Guillaumin, and L. Van Gool · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
J. Donahue, P. Krähenbühl, and T. Darrell · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Earlier work this paper cites.
High quality monocular depth estimation via transfer learning
I. Alhashim and P. Wonka · 2018
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz · 2019
Cited alongside, same era.
Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
D. Hendrycks and T. Dietterich · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
A convnet for the 2020s
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie · 2022
Later among the works it cites.
Test-time prompt tuning for zero-shot generalization in vision-language models
S. Manli, N. Weili, H. De-An, Y. Zhiding, G. Tom, A. Anima, and X. Chaowei · 2022
Later among the works it cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Later among the works it cites.
Continual test-time domain adaptation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Shortcut learning in deep neural networks
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al · 2020
Cited alongside, same era.
Test-time training with self-supervision for generalization under distribution shifts
Y. Sun, X. Wang, Z. Liu, J. Miller, A. Efros, and M. Hardt · 2020
Cited alongside, same era.
Tent: Fully test-time adaptation by entropy minimization
D. Wang, E. Shelhamer, S. Liu, B. Olshausen, and T. Darrell · 2020
Cited alongside, same era.
Label-efficient semantic segmentation with diffusion models
D. Baranchuk, I. Rubachev, A. Voynov, V. Khrulkov, and A. Babenko · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Q. Wang, O. Fink, L. Van Gool, and D. Dai · 2022
Later among the works it cites.
Synthetic data from diffusion models improves imagenet classification
S. Azizi, S. Kornblith, C. Saharia, M. Norouzi, and D. J. Fleet · 2023
Closest in time.
Text-to-image diffusion models are zero-shot classifiers
K. Clark and P. Jaini · 2023
Closest in time.
Test-time adaptation with slot-centric models
M. Prabhudesai, A. Goyal, S. Paul, S. Van Steenkiste, M. S. Sajjadi, G. Aggarwal, T. Kipf, D. Pathak, and K. Fragkiadaki · 2023
Closest in time.
The effectiveness of mae pre-pretraining for billion-scale pretraining
M. Singh, Q. Duval, K. V. Alwala, H. Fan, V. Aggarwal, A. Adcock, A. Joulin, P. Dollár, C. Feichtenhofer, R. Girshick, et al · 2023
Closest in time.
Consistency models
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever · 2023
Closest in time.
Effective data augmentation with diffusion models
B. Trabucco, K. Doherty, M. Gurinas, and R. Salakhutdinov · 2023
Closest in time.
Open-vocabulary panoptic segmentation with text-to-image diffusion models
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello · 2023
Closest in time.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
L. Zhang and M. Agrawala · 2023
Closest in time.
On pitfalls of test-time adaptation
H. Zhao, Y. Liu, A. Alahi, and T. Lin · 2023
Closest in time.