Fetching the paper…
Reading the bibliography…
The creation of high-quality human-labeled image-caption datasets presents a significant bottleneck in the development of Visual-Language Models (VLMs).
Do dall-e and flamingo understand each other?
H. Li, J. Gu, R. Koner, S. Sharifzadeh, and V. Tresp · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Learning from synthetic humans
G. Varol, J. Romero, X. Martin, N. Mahmood, M. J. Black, I. Laptev, and C. Schmid · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros · 2017
Earlier work this paper cites.
Learning semantic segmentation from synthetic data: A geometrically guided input-output adaptation approach
Y. Chen, W. Li, X. Chen, and L. V. Gool · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al · 2020
Cited alongside, same era.
Structured3d: A large photo-realistic dataset for structured 3d modeling
J. Zheng, J. Zhang, J. Li, R. Tang, S. Gao, and Z. Zhou · 2020
Cited alongside, same era.
Label-efficient semantic segmentation with diffusion models
D. Baranchuk, I. Rubachev, A. Voynov, V. Khrulkov, and A. Babenko · 2021
Cited alongside, same era.
High-performance large-scale image recognition without normalization
A. Brock, S. De, S. L. Smith, and K. Simonyan · 2021
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut · 2021
Cited alongside, same era.
Pali: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al · 2022
Later among the works it cites.
Next-generation deep learning based on simulators and synthetic data
C. M. de Melo, A. Torralba, L. Guibas, J. DiCarlo, R. Chellappa, and J. Hodgins · 2022
Later among the works it cites.
Kubric: A scalable dataset generator
K. Greff, F. Belletti, L. Beyer, C. Doersch, Y. Du, D. Duckworth, D. J. Fleet, D. Gnanapragasam, F. Golemo, C. Herrmann, et al · 2022
Later among the works it cites.
Learning video representations of human motion from synthetic data
X. Guo, W. Wu, D. Wang, J. Su, H. Su, W. Gan, J. Huang, and Q. Yang · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Cited alongside, same era.
Semantic segmentation with generative models: Semi-supervised learning and strong out-of-domain generalization
D. Li, J. Yang, K. Kreis, A. Torralba, and S. Fidler · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Imagenet-21k pretraining for the masses
T. Ridnik, E. Ben-Baruch, A. Noy, and L. Zelnik-Manor · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Cited alongside, same era.
Classification by attention: Scene graph classification with prior knowledge
S. Sharifzadeh, S. M. Baharlou, and V. Tresp · 2021
Cited alongside, same era.
Later among the works it cites.
Scaling up vision-language pre-training for image captioning
X. Hu, Z. Gan, J. Wang, Z. Yang, Z. Liu, Y. Lu, and L. Wang · 2022
Later among the works it cites.
Pretrained diffusion models for unified human motion synthesis
J. Ma, S. Bai, and C. Zhou · 2022
Later among the works it cites.
Task2sim: Towards effective pre-training and transfer from synthetic data
S. Mishra, R. Panda, C. P. Phoo, C.-F. R. Chen, L. Karlinsky, K. Saenko, V. Saligrama, and R. S. Feris · 2022
Later among the works it cites.
Improving scene graph classification by exploiting knowledge from texts
S. Sharifzadeh, S. M. Baharlou, M. Schmitt, H. Schütze, and V. Tresp · 2022
Later among the works it cites.
Synthetic data from diffusion models improves imagenet classification
S. Azizi, S. Kornblith, C. Saharia, M. Norouzi, and D. J. Fleet · 2023
Later among the works it cites.
Going beyond nouns with vision & language models using synthetic data
P. Cascante-Bonilla, K. Shehada, J. S. Smith, S. Doveh, D. Kim, R. Panda, G. Varol, A. Oliva, V. Ordonez, R. Feris, et al · 2023
Later among the works it cites.
Muse: Text-to-image generation via masked generative transformers
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. Murphy, W. T. Freeman, M. Rubinstein, et al · 2023
Later among the works it cites.
Scaling laws of synthetic images for model training… for now
L. Fan, K. Chen, D. Krishnan, D. Katabi, P. Isola, and Y. Tian · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
I. Shumailov, Z. Shumaylov, Y. Zhao, Y. Gal, N. Papernot, and R. Anderson · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
M. Gerstgrasser, R. Schaeffer, A. Dey, R. Rafailov, H. Sleight, J. Hughes, T. Korbak, R. Agrawal, D. Pai, A. Gromov, et al · 2024
Closest in time.