Fetching the paper…
Reading the bibliography…
In this research work we present CLIP-GLaSS, a novel zero-shot framework to generate an image (or a caption) corresponding to a given caption (or image).
Generative adversarial networks in computer vision: A survey and taxonomy
Wang, Z., She, Q., and Ward, T. E. (2019) · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V. (2019) · 1906
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. (2019) · 1909
Earlier work this paper cites.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D. N. (2018) · 1962
Earlier work this paper cites.
A fast and elitist multiobjective genetic algorithm: Nsga-ii
Deb, K., Pratap, A., Agarwal, S., and Meyarivan, T. (2002) · 2002
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Socher, R., Ganjoo, M., Sridhar, H., Bastani, O., Manning, C. D., and Ng, A. Y. (2013) · 2013
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Generating images from captions with attention
Mansimov, E., Parisotto, E., Ba, J. L., and Salakhutdinov, R. (2015) · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. (2015) · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A. (2015) · 2015
Cited alongside, same era.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
Yu, F., Seff, A., Zhang, Y., Song, S., Funkhouser, T., and Xiao, J. (2015) · 2015
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
An introductory survey on attention mechanisms in nlp problems
Hu, D. (2019) · 2019
Later among the works it cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T. (2019) · 2019
Later among the works it cites.
Better language models and their implications
Radford, A., Wu, J., Amodei, D., Amodei, D., Clark, J., Brundage, M., and Sutskever, I. (2019) · 2019
Later among the works it cites.
Analyzing and improving the image quality of stylegan
Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., and Aila, T. (2020) · 2020
Later among the works it cites.
Clip-glass repository on github, https://github.com/galatolofederico/clip-glass
Galatolo, F. A. (2021) · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning visual n-grams from web data
Li, A., Jabri, A., Joulin, A., and van der Maaten, L. (2017) · 2017
Cited alongside, same era.
Large scale gan training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K. (2018) · 2018
Cited alongside, same era.
Closest in time.