Fetching the paper…
Reading the bibliography…
Text-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding.
A new algorithm for data compression
P. Gage · 1994
Earlier work this paper cites.
The human visual cortex
K. Grill-Spector and R. Malach · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
T. Kudo and J. Richardson · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Learning a deep convolutional network for image super-resolution
C. Dong, C. C. Loy, K. He, and X. Tang · 2014
Earlier work this paper cites.
Generative adversarial networks
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Draw: A recurrent neural network for image generation
K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Multimodal feature integration in the angular gyrus during episodic and semantic retrieval
H. M. Bonnici, F. R. Richter, Y. Yazar, and J. S. Simons · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Generating images from captions with attention
E. Mansimov, E. Parisotto, J. L. Ba, and R. Salakhutdinov · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee · 2016
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
A. Van Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
A. Caliskan, J. J. Bryson, and A. Narayanan · 2017
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Cited alongside, same era.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas · 2017
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, et al · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Later among the works it cites.
Self-supervised learning: Generative or contrastive
X. Liu, F. Zhang, Z. Hou, Z. Wang, L. Mian, J. Zhang, and J. Tang · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Brundage, S. Avin, J. Clark, H. Toner, P. Eckersley, B. Garfinkel, A. Dafoe, P. Scharre, T. Zeitzoff, B. Filar, et al · 2018
Cited alongside, same era.
Lagging inference networks and posterior collapse in variational autoencoders
J. He, D. Spokoyny, G. Neubig, and T. Berg-Kirkpatrick · 2018
Cited alongside, same era.
Image transformer
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He · 2018
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning · 2019
Cited alongside, same era.
Object-driven text-to-image synthesis via adversarial training
W. Li, P. Zhang, L. Zhang, Q. Huang, X. He, S. Lyu, and J. Gao · 2019
Cited alongside, same era.
Later among the works it cites.
Df-gan: Deep fusion generative adversarial networks for text-to-image synthesis
M. Tao, H. Tang, S. Wu, N. Sebe, F. Wu, and X.-Y. Jing · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Later among the works it cites.
An empirical study of training self-supervised visual transformers
X. Chen, S. Xie, and K. He · 2021
Closest in time.
All nlp tasks are generation tasks: A general pretraining framework
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang · 2021
Closest in time.
Text-to-image generation grounded by fine-grained user attention
J. Y. Koh, J. Baldridge, H. Lee, and Y. Yang · 2021
Closest in time.
M6: A chinese multimodal pretrainer
J. Lin, R. Men, A. Yang, C. Zhou, M. Ding, Y. Zhang, P. Wang, A. Wang, L. Jiang, X. Jia, et al · 2021
Closest in time.
X. Liu, Y. Zheng, Z. Du, M. Ding, Y. Qian, Z. Yang, and J. Tang · 2021
Closest in time.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Closest in time.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Closest in time.
Wudaocorpora: A super large-scale chinese corpora for pre-training language models
S. Yuan, H. Zhao, Z. Du, M. Ding, X. Liu, Y. Cen, X. Zou, and Z. Yang · 2021
Closest in time.
Controllable generation from pre-trained language models via inverse prompting
X. Zou, D. Yin, Q. Zhong, H. Yang, Z. Yang, and J. Tang · 2021
Closest in time.