Fetching the paper…
Reading the bibliography…
While much research has been done in text-to-image synthesis, little work has been done to explore the usage of linguistic structure of the input text.
Tree-transformer: A transformer-based method for correction of tree-structured data
Jacob Harer, Chris Reale, and Peter Chin. 2019 · 1908
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Generative adversarial nets
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Feuding families and former friends: Unsupervised learning for dynamic fictional relationships
Mohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jordan Boyd-Graber, and Hal Daumé III. 2016 · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
End-to-end relation extraction using lstms on sequences and tree structures
Makoto Miwa and Mohit Bansal. 2016 · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C Szegedy, V Vanhoucke, S Ioffe, J Shlens, and Z Wojna. 2016 · 2016
Earlier work this paper cites.
Story comprehension for predicting what happens next
Snigdha Chaturvedi, Haoruo Peng, and Dan Roth. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Weakly-supervised visual grounding of phrases with linguistic structures
Fanyi Xiao, Leonid Sigal, and Yong Jae Lee. 2017 · 2017
Earlier work this paper cites.
Towards bidirectional hierarchical representations for attention-based neural machine translation
Baosong Yang, Derek F Wong, Tong Xiao, Lidia S Chao, and Jingbo Zhu. 2017a · 2017
Earlier work this paper cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris Metaxas. 2017 · 2017
Cited alongside, same era.
Commonsense for generative multi-hop question answering tasks
Lisa Bauer, Yicheng Wang, and Mohit Bansal. 2018 · 2018
Cited alongside, same era.
Graph-to-sequence learning using gated graph neural networks
Daniel Beck, Gholamreza Haffari, and Trevor Cohn. 2018 · 2018
Cited alongside, same era.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan. 2018 · 2018
Cited alongside, same era.
Using syntax to ground referring expressions in natural images
Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency. 2018 · 2018
Cited alongside, same era.
Imagine this! scripts to compositions to videos
Pororogan: An improved story visualization model on pororo-sv dataset
Gangyan Zeng, Zhaohui Li, and Yuan Zhang. 2019 · 2019
Later among the works it cites.
X-lxmert: Paint, caption and answer questions with multi-modal transformers
Jaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi, and Aniruddha Kembhavi. 2020 · 2020
Later among the works it cites.
Effectively unbiased fid and inception score and where to find them
Min Jin Chong and David Forsyth. 2020 · 2020
Later among the works it cites.
Contragan: Contrastive learning for conditional image generation
Minguk Kang and Jaesik Park. 2020 · 2020
Later among the works it cites.
Dense-caption matching and frame-selection gating for temporal localization in videoqa
Hyounghun Kim, Zineng Tang, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, and Aniruddha Kembhavi. 2018 · 2018
Cited alongside, same era.
Inferring semantic layout for hierarchical text-to-image synthesis
Seunghoon Hong, Dingdong Yang, Jongwook Choi, and Honglak Lee. 2018 · 2018
Cited alongside, same era.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein. 2018 · 2018
Cited alongside, same era.
cgans with projection discriminator
Takeru Miyato and Masanori Koyama. 2018 · 2018
Cited alongside, same era.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018 · 2018
Cited alongside, same era.
Incorporating structured commonsense knowledge in story completion
Jiaao Chen, Jianshu Chen, and Zhou Yu. 2019 · 2019
Cited alongside, same era.
Story ending generation with incremental encoding and commonsense knowledge
Jian Guan, Yansen Wang, and Minlie Huang. 2019 · 2019
Cited alongside, same era.
Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara Berg, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Improved-storygan for sequential images visualization
Chunye Li, Liya Kong, and Zhiping Zhou. 2020 · 2020
Later among the works it cites.
Tree-structured attention with hierarchical accumulation
Xuan-Phi Nguyen, Shafiq Joty, Steven CH Hoi, and Richard Socher. 2020 · 2020
Later among the works it cites.
Connecting vision and language with localized narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari. 2020 · 2020
Later among the works it cites.
Character-preserving coherent story visualization
Yun-Zhu Song, Zhi-Rui Tam, Hung-Jen Chen, Huiao-Han Lu, and Hong-Han Shuai. 2020 · 2020
Later among the works it cites.
Transgan: Two transformers can make one strong gan
Yifan Jiang, Shiyu Chang, and Zhangyang Wang. 2021 · 2021
Closest in time.
Text-to-image generation grounded by fine-grained user attention
Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. 2021 · 2021
Closest in time.
Improving generation and evaluation of visual stories via semantic consistency
Adyasha Maharana, Darryl Hannan, and Mohit Bansal. 2021 · 2021
Closest in time.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Closest in time.
Cross-modal contrastive learning for text-to-image generation
Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. 2021 · 2021
Closest in time.
Deepstory: Video story qa by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang. 2017 · 2022
Closest in time.