Fetching the paper…
Reading the bibliography…
There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions.
Evaluating automated and manual acquisition of anaphora resolution strategies
Chinatsu Aone and Scott William · 1995
Earlier work this paper cites.
Using decision trees for coreference resolution
Joseph F McCarthy and Wendy G Lehnert · 1995
Earlier work this paper cites.
Probabilistic coreference in information extraction
Andrew Kehler · 1997
Earlier work this paper cites.
A machine learning approach to coreference resolution of noun phrases
Wee Meng Soon, Hwee Tou Ng, and Daniel Chung Yong Lim · 2001
Earlier work this paper cites.
Narrowing the modeling gap: A cluster-ranking approach to coreference resolution
Altaf Rahman and Vincent Ng · 2011
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
NICE: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-RNN)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L. Yuille · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott E. Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Density estimation using real NVP
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Visual reference resolution using attention memory for visual dialog
Paul Hongsuck Seo, Andreas Lehrmann, Bohyung Han, and Leonid Sigal · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Diverse and accurate image description using a variational auto-encoder with an additive Gaussian encoding space
Liwei Wang, Alexander G. Schwing, and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Imagine this! scripts to compositions to videos
Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, and Aniruddha Kembhavi · 2018
Cited alongside, same era.
Controllable video generation with sparse trajectories
Zekun Hao, Xun Huang, and Serge Belongie · 2018
Cited alongside, same era.
Video generation from text
Yitong Li, Martin Min, Dinghan Shen, David Carlson, and Lawrence Carin · 2018
Cited alongside, same era.
Mocogan: Decomposing motion and content for video generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz · 2018
Cited alongside, same era.
Diverse beam search for improved description of complex scenes
Ashwin K. Vijayakumar, Michael Cogswell, Ramprasaath R. Selvaraju, Qing Sun, Stefan Lee, David J. Crandall, and Dhruv Batra · 2018
Latent normalizing flows for many-to-many cross-domain mappings
Shweta Mahajan, Iryna Gurevych, and Stefan Roth · 2020
Later among the works it cites.
Diverse image captioning with context-object split latent spaces
Shweta Mahajan and Stefan Roth · 2020
Later among the works it cites.
Character-preserving coherent story visualization
Yun-Zhu Song, Zhi Rui Tam, Hung-Jen Chen, Huiao-Han Lu, and Hong-Han Shuai · 2020
Later among the works it cites.
Retrievegan: Image synthesis via differentiable patch retrieval
Hung-Yu Tseng, Hsin-Ying Lee, Lu Jiang, Ming-Hsuan Yang, and Weilong Yang · 2020
Later among the works it cites.
Cogview: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al · 2021
Later among the works it cites.
Integrating visuospatial, linguistic and commonsense structure into story visualization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Every smile is unique: Landmark-guided diverse smile generation
Wei Wang, Xavier Alameda-Pineda, Dan Xu, Pascal Fua, Elisa Ricci, and Nicu Sebe · 2018
Cited alongside, same era.
Sequential latent spaces for modeling the intention during diverse image captioning
Jyoti Aneja, Harsh Agrawal, Dhruv Batra, and Alexander G. Schwing · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Cited alongside, same era.
End-to-end deep reinforcement learning based coreference resolution
Hongliang Fei, Xu Li, Dingcheng Li, and Ping Li · 2019
Cited alongside, same era.
Storygan: A sequential conditional gan for story visualization
Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao · 2019
Cited alongside, same era.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals · 2019
Cited alongside, same era.
Adyasha Maharana and Mohit Bansal · 2021
Later among the works it cites.
Improving generation and evaluation of visual stories via semantic consistency
Adyasha Maharana, Darryl Hannan, and Mohit Bansal · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Godiva: Generating open-domain videos from natural descriptions
Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan · 2021
Later among the works it cites.
Cross-modal contrastive learning for text-to-image generation
Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang · 2021
Later among the works it cites.
Character-centric story visualization via visual planning and token alignment
Hong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama, and Nanyun Peng · 2022
Closest in time.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Closest in time.
Show me what and tell me how: Video synthesis via multimodal conditioning
Ligong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris Metaxas, and Sergey Tulyakov · 2022
Closest in time.
Mugen: A playground for video-audio-text multimodal understanding and generation
Thomas Hayes, Songyang Zhang, Xi Yin, Guan Pang, Sasha Sheng, Harry Yang, Songwei Ge, Isabelle Hu, and Devi Parikh · 2022
Closest in time.
End-to-end neural bridging resolution
Hideo Kobayashi, Yufang Hou, and Vincent Ng · 2022
Closest in time.
Storydall-e: Adapting pretrained text-to-image transformers for story continuation
Adyasha Maharana, Darryl Hannan, and Mohit Bansal · 2022
Closest in time.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al · 2022
Closest in time.
Sdfusion: Multimodal 3d shape completion, reconstruction, and generation
Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexander Schwing, and Liangyan Gui · 2023
Closest in time.