Fetching the paper…
Reading the bibliography…
Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Natural language input for scene generation
Giovanni Adorni and Mauro Di Manzo. 1983 · 1983
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Wordseye: an automatic text-to-scene conversion system
Bob Coyne and Richard Sproat. 2001 · 2001
Earlier work this paper cites.
Fast 3d indoor scene synthesis with discrete and exact layout pattern extraction
Song-Hai Zhang, Shao-Kui Zhang, Wei-Yu Xie, Cheng-Yang Luo, and Hong-Bo Fu. 2020 · 2002
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004 · 2004
Earlier work this paper cites.
Real-time automatic 3d scene generation from natural language voice and text descriptions
Lee M Seversky and Lijun Yin. 2006 · 2006
Earlier work this paper cites.
Interactive furniture layout using interior design guidelines
Paul Merrell, Eric Schkufza, Zeyang Li, Maneesh Agrawala, and Vladlen Koltun. 2011 · 2011
Earlier work this paper cites.
Example-based synthesis of 3d object arrangements
Matthew Fisher, Daniel Ritchie, Manolis Savva, Thomas Funkhouser, and Pat Hanrahan. 2012 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
A fast and accurate dependency parser using neural networks
Danqi Chen and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Sentiment analysis algorithms and applications: A survey
Walaa Medhat, Ahmed Hassan, and Hoda Korashy. 2014 · 2014
Earlier work this paper cites.
Automated simulation creation from military operations documents
John Balint, Jan M Allbeck, and Michael R Hieb. 2015 · 2015
Earlier work this paper cites.
Text to 3D scene generation with rich lexical grounding
Angel Chang, Will Monroe, Manolis Savva, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Activity-centric scene synthesis for functional 3d scene modeling
Matthew Fisher, Manolis Savva, Yangyan Li, Pat Hanrahan, and Matthias Nießner. 2015 · 2015
Earlier work this paper cites.
Generating images from captions with attention
Elman Mansimov, Emilio Parisotto, Jimmy Ba, and Ruslan Salakhutdinov. 2015 · 2015
Cited alongside, same era.
Voxsim: A visual platform for modeling motion language
Nikhil Krishnaswamy and James Pustejovsky. 2016 · 2016
Cited alongside, same era.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Cited alongside, same era.
Pigraphs: learning interaction snapshots from observations
Manolis Savva, Angel X Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nießner. 2016 · 2016
Cited alongside, same era.
Attribute2image: Conditional image generation from visual attributes
Xinchen Yan, Jimei Yang, Kihyuk Sohn, and Honglak Lee. 2016 · 2016
Cited alongside, same era.
A review on deep learning techniques applied to answer selection
Tuan Manh Lai, Trung Bui, and Sheng Li. 2018 · 2018
Later among the works it cites.
Video generation from text
Yitong Li, Martin Renqiang Min, Dinghan Shen, David Carlson, and Lawrence Carin. 2018 · 2018
Later among the works it cites.
Language-driven synthesis of 3d scenes from scene databases
Rui Ma, Akshay Gadi Patil, Matthew Fisher, Manyi Li, Sören Pirk, Binh-Son Hua, Sai-Kit Yeung, Xin Tong, Leonidas Guibas, and Hao Zhang. 2018 · 2018
Later among the works it cites.
Chatpainter: Improving text to image generation using dialogue
Shikhar Sharma, Dendi Suhubdy, Vincent Michalski, Samira Ebrahimi Kahou, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Relationship templates for creating scene variations
Xi Zhao, Ruizhen Hu, Paul Guerrero, Niloy Mitra, and Taku Komura. 2016 · 2016
Cited alongside, same era.
Sceneseer: 3d scene design with natural language
Angel X Chang, Mihail Eric, Manolis Savva, and Christopher D Manning. 2017 · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Cited alongside, same era.
Codraw: Collaborative drawing as a testbed for grounded goal-driven communication
Jin-Hwa Kim, Nikita Kitaev, Xinlei Chen, Marcus Rohrbach, Byoung-Tak Zhang, Yuandong Tian, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
To create what you tell: Generating videos from captions
Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. 2017 · 2017
Cited alongside, same era.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. 2017 · 2017
Cited alongside, same era.
The neural painter: Multi-turn image generation
Ryan Y. Benmalek, Claire Cardie, Serge Belongie, Xiadong He, and Jianfeng Gao. 2018 · 2018
Cited alongside, same era.
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction
Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, Devon Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, and Graham Taylor. 2019 · 2019
Later among the works it cites.
Text2scene: Generating compositional scenes from textual descriptions
Fuwen Tan, Song Feng, and Vicente Ordonez. 2019 · 2019
Later among the works it cites.
Planit: Planning and instantiating indoor scenes with relation graph and spatial prior networks
Kai Wang, Yu-An Lin, Ben Weissmann, Manolis Savva, Angel X Chang, and Daniel Ritchie. 2019 · 2019
Later among the works it cites.
Scenegraphnet: Neural message passing for 3d indoor scene augmentation
Yang Zhou, Zachary While, and Evangelos Kalogerakis. 2019 · 2019
Later among the works it cites.
Blender script
Blender. 2020 · 2020
Closest in time.
Intelligent home 3d: Automatic 3d-house design from linguistic descriptions only
Qi Chen, Qi Wu, Rui Tang, Yuhan Wang, Shuai Wang, and Mingkui Tan. 2020 · 2020
Closest in time.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Rohit Girdhar and Deva Ramanan. 2020 · 2020
Closest in time.
What is artificial general intelligence?
Nick Heath. 2020 · 2020
Closest in time.
Learning spatial knowledge for text to 3d scene generation
Angel Chang, Manolis Savva, and Christopher D Manning. 2014 · 2038
Closest in time.