Fetching the paper…
Reading the bibliography…
Generating shapes using natural language can enable new ways of imagining and creating the things around us.
WordNet: An Electronic Lexical Database
Christiane Fellbaum · 1998
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
Christoph Lampert, Hannes Nickisch, and Stefan Harmeling · 2009
Earlier work this paper cites.
Zero-shot learning with semantic output codes
Mark Palatucci, Dean Pomerleau, Geoffrey E Hinton, and Tom M Mitchell · 2009
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li Fei-Fei · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V Le, Christopher D Manning, and Andrew Y Ng · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository, 2015
Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu · 2015
Earlier work this paper cites.
Nice: Non-linear independent components estimation, 2015
Laurent Dinh, David Krueger, and Yoshua Bengio · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
3d-r2n2: A unified approach for single and multi-view 3d object reconstruction
Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese · 2016
Earlier work this paper cites.
Density estimation using real nvp
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens Van Der Maaten, Allan Jabri, and Nicolas Vasilache · 2016
Earlier work this paper cites.
Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling
Jiajun Wu, Chengkai Zhang, Tianfan Xue, William T. Freeman, and Joshua B. Tenenbaum · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Masked autoregressive flow for density estimation
George Papamakarios, Theo Pavlakou, and Iain Murray · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Learning representations and generative models for 3d point clouds
Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas · 2018
Cited alongside, same era.
Learning to refer to 3d objects with natural language
Panos Achlioptas, Judy Fan, Robert X. D. Hawkins, Noah D. Goodman, and Leonidas J. Guibas · 2018
Cited alongside, same era.
Text2shape: Generating shapes from natural language by learning joint embeddings
Kevin Chen, Christopher B Choy, Manolis Savva, Angel X Chang, Thomas Funkhouser, and Silvio Savarese · 2018
Cited alongside, same era.
Glow: Generative flow with invertible 1x1 convolutions, 2018
Diederik P. Kingma and Prafulla Dhariwal · 2018
Cited alongside, same era.
Chun-Liang Li, Manzil Zaheer, Yang Zhang, Barnabas Poczos, and Ruslan Salakhutdinov · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Holistic static and animated 3d scene generation from diverse text descriptions
Faria Huq, Anindya Iqbal, and Nafees Ahmed · 2020
Later among the works it cites.
Videoflow: A conditional flow-based model for stochastic video generation, 2020
Manoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, and Durk Kingma · 2020
Later among the works it cites.
Polygen: An autoregressive generative model of 3d meshes
Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami, and Peter W. Battaglia · 2020
Later among the works it cites.
Convolutional occupancy networks
Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, and Andreas Geiger · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generating animated videos of human activities from natural language descriptions
Angela S. Lin, Lemeng Wu, Rodolfo Corona, Kevin W. H. Tai, Qixing Huang, and Raymond J. Mooney · 2018
Cited alongside, same era.
Pixel2mesh: Generating 3d mesh models from single rgb images
Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang · 2018
Cited alongside, same era.
Foldingnet: Point cloud auto-encoder via deep grid deformation
Yaoqing Yang, Chen Feng, Yiru Shen, and Dong Tian · 2018
Cited alongside, same era.
Learning implicit fields for generative shape modeling
Zhiqin Chen and Hao Zhang · 2019
Cited alongside, same era.
Y2seq2seq: Cross-modal representation learning for 3d shape and text by joint reconstruction and prediction of view and word sequences
Zhizhong Han, Mingyang Shang, Xiyang Wang, Yu-Shen Liu, and Matthias Zwicker · 2019
Cited alongside, same era.
Flowavenet : A generative flow for raw audio, 2019
Sungwon Kim, Sang gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon · 2019
Cited alongside, same era.
An acceleration framework for high resolution image synthesis
Jinlin Liu, Yuan Yao, and Jianqiang Ren · 2019
Cited alongside, same era.
Later among the works it cites.
C-flow: Conditional generative flow models for images and 3d point clouds, 2020
Albert Pumarola, Stefan Popov, Francesc Moreno-Noguer, and Vittorio Ferrari · 2020
Later among the works it cites.
Learning visual representations with caption annotations
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus · 2020
Later among the works it cites.
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2021
Closest in time.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Closest in time.
Clip2video: Mastering video-text retrieval via image clip
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Closest in time.
Clipdraw: Exploring text-to-drawing synthesis through language-image encoders, 2021
Kevin Frans, L. B. Soros, and Olaf Witkowski · 2021
Closest in time.
Text-guided graph neural networks for referring 3d instance segmentation
Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, and Tyng-Luh Liu · 2021
Closest in time.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Closest in time.
Segmentation in style: Unsupervised semantic image segmentation with stylegan and clip
Daniil Pakhomov, Sanchit Hira, Narayani Wagle, Kemar E Green, and Nassir Navab · 2021
Closest in time.
Styleclip: Text-driven manipulation of stylegan imagery
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski · 2021
Closest in time.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Closest in time.
Zero-shot text-to-image generation, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Closest in time.
How much can clip benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2021
Closest in time.
Part2word: Learning joint embedding of point clouds and text by matching parts to words, 2021
Chuan Tang, Xi Yang, Bojian Wu, Zhizhong Han, and Yi Chang · 2021
Closest in time.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2021
Closest in time.