Fetching the paper…
Reading the bibliography…
Current state-of-the-art methods for text-to-shape generation either require supervised training using a labeled dataset of pre-defined 3D shapes, or perform expensive inference-time optimization of implicit neural representations.
An acceleration framework for high resolution image synthesis
Jinlin Liu, Yuan Yao, and Jianqiang Ren · 1909
Earlier work this paper cites.
Volume rendering
Robert A. Drebin, Loren Carpenter, and Pat Hanrahan · 1988
Earlier work this paper cites.
Spherical sampling by archimedes’ theorem
Min-Zhi Shao and Norman Badler · 1996
Earlier work this paper cites.
GPU gems: programming techniques, tips, and tricks for real-time graphics , volume 590
Randima Fernando et al · 2004
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomás Mikolov · 2013
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li Fei-Fei · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Angel X. Chang, Thomas A. Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qi-Xing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu · 2015
Earlier work this paper cites.
Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling
Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache · 2016
Earlier work this paper cites.
Fast bilateral filtering for denoising large 3d images
Giuseppe Papari, Nasiru Idowu, and Trond Varslot · 2016
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Density estimation using real NVP
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2017
Earlier work this paper cites.
Learning representations and generative models for 3d point clouds
Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas J. Guibas · 2018
Earlier work this paper cites.
Atlasnet: Multi-atlas non-linear deep networks for medical image segmentation
Maria Vakalopoulou, Guillaume Chassagnon, Norbert Bus, Rafael Marini, Evangelia I. Zacharaki, Marie-Pierre Revel, and Nikos Paragios · 2018
Earlier work this paper cites.
Text2shape: Generating shapes from natural language by learning joint embeddings
Kevin Chen, Christopher B. Choy, Manolis Savva, Angel X. Chang, Thomas A. Funkhouser, and Silvio Savarese · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Point cloud GAN
Chun-Liang Li, Manzil Zaheer, Yang Zhang, Barnabás Póczos, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Pointflow: 3d point cloud generation with continuous normalizing flows
Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge J. Belongie, and Bharath Hariharan · 2019
Cited alongside, same era.
LXMERT: learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Cited alongside, same era.
Occupancy networks: Learning 3d reconstruction in function space
Lars M. Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger · 2019
Cited alongside, same era.
Escaping plato’s cave: 3d shape from adversarial rendering
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi · 2021
Later among the works it cites.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal · 2021
Later among the works it cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Björn Ommer · 2021
Later among the works it cites.
Clipdraw: Exploring text-to-drawing synthesis through language-image encoders
Kevin Frans, Lisa B. Soros, and Olaf Witkowski · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Philipp Henzler, Niloy J. Mitra, and Tobias Ritschel · 2019
Cited alongside, same era.
Polygen: An autoregressive generative model of 3d meshes
Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami, and Peter W. Battaglia · 2020
Cited alongside, same era.
Learning visual representations with caption annotations
Mert Bülent Sariyildiz, Julien Perez, and Diane Larlus · 2020
Cited alongside, same era.
Neural voxel renderer: Learning an accurate and controllable rendering tool
Konstantinos Rematas and Vittorio Ferrari · 2020
Cited alongside, same era.
Exploring versatile generative language model via parameter-efficient transfer learning
Zhaojiang Lin, Andrea Madotto, and Pascale Fung · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Real-time scene text detection with differentiable binarization
Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen, and Xiang Bai · 2020
Cited alongside, same era.
Aditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang, Chin-Yi Cheng, Marco Fumero, and Kamal Rahimi Malekshan · 2022
Later among the works it cites.
Zero-shot text-guided object generation with dream fields
Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, and Ben Poole · 2022
Later among the works it cites.
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi · 2022
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao · 2022
Later among the works it cites.
FILIP: fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2022
Later among the works it cites.
Nerf: representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2022
Later among the works it cites.
Dreamfusion: Text-to-3d using 2d diffusion, 2022
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall · 2022
Later among the works it cites.
Clip-nerf: Text-and-image driven manipulation of neural radiance fields
Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Parts2words: Learning joint embedding of point clouds and texts by bidirectional matching between parts and words, 2023
Chuan Tang, Xi Yang, Bojian Wu, Zhizhong Han, and Yi Chang · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang and Maneesh Agrawala · 2023
Closest in time.