Fetching the paper…
Reading the bibliography…
In this paper, we present a new text-guided 3D shape generation approach DreamStone that uses images as a stepping stone to bridge the gap between text and shape modalities for generating 3D shapes without requiring paired text and 3D data.
ShapeNet: An Information-Rich 3D Model Repository
A. X. Chang, T. Funkhouser, L. J. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, J. Xiao, L. Yi, and F. Yu · 2015
Earlier work this paper cites.
Generative adversarial text to image synthesis
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee · 2016
Earlier work this paper cites.
Learning what and where to draw
S. E. Reed, Z. Akata, S. Mohan, S. Tenka, B. Schiele, and H. Lee · 2016
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas · 2017
Earlier work this paper cites.
Text2shape: Generating shapes from natural language by learning joint embeddings
K. Chen, C. B. Choy, M. Savva, A. X. Chang, T. Funkhouser, and S. Savarese · 2018
Earlier work this paper cites.
AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He · 2018
Earlier work this paper cites.
StackGAN++: Realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas · 2018
Earlier work this paper cites.
Controllable text-to-image generation
B. Li, X. Qi, T. Lukasiewicz, and P. H. S. Torr · 2019
Earlier work this paper cites.
MirrorGAN: Learning text-to-image generation by redescription
T. Qiao, J. Zhang, D. Xu, and D. Tao · 2019
Earlier work this paper cites.
Bridge-GAN: Interpretable representation learning for text-to-image synthesis
M. Yuan and Y. Peng · 2019
Earlier work this paper cites.
Gamesh: Guided and augmented meshing for deep point networks
N. Agarwal and M. Gopi · 2020
Earlier work this paper cites.
ManiGAN: Text-guided image manipulation
B. Li, X. Qi, T. Lukasiewicz, and P. H. S. Torr · 2020
Earlier work this paper cites.
Differentiable volumetric rendering: Learning implicit 3D representations without 3D supervision
M. Niemeyer, L. Mescheder, M. Oechsle, and A. Geiger · 2020
Earlier work this paper cites.
Network-to-network translation with conditional invertible neural networks
R. Rombach, P. Esser, and B. Ommer · 2020
Earlier work this paper cites.
Efficient neural architecture for text-to-image synthesis
D. M. Souza, J. Wehrmann, and D. D. Ruiz · 2020
Earlier work this paper cites.
Conditional image generation and manipulation for user-specified content
D. Stap, M. Bleeker, S. Ibrahimi, and M. ter Hoeve · 2020
Earlier work this paper cites.
Text to image synthesis with bidirectional generative adversarial network
Z. Wang, Z. Quan, Z.-J. Wang, X. Hu, and Y. Chen · 2020
Cited alongside, same era.
Cogview: Mastering text-to-image generation via transformers
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang, et al · 2021
Cited alongside, same era.
Semantics-guided latent space exploration for shape generation
T. Jahan, Y. Guan, and O. van Kaick · 2021
Cited alongside, same era.
ClipMatrix: Text-controlled creation of 3D textured meshes
N. Jetchev · 2021
Cited alongside, same era.
FuseDream: Training-free text-to-image generation with improved CLIP+ GAN space optimization
X. Liu, C. Gong, L. Wu, S. Zhang, H. Su, and Q. Liu · 2021
Cited alongside, same era.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Zou · 2022
Later among the works it cites.
Towards implicit text-guided 3 d d shape generation
Z. Liu, Y. Wang, X. Qi, and C.-W. Fu · 2022
Later among the works it cites.
Text2mesh: Text-driven neural stylization for meshes
O. Michel, R. Bar-On, R. Liu, S. Benaim, and R. Hanocka · 2022
Later among the works it cites.
Clip-mesh: Generating textured meshes from text using pretrained image-text models
N. Mohammad Khalid, T. Xie, E. Belilovsky, and T. Popa · 2022
Later among the works it cites.
Extracting triangular 3D models, materials, and lighting from images
J. Munkberg, J. Hasselgren, T. Shen, J. Gao, W. Chen, A. Evans, T. Müller, and S. Fidler · 2022
Later among the works it cites.
GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
StyleCLIP: Text-driven manipulation of StyleGAN imagery
O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Cited alongside, same era.
Common objects in 3D: Large-scale learning and evaluation of real-life 3D category reconstruction
J. Reizenstein, R. Shapovalov, P. Henzler, L. Sbordone, P. Labatut, and D. Novotny · 2021
Cited alongside, same era.
Cycle-consistent inverse GAN for text-to-image synthesis
H. Wang, G. Lin, S. Hoi, and C. Miao · 2021
Cited alongside, same era.
TediGAN: Text-guided diverse face image generation and manipulation
W. Xia, Y. Yang, J.-H. Xue, and B. Wu · 2021
Cited alongside, same era.
An effective loss function for generating 3D models from single 2D image without rendering
N. Zubić and P. Liò · 2021
Cited alongside, same era.
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with CLIP latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, et al · 2022
Later among the works it cites.
CLIP-Forge: Towards zero-shot text-to-shape generation
A. Sanghi, H. Chu, J. G. Lambourne, Y. Wang, C.-Y. Cheng, and M. Fumero · 2022
Later among the works it cites.
Stable-DreamFusion: Text-to-3D with stable-diffusion, 2022
J. Tang · 2022
Later among the works it cites.
CLIP-NeRF: Text-and-image driven manipulation of neural radiance fields
C. Wang, M. Chai, M. He, D. Chen, and J. Liao · 2022
Later among the works it cites.
CLIP-GEN: Language-free training of a text-to-image generator with CLIP
Z. Wang, W. Liu, Q. He, X. Wu, and Z. Yi · 2022
Later among the works it cites.
LAFITE: Towards language-free training for text-to-image generation
Y. Zhou, R. Zhang, C. Chen, C. Li, C. Tensmeyer, T. Yu, J. Gu, J. Xu, and T. Sun · 2022
Later among the works it cites.
ISS: Image as stetting stone for text-guided 3D shape generation
Z. Liu, P. Dai, R. Li, X. Qi, and C.-W. Fu · 2023
Closest in time.
DreamFusion: Text-to-3D using 2D diffusion
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2023
Closest in time.