Fetching the paper…
Reading the bibliography…
The field of visual computing is rapidly advancing due to the emergence of generative artificial intelligence (AI), which unlocks unprecedented capabilities for the generation, editing, and reconstruction of images, videos, and 3D scenes.
Semantic image synthesis with spatially-adaptive normalization, 2019
Park T., Liu M.-Y., Wang T.-C., Zhu J.-Y · 1903
Earlier work this paper cites.
Plug-and-play diffusion features for text-driven image-to-image translation
Tumanyan N., Geyer M., Bagon S., Dekel T · 1930
Earlier work this paper cites.
Multi-concept customization of text-to-image diffusion
Kumari N., Zhang B., Zhang R., Shechtman E., Zhu J.-Y · 1941
Earlier work this paper cites.
Stochastic differential equations in a differentiable manifold
Itô K · 1950
Earlier work this paper cites.
On a formula concerning stochastic differentials
Itô K · 1951
Earlier work this paper cites.
Reverse-time diffusion equation models
Anderson B. D · 1982
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
Ho J., Jain A., Abbeel P · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng J., Dong W., Socher R., Li L.-J., Li K., Fei-Fei L · 2009
Earlier work this paper cites.
Adversarial score matching and improved sampling for image generation, 2020
Jolicoeur-Martineau A., Piché-Taillefer R., des Combes R. T., Mitliagkas I · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent P · 2011
Earlier work this paper cites.
A dataset of 101 human action classes from videos in the wild
Soomro K., Zamir A. R., Shah M · 2012
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Ionescu C., Papava D., Olaru V., Sminchisescu C · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin T.-Y., Maire M., Belongie S., Hays J., Perona P., Ramanan D., Dollár P., Zitnick C. L · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Chang A. X., Funkhouser T., Guibas L., Hanrahan P., Huang Q., Li Z., Savarese S., Savva M., Song S., Su H., et al · 2015
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Loper M., Mahmood N., Romero J., Pons-Moll G., Black M. J · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger O., Fischer P., Brox T · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision, 2015
Szegedy C., Vanhoucke V., Ioffe S., Shlens J., Wojna Z · 2015
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Kempka M., Wydmuch M., Runc G., Toczek J., Jaśkowski W · 2016
Earlier work this paper cites.
The kit motion-language dataset
Plappert M., Mandery C., Asfour T · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis, 2016
Reed S., Akata Z., Yan X., Logeswaran L., Schiele B., Lee H · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans T., Goodfellow I., Zaremba W., Cheung V., Radford A., Chen X · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu J., Mei T., Yao T., Rui Y · 2016
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
Chang A., Dai A., Funkhouser T., Halber M., Niessner M., Savva M., Song S., Zeng A., Zhang Y · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira J., Zisserman A · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Dai A., Chang A. X., Savva M., Halber M., Funkhouser T., Nießner M · 2017
Earlier work this paper cites.
Carla: An open urban driving simulator
Dosovitskiy A., Ros G., Codevilla F., Lopez A., Koltun V · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel M., Ramsauer H., Unterthiner T., Nessler B., Hochreiter S · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson J., Hariharan B., Van Der Maaten L., Fei-Fei L., Lawrence Zitnick C., Girshick R · 2017
Earlier work this paper cites.
Monocular 3d human pose estimation in the wild using improved cnn supervision
Mehta D., Rhodin H., Casas D., Fua P., Sotnychenko O., Xu W., Theobalt C · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi C. R., Yi L., Su H., Guibas L. J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A. N., Kaiser Ł., Polosukhin I · 2017
Earlier work this paper cites.
Hp-gan: Probabilistic 3d human motion prediction via gan
Barsoum E., Kender J., Liu Z · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin J., Chang M.-W., Lee K., Toutanova K · 2018
Earlier work this paper cites.
Investigating the use of recurrent motion modelling for speech gesture generation
Ferstl Y., McDonnell R · 2018
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks, 2018
Isola P., Zhu J.-Y., Zhou T., Efros A. A · 2018
Earlier work this paper cites.
Learning category-specific mesh reconstruction from image collections
Kanazawa A., Tulsiani S., Efros A. A., Malik J · 2018
Earlier work this paper cites.
Photoshape: Photorealistic materials for large-scale shape collections
Park K., Rematas K., Farhadi A., Seitz S. M · 2018
Earlier work this paper cites.
Pix3d: Dataset and methods for single-image 3d shape modeling
Sun X., Wu J., Zhang X., Zhang Z., Zhang C., Xue T., Tenenbaum J. B., Freeman W. T · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner T., Van Steenkiste S., Kurach K., Marinier R., Michalski M., Gelly S · 2018
Earlier work this paper cites.
Monoperfcap: Human performance capture from monocular video
Xu W., Chatterjee A., Zollhöfer M., Rhodin H., Mehta D., Seidel H., Theobalt C · 2018
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
Zhou T., Tucker R., Flynn J., Fyffe G., Snavely N · 2018
Earlier work this paper cites.
Protecting world leaders against deep fakes
Agarwal S., Farid H., Gu Y., He M., Nagano K., Li H · 2019
Earlier work this paper cites.
Shapeglot: Learning language for shape differentiation
Achlioptas P., Fan J., Hawkins R., Goodman N., Guibas L. J · 2019
Earlier work this paper cites.
Text2shape: Generating shapes from natural language by learning joint embeddings
Chen K., Choy C. B., Savva M., Chang A. X., Funkhouser T., Savarese S · 2019
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
Deng J., Guo J., Xue N., Zafeiriou S · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby N., Giurgiu A., Jastrzebski S., Morrone B., de Laroussilhe Q., Gesmundo A., Attariyan M., Gelly S · 2019
Earlier work this paper cites.
Point-voxel cnn for efficient 3d deep learning
Liu Z., Tang H., Lin Y., Han S · 2019
Earlier work this paper cites.
Amass: Archive of motion capture as surface shapes
Mahmood N., Ghorbani N., Troje N. F., Pons-Moll G., Black M. J · 2019
Earlier work this paper cites.
Hologan: Unsupervised learning of 3d representations from natural images
Nguyen-Phuoc T., Li C., Theis L., Richardt C., Yang Y.-L · 2019
Earlier work this paper cites.
Expressive body capture: 3D hands, face, and body from a single image
Pavlakos G., Choutas V., Ghorbani N., Bolkart T., Osman A. A. A., Tzionas D., Black M. J · 2019
Earlier work this paper cites.
Faceforensics++: Learning to detect manipulated facial images
Rossler A., Cozzolino D., Verdoliva L., Riess C., Thies J., Nießner M · 2019
Earlier work this paper cites.
The replica dataset: A digital replica of indoor spaces
Straub J., Whelan T., Ma L., Chen Y., Wijmans E., Green S., Engel J. J., Mur-Artal R., Ren C., Verma S., et al · 2019
Earlier work this paper cites.
Ego-pose estimation and forecasting as real-time pd control
Yuan Y., Kitani K · 2019
Earlier work this paper cites.
Deepdeform: Learning non-rigid rgb-d reconstruction with semi-supervised data
Bozic A., Zollhofer M., Theobalt C., Nießner M · 2020
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal P., Jun H., Payne C., Kim J. W., Radford A., Sutskever I · 2020
Earlier work this paper cites.
Deepcap: Monocular human performance capture using weak supervision
Habermann M., Xu W., Zollhoefer M., Pons-Moll G., Theobalt C · 2020
Earlier work this paper cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
Krantz J., Wijmans E., Majumdar A., Batra D., Lee S · 2020
Earlier work this paper cites.
Character controllers using motion vaes
Ling H. Y., Zinno F., Cheng G., Van De Panne M · 2020
Earlier work this paper cites.
Convolutional occupancy networks
Peng S., Niemeyer M., Mescheder L., Pollefeys M., Geiger A · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song J., Meng C., Ermon S · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song Y., Sohl-Dickstein J., Kingma D. P., Kumar A., Ermon S., Poole B · 2020
Earlier work this paper cites.
Neural dense non-rigid structure from motion with latent space constraints
Sidhu V., Tretschk E., Golyanik V., Agudo A., Theobalt C · 2020
Earlier work this paper cites.
GRAB: A dataset of whole-body human grasping of objects
Taheri O., Ghorbani N., Black M. J., Tzionas D · 2020
Earlier work this paper cites.
imghum: Implicit generative models of 3d human shape and articulated pose, 2021
Alldieck T., Xu H., Sminchisescu C · 2021
Earlier work this paper cites.
Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data
Baruch G., Chen Z., Dehghan A., Dimry T., Feigin Y., Fu P., Gebauer T., Joffe B., Kurz D., Schwartz A., et al · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani R., Hudson D. A., Adeli E., Altman R., Arora S., von Arx S., Bernstein M. S., Bohg J., Bosselut A., Brunskill E., et al · 2021
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain M., Nagrani A., Varol G., Zisserman A · 2021
Earlier work this paper cites.
Bińkowski M., Sutherland D. J., Arbel M., Gretton A · 2021
Earlier work this paper cites.
Deep generative modelling: A comparative review of vaes, gans, normalizing flows, energy-based and autoregressive models
Bond-Taylor S., Leach A., Long Y., Willcocks C. G · 2021
Earlier work this paper cites.
pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis
Chan E. R., Monteiro M., Kellnhofer P., Wu J., Wetzstein G · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron M., Touvron H., Misra I., Jégou H., Mairal J., Bojanowski P., Joulin A · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis, 2021
Dhariwal P., Nichol A · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser P., Rombach R., Ommer B · 2021
Earlier work this paper cites.
3d-front: 3d furnished rooms with layouts and semantics
Fu H., Cai B., Gao L., Zhang L.-X., Wang J., Li C., Zeng Q., Sun C., Jia R., Zhao B., et al · 2021
Earlier work this paper cites.
VideoForensicsHQ: Detecting high-quality manipulated face videos
Fox G., Liu W., Kim H., Seidel H.-P., Elgharib M., Theobalt C · 2021
Earlier work this paper cites.
Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors
Guzov V., Mir A., Sattler T., Pons-Moll G · 2021
Earlier work this paper cites.
A review on generative adversarial networks: Algorithms, theory, and applications
Gui J., Sun Z., Wen Y., Tao D., Ye J · 2021
Earlier work this paper cites.
Real-time deep dynamic characters
Habermann M., Liu L., Xu W., Zollhoefer M., Pons-Moll G., Theobalt C · 2021
Earlier work this paper cites.
Vlgrammar: Grounded grammar induction of vision and language
Hong Y., Li Q., Zhu S.-C., Huang S · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu E. J., Shen Y., Wallis P., Allen-Zhu Z., Li Y., Wang S., Wang L., Chen W · 2021
Earlier work this paper cites.
Learning speech-driven 3d conversational gestures from video
Habibie I., Xu W., Mehta D., Liu L., Seidel H.-P., Pons-Moll G., Elgharib M., Theobalt C · 2021
Earlier work this paper cites.
Layered neural atlases for consistent video editing
Kasten Y., Ofri D., Wang O., Dekel T · 2021
Earlier work this paper cites.
Diffusion probabilistic models for 3d point cloud generation
Luo S., Hu W · 2021
Earlier work this paper cites.
Neural actor: Neural free-view synthesis of human actors with pose control
Liu L., Habermann M., Rudnev V., Sarkar K., Gu J., Theobalt C · 2021
Earlier work this paper cites.
Infinite nature: Perpetual view generation of natural scenes from a single image
Liu A., Tucker R., Jampani V., Makadia A., Snavely N., Kanazawa A · 2021
Earlier work this paper cites.
4dcomplete: Non-rigid motion estimation beyond the observable surface
Li Y., Takehara H., Taketomi T., Zheng B., Nießner M · 2021
Earlier work this paper cites.
Ai choreographer: Music conditioned 3d dance generation with aist++, 2021
Li R., Yang S., Ross D. A., Kanazawa A · 2021
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Meng C., He Y., Song Y., Song J., Wu J., Zhu J.-Y., Ermon S · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Nichol A. Q., Dhariwal P · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol A., Dhariwal P., Ramesh A., Shyam P., Mishkin P., McGrew B., Sutskever I., Chen M · 2021
Earlier work this paper cites.
BABEL: Bodies, action and behavior with english labels
Punnakkal A. R., Chandrasekaran A., Athanasiou N., Quiros-Ramirez A., Black M. J · 2021
Earlier work this paper cites.
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans
Peng S., Zhang Y., Xu Y., Wang Q., Shuai Q., Bao H., Zhou X · 2021
Earlier work this paper cites.
Humor: 3d human motion model for robust pose estimation
Rempe D., Birdal T., Hertzmann A., Yang J., Sridhar S., Guibas L. J · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford A., Kim J. W., Hallacy C., Ramesh A., Goh G., Agarwal S., Sastry G., Askell A., Mishkin P., Clark J., et al · 2021
Earlier work this paper cites.
Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Reizenstein J., Shapovalov R., Henzler P., Sbordone L., Labatut P., Novotny D · 2021
Earlier work this paper cites.
Buildingnet: Learning to label 3d buildings
Selvaraju P., Nabail M., Loizou M., Maslioukova M., Averkiou M., Andreou A., Chaudhuri S., Kalogerakis E · 2021
Earlier work this paper cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Srinivasan K., Raman K., Chen J., Bendersky M., Najork M · 2021
Earlier work this paper cites.
LAION-400M: open dataset of clip-filtered 400 million image-text pairs
Schuhmann C., Vencu R., Beaumont R., Kaczmarczyk R., Mullis C., Katta A., Coombes T., Jitsev J., Komatsuzaki A · 2021
Earlier work this paper cites.
Lasr: Learning articulated shape reconstruction from a monocular video
Yang G., Sun D., Jampani V., Vlasic D., Cole F., Chang H., Ramanan D., Freeman W. T., Liu C · 2021
Earlier work this paper cites.
pixelNeRF: Neural radiance fields from one or few images
Yu A., Ye V., Tancik M., Kanazawa A · 2021
Earlier work this paper cites.
3d shape generation and completion through point-voxel diffusion
Zhou L., Du Y., Wu J · 2021
Earlier work this paper cites.
Blended diffusion for text-driven editing of natural images
Avrahami O., Lischinski D., Fried O · 2022
Earlier work this paper cites.
Gaudi: A neural architect for immersive 3d scene generation
Bautista M. A., Guo P., Abnar S., Talbott W., Toshev A., Chen Z., Dinh L., Zhai S., Goh H., Ulbricht D., et al · 2022
Earlier work this paper cites.
Generating long videos of dynamic scenes
Brooks T., Hellsten J., Aittala M., Wang T.-C., Aila T., Lehtinen J., Liu M.-Y., Efros A., Karras T · 2022
Earlier work this paper cites.
Generative neural articulated radiance fields
Bergman A., Kellnhofer P., Yifan W., Chan E., Lindell D., Wetzstein G · 2022
Earlier work this paper cites.
Retrieval-augmented diffusion models
Blattmann A., Rombach R., Oktay K., Müller J., Ommer B · 2022
Earlier work this paper cites.
Text2live: Text-driven layered image and video editing
Bar-Tal O., Ofri-Amar D., Fridman R., Kasten Y., Dekel T · 2022
Cited alongside, same era.
Behave: Dataset and method for tracking human object interactions
Bhatnagar B. L., Xie X., Petrov I., Sminchisescu C., Theobalt C., Pons-Moll G · 2022
Cited alongside, same era.
Abo: Dataset and benchmarks for real-world 3d object understanding
Collins J., Goel S., Deng K., Luthra A., Xu L., Gundogdu E., Zhang X., Vicente T. F. Y., Dideriksen T., Arora H., et al · 2022
Cited alongside, same era.
Efficient geometry-aware 3d generative adversarial networks
Chan E. R., Lin C. Z., Chan M. A., Nagano K., Pan B., De Mello S., Gallo O., Guibas L. J., Tremblay J., Khamis S., et al · 2022
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
Deitke M., Schwenk D., Salvador J., Weihs L., Michel O., VanderBilt E., Schmidt L., Ehsani K., Kembhavi A., Farhadi A · 2022
Cited alongside, same era.
Avatarcraft: Transforming text into neural human avatars with parameterized shape and pose control
Jiang R., Wang C., Zhang J., Chai M., He M., Chen D., Liao J · 2023
Closest in time.
Taming encoder for zero fine-tuning image customization with text-to-image diffusion models
Jia X., Zhao Y., Chan K. C., Li Y., Zhang H., Gong B., Hou T., Wang H., Su Y.-C · 2023
Closest in time.
Dreamhuman: Animatable 3d avatars from text
Kolotouros N., Alldieck T., Zanfir A., Bazavan E. G., Fieraru M., Sminchisescu C · 2023
Closest in time.
Neuralfield-ldm: Scene generation with hierarchical latent diffusion models
Kim S. W., Brown B., Yin K., Kreis K., Schwarz K., Li D., Rombach R., Torralba A., Fidler S · 2023
Closest in time.
Diffusion models in practice. part 1: The tools of the trade
Kochanowicz J., Domagała M., Stachowiak D., Dziedzic K · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Genie: Higher-order denoising diffusion solvers
Dockhorn T., Vahdat A., Kreis K · 2022
Cited alongside, same era.
Plenoxels: Radiance fields without neural networks
Fridovich-Keil S., Yu A., Tancik M., Chen Q., Recht B., Kanazawa A · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal R., Alaluf Y., Atzmon Y., Patashnik O., Bermano A. H., Chechik G., Cohen-Or D · 2022
Cited alongside, same era.
Generating diverse and natural 3d human motions from text
Guo C., Zou S., Zuo X., Wang S., Ji W., Li X., Cheng L · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho J., Chan W., Saharia C., Whang J., Gao R., Gritsenko A., Kingma D. P., Poole B., Norouzi M., Fleet D. J., et al · 2022
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong W., Ding M., Zheng W., Liu X., Tang J · 2022
Cited alongside, same era.
Neural wavelet-domain diffusion for 3d shape generation
Hui K.-H., Li R., Hu J., Fu C.-W · 2022
Cited alongside, same era.
Closest in time.
Dreampose: Fashion image-to-video synthesis via stable diffusion
Karras J., Holynski A., Wang T.-C., Kemelmacher-Shlizerman I · 2023
Closest in time.
Flame: Free-form language-based motion synthesis & editing
Kim J., Kim J., Choi S · 2023
Closest in time.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan L., Movsisyan A., Tadevosyan V., Henschel R., Wang Z., Navasardyan S., Shi H · 2023
Closest in time.
Holofusion: Towards photo-realistic 3d generative modeling
Karnewar A., Mitra N. J., Vedaldi A., Novotny D · 2023
Closest in time.
Gmd: Controllable human motion synthesis via guided diffusion models
Karunratanakul K., Preechakul K., Suwajanakorn S., Tang S · 2023
Closest in time.
Nersemble: Multi-view radiance field reconstruction of human heads, 2023
Kirschstein T., Qian S., Giebenhain S., Walter T., Nießner M · 2023
Closest in time.
Nifty: Neural object interaction fields for guided human motion synthesis, 2023
Kulkarni N., Rempe D., Genova K., Kundu A., Johnson J., Fouhey D., Guibas L · 2023
Closest in time.
Holodiffusion: Training a 3d diffusion model using 2d images
Karnewar A., Vedaldi A., Novotny D., Mitra N. J · 2023
Closest in time.
Imagic: Text-based real image editing with diffusion models
Kawar B., Zada S., Lang O., Tov O., Chang H., Dekel T., Mosseri I., Irani M · 2023
Closest in time.
Deepfloyd if
Lab D · 2023
Closest in time.
Li X., Chu W., Wu Y., Yuan W., Liu F., Zhang Q., Li F., Feng H., Ding E., Wang J · 2023
Closest in time.
Videofusion: Decomposed diffusion models for high-quality video generation
Luo Z., Chen D., Zhang Y., Huang Y., Wang L., Shen Y., Zhao D., Zhou J., Tan T · 2023
Closest in time.
Diffusion hyperfeatures: Searching through time and space for semantic correspondence
Luo G., Dunlap L., Park D. H., Holynski A., Darrell T · 2023
Closest in time.
Nap: Neural 3d articulation prior
Lei J., Deng C., Shen B., Guibas L., Daniilidis K · 2023
Closest in time.
Diffusion-SDF: Text-to-shape via voxelized diffusion
Li M., Duan Y., Zhou J., Lu J · 2023
Closest in time.
Magic3d: High-resolution text-to-3d content creation
Lin C.-H., Gao J., Tang L., Takikawa T., Zeng X., Huang X., Kreis K., Fidler S., Liu M.-Y., Lin T.-Y · 2023
Closest in time.
Shape-aware text-driven layered video editing
Lee Y.-C., Jang J.-Z. G., Chen Y.-T., Qiu E., Huang J.-B · 2023
Closest in time.
Syncdreamer: Learning to generate multiview-consistent images from a single-view image
Liu Y., Lin C., Zeng Z., Long X., Liu L., Komura T., Wang W · 2023
Closest in time.
Li Z., Tucker R., Snavely N., Holynski A · 2023
Closest in time.
The carbon footprint of gpt-4
Ludvigsen K. G. A · 2023
Closest in time.
Zero-1-to-3: Zero-shot one image to 3d object
Liu R., Wu R., Hoorick B. V., Tokmakov P., Zakharov S., Vondrick C · 2023
Closest in time.
Snapfusion: Text-to-image diffusion model on mobile devices within two seconds
Li Y., Wang H., Jin Q., Hu J., Chemerys P., Fu Y., Wang Y., Tulyakov S., Ren J · 2023
Closest in time.
Object motion guided human motion synthesis
Li J., Wu J., Liu C. K · 2023
Closest in time.
A large-scale outdoor multi-modal dataset and benchmark for novel view synthesis and implicit scene reconstruction
Lu C., Yin F., Chen X., Liu W., Chen T., Yu G., Fan J · 2023
Closest in time.
Tada! text to animatable digital avatars
Liao T., Yi H., Xiu Y., Tang J., Huang Y., Thies J., Black M. J · 2023
Closest in time.
Magicedit: High-fidelity and temporally coherent video editing
Liew J. H., Yan H., Zhang J., Xu Z., Feng J · 2023
Closest in time.
Intergen: Diffusion-based multi-human motion generation under complex interactions
Liang H., Zhang W., Li W., Yu J., Xu L · 2023
Closest in time.
Motion-x: A large-scale 3d expressive whole-body human motion dataset, 2023
Lin J., Zeng A., Lu S., Cai Y., Zhang R., Wang H., Zhang L · 2023
Closest in time.
Video-p2p: Video editing with cross-attention control
Liu S., Zhang Y., Li W., Lin Z., Jia J · 2023
Closest in time.
Generative ai meets 3d: A survey on text-to-3d in aigc era
Li C., Zhang C., Waghwase A., Lee L.-H., Rameau F., Yang Y., Bae S.-H., Hong C. S · 2023
Closest in time.
Null-text inversion for editing real images using guided diffusion models
Mokady R., Hertz A., Aberman K., Pritch Y., Cohen-Or D · 2023
Closest in time.
Midjourney
Midjourney · 2023
Closest in time.
Avatarstudio: Text-driven editing of 3d dynamic human head avatars
Mendiratta M., Pan X., Elgharib M., Teotia K., R M. B., Tewari A., Golyanik V., Kortylewski A., Theobalt C · 2023
Closest in time.
On distillation of guided diffusion models, 2023
Meng C., Rombach R., Gao R., Kingma D. P., Ermon S., Ho J., Salimans T · 2023
Closest in time.
Latent-nerf for shape-guided generation of 3d shapes and textures
Metzer G., Richardson E., Patashnik O., Giryes R., Cohen-Or D · 2023
Closest in time.
Plotting behind the scenes: Towards learnable game engines
Menapace W., Siarohin A., Lathuilière S., Achlioptas P., Golyanik V., Ricci E., Tulyakov S · 2023
Closest in time.
Diffrf: Rendering-guided 3d radiance field diffusion
Müller N., Siddiqui Y., Porzi L., Bulo S. R., Kontschieder P., Nießner M · 2023
Closest in time.
Dragondiffusion: Enabling drag-style manipulation on diffusion models
Mou C., Wang X., Song J., Shan Y., Zhang J · 2023
Closest in time.
Mou C., Wang X., Xie L., Zhang J., Qi Z., Shan Y., Qie X · 2023
Closest in time.
DALL·E 2 — openai.com
OpenAI · 2023
Closest in time.
DALL·E 3 — openai.com
OpenAI · 2023
Closest in time.
Codef: Content deformation fields for temporally consistent video processing
Ouyang H., Wang Q., Xiao Y., Bai Q., Zhang J., Zheng K., Zhou X., Chen Q., Shen Y · 2023
Closest in time.
Zero-shot image-to-image translation
Parmar G., Kumar Singh K., Zhang R., Li Y., Lu J., Zhu J.-Y · 2023
Closest in time.
Drag your gan: Interactive point-based manipulation on the generative image manifold
Pan X., Tewari A., Leimkühler T., Liu L., Meka A., Theobalt C · 2023
Closest in time.
Compositional 3d scene generation using locally conditioned diffusion
Po R., Wetzstein G · 2023
Closest in time.
Fatezero: Fusing attentions for zero-shot text-based video editing
Qi C., Cun X., Zhang Y., Lei C., Wang X., Shan Y., Chen Q · 2023
Closest in time.
Trace and pace: Controllable pedestrian animation via guided trajectory diffusion
Rempe D., Luo Z., Bin Peng X., Yuan Y., Kitani K., Kreis K., Fidler S., Litany O · 2023
Closest in time.
Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models
Ruiz N., Li Y., Jampani V., Wei W., Hou T., Pritch Y., Wadhwa N., Rubinstein M., Aberman K · 2023
Closest in time.
3d neural field generation using triplane diffusion
Shue J. R., Chan E. R., Po R., Ankner Z., Wu J., Wetzstein G · 2023
Closest in time.
Song Y., Dhariwal P., Chen M., Sutskever I · 2023
Closest in time.
Vox-e: Text-guided voxel editing of 3d objects, 2023
Sella E., Fiebelman G., Hedman P., Averbuch-Elor H · 2023
Closest in time.
Clip-sculptor: Zero-shot generation of high-fidelity and diverse shapes from natural language
Sanghi A., Fu R., Liu V., Willis K. D., Shayani H., Khasahmadi A. H., Sridhar S., Ritchie D · 2023
Closest in time.
Sketchfab — sketchfab.com
Sketchfab · 2023
Closest in time.
Styledrop: Text-to-image generation in any style
Sohn K., Ruiz N., Lee K., Chin D. C., Blok I., Chang H., Barber J., Jiang L., Entis G., Li Y., Hao Y., Essa I., Rubinstein M., Krishnan D · 2023
Closest in time.
Viewset diffusion: (0-)image-conditioned 3d generative models from 2d data
Szymanowicz S., Rupprecht C., Vedaldi A · 2023
Closest in time.
Control4d: Dynamic portrait editing by learning 4d gan from 2d diffusion-based editor
Shao R., Sun J., Peng C., Zheng Z., Zhou B., Zhang H., Liu Y · 2023
Closest in time.
Text-to-4d dynamic scene generation
Singer U., Sheynin S., Polyak A., Ashual O., Makarov I., Kokkinos F., Goyal N., Vedaldi A., Parikh D., Johnson J., Taigman Y · 2023
Closest in time.
Human Motion Diffusion as a Generative Prior
Shafir Y., Tevet G., Kapon R., Bermano A. H · 2023
Closest in time.
Mvdream: Multi-view diffusion for 3d generation
Shi Y., Wang P., Ye J., Mai L., Li K., Yang X · 2023
Closest in time.
Instantbooth: Personalized text-to-image generation without test-time finetuning
Shi J., Xiong W., Lin Z., Jung H. J · 2023
Closest in time.
Dragdiffusion: Harnessing diffusion models for interactive point-based image editing
Shi Y., Xue C., Pan J., Zhang W., Tan V. Y., Bai S · 2023
Closest in time.
Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering
Shao R., Zheng Z., Tu H., Liu B., Zhang H., Liu Y · 2023
Closest in time.
Edge: Editable dance generation from music
Tseng J., Castellon R., Liu K · 2023
Closest in time.
Emergent correspondence from image diffusion
Tang L., Jia M., Wang Q., Phoo C. P., Hariharan B · 2023
Closest in time.
Realfill: Reference-driven generation for authentic image completion
Tang L., Ruiz N., Chu Q., Li Y., Holynski A., Jacobs D. E., Hariharan B., Pritch Y., Wadhwa N., Aberman K., et al · 2023
Closest in time.
Human motion diffusion model
Tevet G., Raab S., Gordon B., Shafir Y., Cohen-or D., Bermano A. H · 2023
Closest in time.
Diffusion with forward models: Solving stochastic inverse problems without direct supervision
Tewari A., Yin T., Cazenavette G., Rezchikov S., Tenenbaum J. B., Durand F., Freeman W. T., Sitzmann V · 2023
Closest in time.
Sketch-guided text-to-image diffusion models
Voynov A., Aberman K., Cohen-Or D · 2023
Closest in time.
p + p+ : Extended textual conditioning in text-to-image generation
Voynov A., Chu Q., Cohen-Or D., Aberman K · 2023
Closest in time.
Edict: Exact diffusion inversion via coupled transformations
Wallace B., Gokul A., Naik N · 2023
Closest in time.
Sunstage: Portrait reconstruction and relighting using the sun as a light stage
Wang Y., Holynski A., Zhang X., Zhang X · 2023
Closest in time.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Wang Z., Lu C., Wang Y., Bao F., Li C., Su H., Zhu J · 2023
Closest in time.
Modelscope text-to-video technical report
Wang J., Yuan H., Chen D., Zhang Y., Wang X., Zhang S · 2023
Closest in time.
Videocomposer: Compositional video synthesis with motion controllability
Wang X., Yuan H., Zhang S., Chen D., Wang J., Zhang Y., Shen Y., Zhao D., Zhou J · 2023
Closest in time.
Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation
Wu T., Zhang J., Fu X., Wang Y., Ren J., Pan L., Wu W., Yang L., Wang J., Qian C., et al · 2023
Closest in time.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation
Wei Y., Zhang Y., Ji Z., Bai J., Zhang L., Zuo W · 2023
Closest in time.
RODIN: A generative model for sculpting 3d digital avatars using diffusion
Wang T., Zhang B., Zhang T., Gu S., Bao J., Baltrusaitis T., Shen J., Chen D., Wen F., Chen Q., Guo B · 2023
Closest in time.
Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding
Xue L., Gao M., Xing C., Martín-Martín R., Wu J., Xiong C., Xu R., Niebles J. C., Savarese S · 2023
Closest in time.
Fastcomposer: Tuning-free multi-subject image generation with localized attention
Xiao G., Yin T., Freeman W. T., Durand F., Han S · 2023
Closest in time.
Scannet++: A high-fidelity dataset of 3d indoor scenes
Yeshwanth C., Liu Y.-C., Nießner M., Dai A · 2023
Closest in time.
Artic3d: Learning robust articulated 3d shapes from noisy web image collections
Yao C.-H., Raj A., Hung W.-C., Li Y., Rubinstein M., Yang M.-H., Jampani V · 2023
Closest in time.
Physdiff: Physics-guided human motion diffusion model
Yuan Y., Song J., Iqbal U., Vahdat A., Kautz J · 2023
Closest in time.
Video probabilistic diffusion models in projected latent space
Yu S., Sohn K., Kim S., Shin J · 2023
Closest in time.
Emog: Synthesizing emotive co-speech 3d gesture with diffusion model
Yin L., Wang Y., He T., Liu J., Zhao W., Li B., Jin X., Lin J · 2023
Closest in time.
Diffusestylegesture: Stylized audio-driven co-speech gesture generation with diffusion models
Yang S., Wu Z., Li M., Zhang Z., Hao L., Bao W., Cheng M., Xiao L · 2023
Closest in time.
Rerender a video: Zero-shot text-guided video-to-video translation
Yang S., Zhou Y., Liu Z., Loy C. C · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Zhang L., Agrawala M · 2023
Closest in time.
Zou Z.-X., Cheng W., Cao Y.-P., Huang S.-S., Shan Y., Zhang S.-H · 2023
Closest in time.
Teca: Text-guided generation and editing of compositional 3d avatars
Zhang H., Feng Y., Kulits P., Wen Y., Thies J., Black M. J · 2023
Closest in time.
4D Facial Expression Diffusion Model
Zou K., Faisan S., Yu B., Valette S., Seo H · 2023
Closest in time.
Diffmotion: Speech-driven gesture synthesis using denoising diffusion model
Zhang F., Ji N., Gao F., Li Y · 2023
Closest in time.
Tedi: Temporally-entangled diffusion for long-term motion synthesis
Zhang Z., Liu R., Aberman K., Hanocka R · 2023
Closest in time.
Zhao Z., Liu W., Chen X., Zeng X., Wang R., Cheng P., Fu B., Chen T., Yu G., Gao S · 2023
Closest in time.
Locally attentional SDF diffusion for controllable 3d shape generation
Zheng X.-Y., Pan H., Wang P.-S., Tong X., Liu Y., Shum H.-Y · 2023
Closest in time.
Dreamface: Progressive generation of animatable 3d faces under text guidance
Zhang L., Qiu Q., Lin H., Zhang Q., Shi C., Yang W., Shi Y., Yang S., Xu L., Yu J · 2023
Closest in time.
Sparsefusion: Distilling view-conditioned diffusion for 3d reconstruction
Zhou Z., Tulsiani S · 2023
Closest in time.
3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models
Zhang B., Tang J., Nießner M., Wonka P · 2023
Closest in time.
Dreameditor: Text-driven 3d scene editing with neural fields
Zhuang J., Wang C., Liu L., Lin L., Li G · 2023
Closest in time.
Multimodal image synthesis and editing: The generative ai era
Zhan F., Yu Y., Wu R., Zhang J., Lu S., Liu L., Kortylewski A., Theobalt C., Xing E · 2023
Closest in time.
A survey of large language models
Zhao W. X., Zhou K., Li J., Tang T., Wang X., Hou Y., Min Y., Zhang B., Zhang J., Dong Z., Du Y., Yang C., Chen Y., Chen Z., Jiang J., Ren R., Li Y., Tang X., Liu Z., Liu P., Nie J., rong Wen J · 2023
Closest in time.
Action2motion: Conditioned generation of 3d human motions
Guo C., Zuo X., Wang S., Zou S., Sun Q., Deng A., Gong M., Cheng L · 2029
Closest in time.