Fetching the paper…
Reading the bibliography…
As information exists in various modalities in real world, effective interaction and fusion among multimodal information plays a key role for the creation and perception of multimodal data in computer vision and deep learning research.
Distributional structure
Z. S. Harris · 1954
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
A morphable model for the synthesis of 3d faces
V. Blanz et al · 1999
Earlier work this paper cites.
Facial action coding system (facs) a human face
P. Ekman et al · 2002
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng et al · 2009
Earlier work this paper cites.
Caltech-ucsd birds 200
P. Welinder et al · 2010
Earlier work this paper cites.
Indoor scene segmentation using a structured light sensor
N. Silberman and R. Fergus · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov et al · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio et al · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause et al · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow et al · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
M. Mirza and S. Osindero · 2014
Earlier work this paper cites.
Deep autoregressive networks
K. Gregor et al · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin et al · 2014
Earlier work this paper cites.
Semantic image segmentation with deep convolutional nets and fully connected crfs
L.-C. Chen et al · 2014
Earlier work this paper cites.
Generating images from captions with attention
E. Mansimov et al · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
X. Shi et al · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein et al · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
D. Rezende and S. Mohamed · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Z. Liu et al · 2015
Earlier work this paper cites.
Deep human parsing with active template regression
X. Liang et al · 2015
Earlier work this paper cites.
Generative adversarial text to image synthesis
S. Reed et al · 2016
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar et al · 2016
Earlier work this paper cites.
Visually indicated sounds
A. Owens et al · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson et al · 2016
Earlier work this paper cites.
Density estimation using real nvp
L. Dinh et al · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
A. Van Oord et al · 2016
Earlier work this paper cites.
Discrete variational autoencoders
J. T. Rolfe · 2016
Earlier work this paper cites.
Discriminative regularization for generative models
A. Lamb et al · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
A. B. L. Larsen et al · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
E. Jang et al · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
A. Van den Oord et al · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
M. Cordts et al · 2016
Earlier work this paper cites.
Deepfashion: Powering robust clothes recognition and retrieval with rich annotations
Z. Liu et al · 2016
Earlier work this paper cites.
Lip reading in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans et al · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
P. Isola et al · 2017
Earlier work this paper cites.
Learning from simulated and unsupervised images through adversarial training
A. Shrivastava et al · 2017
Earlier work this paper cites.
J. S. Chung et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. v. d. Oord et al · 2017
Earlier work this paper cites.
StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks
H. Zhang et al · 2017
Earlier work this paper cites.
Pose guided person image generation
L. Ma et al · 2017
Earlier work this paper cites.
Toward multimodal image-to-image translation
J.-Y. Zhu et al · 2017
Earlier work this paper cites.
Learning word-like units from joint audio-visual analysis
D. Harwath and J. R. Glass · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
S. Suwajanakorn et al · 2017
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
T. Karras et al · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
J.-Y. Zhu et al · 2017
Earlier work this paper cites.
Plug & play generative networks: Conditional iterative generation of images in latent space
A. Nguyen et al · 2017
Earlier work this paper cites.
One-sided unsupervised domain mapping
S. Benaim and L. Wolf · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
B. Zhou et al · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna et al · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
A. Nagrani et al · 2017
Earlier work this paper cites.
Lip reading sentences in the wild
J. Son Chung et al · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel et al · 2017
Earlier work this paper cites.
Dilated residual networks
F. Yu et al · 2017
Earlier work this paper cites.
High-resolution image synthesis and semantic manipulation with conditional gans
T.-C. Wang et al · 2018
Earlier work this paper cites.
Diverse image-to-image translation via disentangled representations
H.-Y. Lee et al · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
T. Xu et al · 2018
Earlier work this paper cites.
Sketchygan: Towards diverse and realistic sketch to image synthesis
W. Chen and J. Hays · 2018
Earlier work this paper cites.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
H. Zhang et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2018
Earlier work this paper cites.
Vision as an interlingua: Learning multilingual semantic embeddings of untranscribed speech
D. Harwath et al · 2018
Earlier work this paper cites.
Talking face generation by conditional recurrent adversarial network
Y. Song et al · 2018
Earlier work this paper cites.
Image generation from scene graphs
J. Johnson et al · 2018
Earlier work this paper cites.
Photographic text-to-image synthesis with a hierarchically-nested adversarial network
Z. Zhang et al · 2018
Earlier work this paper cites.
Perceptual adversarial networks for image-to-image transformation
C. Wang et al · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord et al · 2018
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
A. Brock et al · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang et al · 2018
Earlier work this paper cites.
Generating high fidelity images with subscale pixel networks and multidimensional upscaling
J. Menick and N. Kalchbrenner · 2018
Earlier work this paper cites.
Vggface2: A dataset for recognising faces across pose and age
Q. Cao et al · 2018
Earlier work this paper cites.
Image transformer
N. Parmar et al · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez et al · 2018
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
H. Caesar et al · 2018
Earlier work this paper cites.
Visual object networks: Image generation with disentangled 3d representations
J.-Y. Zhu et al · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
J. S. Chung et al · 2018
Earlier work this paper cites.
Inferring semantic layout for hierarchical text-to-image synthesis
S. Hong et al · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
T. Xiao et al · 2018
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
T. Park et al · 2019
Earlier work this paper cites.
Pastegan: A semi-parametric method to generate image from scene graph
Y. Li et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford et al · 2019
Earlier work this paper cites.
Mirrorgan: Learning text-to-image generation by redescription
T. Qiao et al · 2019
Earlier work this paper cites.
A deep collaborative framework for face photo–sketch synthesis
M. Zhu et al · 2019
Earlier work this paper cites.
Image synthesis from reconfigurable layout and style
W. Sun and T. Wu · 2019
Earlier work this paper cites.
Image generation from layout
B. Zhao et al · 2019
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
L. Chen et al · 2019
Earlier work this paper cites.
Multi-channel attention selection gan with cascaded semantic guidance for cross-view image translation
H. Tang et al · 2019
Earlier work this paper cites.
Semantics-enhanced adversarial nets for text-to-image synthesis
H. Tan et al · 2019
Earlier work this paper cites.
Controllable text-to-image generation
B. Li et al · 2019
Earlier work this paper cites.
Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis
M. Zhu et al · 2019
Earlier work this paper cites.
Talking face generation by adversarially disentangled audio-visual representation
H. Zhou et al · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
T. Karras et al · 2019
Earlier work this paper cites.
Semantics disentangling for text-to-image generation
G. Yin et al · 2019
Earlier work this paper cites.
Adversarial learning of semantic relevance in text to image synthesis
M. Cha et al · 2019
Earlier work this paper cites.
Travelgan: Image-to-image translation by transformation vector learning
M. Amodio and S. Krishnaswamy · 2019
Earlier work this paper cites.
Cycle in cycle generative adversarial networks for keypoint-guided image generation
H. Tang et al · 2019
Earlier work this paper cites.
Dual adversarial inference for text-to-image synthesis
Q. Lao et al · 2019
Earlier work this paper cites.
Cycle-consistent diverse image synthesis from natural language
Z. Chen and Y. Luo · 2019
Cited alongside, same era.
Geometry-consistent generative adversarial networks for one-sided unsupervised domain mapping
H. Fu et al · 2019
Cited alongside, same era.
Generating diverse high-fidelity images with vq-vae-2
A. Razavi et al · 2019
Cited alongside, same era.
vq-wav2vec: Self-supervised learning of discrete speech representations
A. Baevski et al · 2019
Cited alongside, same era.
A. Noguchi and T. Harada · 2019
Cited alongside, same era.
Transgan: Two pure transformers can make one strong gan, and that can scale up
Y. Jiang et al · 2021
Closest in time.
Generative adversarial transformers
D. A. Hudson and C. L. Zitnick · 2021
Closest in time.
Multimodal conditional image synthesis with product-of-experts gans
X. Huang et al · 2021
Closest in time.
Efficient semantic image synthesis via class-adaptive normalization
Z. Tan et al · 2021
Closest in time.
Diverse semantic image synthesis via probability distribution modeling
Z. Tan et al · 2021
Closest in time.
Parallel and flexible sampling from autoregressive models via langevin dynamics
V. Jayaram and J. Thickstun · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Nguyen-Phuoc et al · 2019
Cited alongside, same era.
Learning to predict layout-to-image conditional convolutions for semantic image synthesis
X. Liu et al · 2019
Cited alongside, same era.
Semantic object accuracy for generative text-to-image synthesis
T. Hinz et al · 2019
Cited alongside, same era.
Object-driven text-to-image synthesis via adversarial training
W. Li et al · 2019
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
P. Esser et al · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
J. Ho et al · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall et al · 2020
Cited alongside, same era.
Score-based generative modeling with critically-damped langevin diffusion
T. Dockhorn et al · 2021
Closest in time.
The creation and detection of deepfakes: A survey
Y. Mirsky and W. Lee · 2021
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia et al · 2022
Closest in time.
Make-a-video: Text-to-video generation without text-video data
U. Singer et al · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh et al · 2022
Closest in time.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Closest in time.
Gan inversion: A survey
W. Xia et al · 2022
Closest in time.
Prompt-to-prompt image editing with cross attention control
A. Hertz et al · 2022
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz et al · 2022
Closest in time.
An image is worth one word: Personalizing text-to-image generation using textual inversion
R. Gal et al · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
R. Rombach et al · 2022
Closest in time.
Magic3d: High-resolution text-to-3d content creation
C.-H. Lin et al · 2022
Closest in time.
Ide-3d: Interactive disentangled editing for high-resolution 3d-aware portrait synthesis
J. Sun et al · 2022
Closest in time.
Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory
S. J. Park et al · 2022
Closest in time.
Semantic image synthesis via diffusion models
W. Wang et al · 2022
Closest in time.
Diffusion-based scene graph to image generation with masked contrastive pre-training
L. Yang et al · 2022
Closest in time.
Instaformer: Instance-aware image-to-image translation with transformer
S. Kim et al · 2022
Closest in time.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu et al · 2022
Closest in time.
Make-a-scene: Scene-based text-to-image generation with human priors
O. Gafni et al · 2022
Closest in time.
Bailando: 3d dance generation by actor-critic gpt with choreographic memory
L. Siyao et al · 2022
Closest in time.
Sinnerf: Training neural radiance fields on complex scenes from a single image
D. Xu et al · 2022
Closest in time.
Clip-nerf: Text-and-image driven manipulation of neural radiance fields
C. Wang et al · 2022
Closest in time.
Learning dynamic facial radiance fields for few-shot talking head synthesis
S. Shen et al · 2022
Closest in time.
Semantic-aware implicit neural audio-driven video portrait generation
X. Liu et al · 2022
Closest in time.
Bi-level feature alignment for versatile image translation and manipulation
F. Zhan et al · 2022
Closest in time.
Semantic layout manipulation with high-resolution sparse attention
H. Zheng et al · 2022
Closest in time.
Mind reader: Reconstructing complex images from brain activities
S. Lin et al · 2022
Closest in time.
High-resolution image reconstruction with latent diffusion models from human brain activity
Y. Takagi and S. Nishimoto · 2022
Closest in time.
User-controllable latent transformer for stylegan image layout editing
Y. Endo · 2022
Closest in time.
Towards counterfactual image manipulation via clip
Y. Yu et al · 2022
Closest in time.
Imagic: Text-based real image editing with diffusion models
B. Kawar et al · 2022
Closest in time.
Cascaded diffusion models for high fidelity image generation
J. Ho et al · 2022
Closest in time.
Improved vector quantized diffusion models
Z. Tang et al · 2022
Closest in time.
Exploring transformer backbones for image diffusion models
P. Chahal · 2022
Closest in time.
Text2human: Text-driven controllable human image generation
Y. Jiang et al · 2022
Closest in time.
Compositional visual generation with composable diffusion models
N. Liu et al · 2022
Closest in time.
Retrieval-augmented diffusion models
A. Blattmann et al · 2022
Closest in time.
Sine: Single image editing with text-to-image diffusion models
Z. Zhang et al · 2022
Closest in time.
Auto-regressive image synthesis with integrated quantization
F. Zhan et al · 2022
Closest in time.
Divae: Photorealistic images synthesis with denoising diffusion decoder
J. Shi et al · 2022
Closest in time.
NÜwa-lip: Language guided image inpainting with defect-free vqgan
M. Ni et al · 2022
Closest in time.
Autoregressive image generation using residual quantization
D. Lee et al · 2022
Closest in time.
Get3d: A generative model of high quality 3d textured shapes learned from images
J. Gao et al · 2022
Closest in time.
Maskgit: Masked generative image transformer
H. Chang et al · 2022
Closest in time.
Asset: autoregressive semantic scene editing with transformers at high resolutions
D. Liu et al · 2022
Closest in time.
Neural fields in visual computing and beyond
Y. Xie et al · 2022
Closest in time.
Zero-shot text-guided object generation with dream fields
A. Jain et al · 2022
Closest in time.
Avatarclip: Zero-shot text-driven generation and animation of 3d avatars
F. Hong et al · 2022
Closest in time.
Ref-nerf: Structured view-dependent appearance for neural radiance fields
D. Verbin et al · 2022
Closest in time.
Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis
J. Gu et al · 2022
Closest in time.
Efficient geometry-aware 3d generative adversarial networks
E. R. Chan et al · 2022
Closest in time.
Sem2nerf: Converting single-view semantic masks to neural radiance fields
Y. Chen et al · 2022
Closest in time.
Fenerf: Face editing in neural radiance fields
J. Sun et al · 2022
Closest in time.
Deep image synthesis from intuitive user input: A review and perspectives
Y. Xue et al · 2022
Closest in time.
Panoptic scene graph generation
J. Yang et al · 2022
Closest in time.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann et al · 2022
Closest in time.
Text-driven stylization of video objects
S. Loeschcke et al · 2022
Closest in time.
Text2mesh: Text-driven neural stylization for meshes
O. Michel et al · 2022
Closest in time.
Text to mesh without 3d supervision using limit subdivision
N. Khalid et al · 2022
Closest in time.
3d photo stylization: Learning to generate stylized novel views from a single image
F. Mu et al · 2022
Closest in time.
3d-aware indoor scene synthesis with depth priors
Z. Shi et al · 2022
Closest in time.
Styleformer: Transformer based generative adversarial networks with style vector
J. Park and Y. Kim · 2022
Closest in time.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
M. Ding et al · 2022
Closest in time.
Expressive talking head generation with granular audio-visual control
B. Liang et al · 2022
Closest in time.
Eamm: One-shot emotional talking face via audio-based emotion-aware motion model
X. Ji et al · 2022
Closest in time.
One-shot talking face generation from single-speaker audio-visual correlation learning
S. Wang et al · 2022
Closest in time.
Point-nerf: Point-based neural radiance fields
Q. Xu et al · 2022
Closest in time.
Adding conditional control to text-to-image diffusion models
L. Zhang and M. Agrawala · 2023
Closest in time.
Layoutdiffuse: Adapting foundational diffusion models for layout-to-image generation
J. Cheng et al · 2023
Closest in time.
Dreamfusion: Text-to-3d using 2d diffusion
B. Poole et al · 2023
Closest in time.
Y. Takagi and S. Nishimoto · 2023
Closest in time.
Drag your gan: Interactive point-based manipulation on the generative image manifold
X. Pan et al · 2023
Closest in time.
3d-aware conditional image synthesis
K. Deng et al · 2023
Closest in time.
Scaling up gans for text-to-image synthesis
M. Kang et al · 2023
Closest in time.
Unicontrol: A unified diffusion model for controllable visual generation in the wild
C. Qin et al · 2023
Closest in time.
Instructpix2pix: Learning to follow image editing instructions
T. Brooks et al · 2023
Closest in time.
Difftalk: Crafting diffusion models for generalized audio-driven portraits animation
S. Shen et al · 2023
Closest in time.
Edge: Editable dance generation from music
J. Tseng et al · 2023
Closest in time.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
L. Ruan et al · 2023
Closest in time.
Muse: Text-to-image generation via masked generative transformers
H. Chang et al · 2023
Closest in time.
Or-nerf: Object removing from 3d scenes guided by multiview segmentation with neural radiance fields
Y. Yin et al · 2023
Closest in time.
Sked: Sketch-guided text-based 3d editing
A. Mikaeili et al · 2023
Closest in time.
Sine: Semantic-driven image-based nerf editing with prior-guided editing field
C. Bao et al · 2023
Closest in time.
Removing objects from neural radiance fields
S. Weder et al · 2023
Closest in time.
Local 3d editing via 3d distillation of clip knowledge
J. Hyung et al · 2023
Closest in time.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Z. Wang et al · 2023
Closest in time.
Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis
Z. Ye et al · 2023
Closest in time.
Geneface++: Generalized and stable real-time audio-driven 3d talking face generation
Z. Ye et al · 2023
Closest in time.
Multidiffusion: Fusing diffusion paths for controlled image generation
O. Bar-Tal et al · 2023
Closest in time.
Regularized vector quantization for tokenized image synthesis
J. Zhang et al · 2023
Closest in time.
Audio-driven talking face generation with diverse yet realistic facial animations
R. Wu et al · 2023
Closest in time.
X. Wu et al · 2023
Closest in time.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Y. Kirstain et al · 2023
Closest in time.
Consistency models
Y. Song et al · 2023
Closest in time.
3d generation on imagenet
I. Skorokhodov et al · 2023
Closest in time.