Fetching the paper…
Reading the bibliography…
The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information.
Visualizing high-dimensional data using t-629 sne
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
J. Salamon, C. Jacoby, and J. P. Bello · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
K. J. Piczak · 2015
Earlier work this paper cites.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
F. Yu, Y. Zhang, S. Song, A. Seff, and J. Xiao · 2015
Earlier work this paper cites.
A neural algorithm of artistic style
L. Gatys, A. Ecker, and M. Bethge · 2016
Earlier work this paper cites.
Image style transfer using convolutional neural networks
L. A. Gatys, A. S. Ecker, and M. Bethge · 2016
Earlier work this paper cites.
Large-scale classification of fine-art paintings: Learning the right metric on the right feature
B. Saleh and A. Elgammal · 2016
Earlier work this paper cites.
See, hear, and read: Deep aligned representations
Y. Aytar, C. Vondrick, and A. Torralba · 2017
Earlier work this paper cites.
Deep cross-modal audio-visual generation
L. Chen, S. Srivastava, Z. Duan, and C. Xu · 2017
Earlier work this paper cites.
Semantic image synthesis via adversarial learning
H. Dong, S. Yu, C. Wu, and Y. Guo · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, et al · 2017
Earlier work this paper cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
L. N. Smith · 2017
Earlier work this paper cites.
Cmcgan: A uniform framework for cross-modal visual-audio mutual generation
W. Hao, Z. Zhang, and H. Guan · 2018
Earlier work this paper cites.
Learnable pins: Cross-modal embeddings for person identity
A. Nagrani, S. Albanie, and A. Zisserman · 2018
Earlier work this paper cites.
Text-adaptive generative adversarial networks: Manipulating images with natural language
S. Nam, Y. Kim, and S. J. Kim · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Image generation associated with music data
Y. Qiu and H. Kataoka · 2018
Cited alongside, same era.
Cross-modal embeddings for video and audio retrieval
D. Surís, A. Duarte, A. Salvador, J. Torres, and X. Giró-i Nieto · 2018
Cited alongside, same era.
Image2stylegan: How to embed images into the stylegan latent space?
R. Abdal, Y. Qin, and P. Wonka · 2019
Cited alongside, same era.
Arcface: Additive angular margin loss for deep face recognition
J. Deng, J. Guo, N. Xue, and S. Zafeiriou · 2019
Crossing you in style: Cross-modal style transfer from music to visual arts
C.-C. Lee, W.-Y. Lin, Y.-T. Shih, P.-Y. Kuo, and L. Su · 2020
Later among the works it cites.
Manigan: Text-guided image manipulation
B. Li, X. Qi, T. Lukasiewicz, and P. H. Torr · 2020
Later among the works it cites.
Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis
Z. Sun, P. Sarma, W. Sethares, and Y. Liang · 2020
Later among the works it cites.
Contrastive learning of medical visual representations from paired images and text
Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P. Langlotz · 2020
Later among the works it cites.
Distilling audio-visual knowledge by compositional contrastive learning
Y. Chen, Y. Xian, A. Koepke, Y. Shan, and Z. Akata · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction
A. El-Nouby, S. Sharma, H. Schulz, D. Hjelm, L. E. Asri, S. E. Kahou, Y. Bengio, and G. W. Taylor · 2019
Cited alongside, same era.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Cited alongside, same era.
Speech2face: Learning the face behind a voice
T.-H. Oh, T. Dekel, C. Kim, I. Mosseri, W. T. Freeman, M. Rubinstein, and W. Matusik · 2019
Cited alongside, same era.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Cited alongside, same era.
Semantic image synthesis with spatially-adaptive normalization
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Closest in time.
Audioclip: Extending clip to image, text and audio, 2021
A. Guzhov, F. Raue, J. Hees, and A. Dengel · 2021
Closest in time.
Esresne (x) t-fbsp: Learning robust time-frequency transformation of audio
A. Guzhov, F. Raue, J. Hees, and A. Dengel · 2021
Closest in time.
Language-guided global image editing via cross-modal cyclic mechanism
W. Jiang, N. Xu, J. Wang, C. Gao, J. Shi, Z. Lin, and S. Liu · 2021
Closest in time.
Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label features from multi-modal embeddings
P. Mazumder, P. Singh, K. K. Parida, and V. P. Namboodiri · 2021
Closest in time.
Sound event detection: A tutorial
A. Mesaros, T. Heittola, T. Virtanen, and M. D. Plumbley · 2021
Closest in time.
Styleclip: Text-driven manipulation of stylegan imagery
O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski · 2021
Closest in time.
Encoding in style: a stylegan encoder for image-to-image translation
E. Richardson, Y. Alaluf, O. Patashnik, Y. Nitzan, Y. Azar, S. Shapiro, and D. Cohen-Or · 2021
Closest in time.
Designing an encoder for stylegan image manipulation
O. Tov, Y. Alaluf, Y. Nitzan, O. Patashnik, and D. Cohen-Or · 2021
Closest in time.
Wav2clip: Learning robust audio representations from clip, 2021
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello · 2021
Closest in time.
Tedigan: Text-guided diverse face image generation and manipulation
W. Xia, Y. Yang, J.-H. Xue, and B. Wu · 2021
Closest in time.
Deep audio-visual learning: A survey
H. Zhu, M.-D. Luo, R. Wang, A.-H. Zheng, and R. He · 2021
Closest in time.