Fetching the paper…
Reading the bibliography…
We present Compound Conditioned ControlNet, C3Net, a novel generative neural architecture taking conditions from multiple modalities and synthesizing multimodal contents simultaneously (e.g., image, text, audio).
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models, 2014
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation, 2015
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of SPIDEr
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium, 2018
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2018
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms, 2019
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Improved techniques for training score-based generative models, 2020
Yang Song and Stefano Ermon · 2020
Earlier work this paper cites.
Masked autoencoders are scalable vision learners, 2021
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Earlier work this paper cites.
Align and prompt: Video-and-language pre-training with entity prompts, 2021
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven C. H. Hoi · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs, 2021
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations, 2021
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Cited alongside, same era.
Blended diffusion for text-driven editing of natural images
Omri Avrahami, Dani Lischinski, and Ohad Fried · 2022
Cited alongside, same era.
Beats: Audio pre-training with acoustic tokenizers, 2022
Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, and Furu Wei · 2022
Cited alongside, same era.
Structure and content-guided video synthesis with diffusion models, 2023
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis · 2023
Closest in time.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models, 2023
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao · 2023
Closest in time.
Understanding and constructing latent modality structures in multi-modal representation learning, 2023
Qian Jiang, Changyou Chen, Han Zhao, Liqun Chen, Qing Ping, Son Dinh Tran, Yi Xu, Belinda Zeng, and Trishul Chilimbi · 2023
Closest in time.
Audiogen: Textually guided audio generation, 2023
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2023
Closest in time.
Masked vision and language pre-training with unimodal and multimodal contrastive losses for medical visual question answering, 2023
Pengfei Li, Gang Liu, Jinlong He, Zixu Zhao, and Shenjun Zhong · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clap: Learning audio concepts from natural language supervision, 2022
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · 2022
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning, 2022
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models, 2022
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans · 2022
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang · 2022
Cited alongside, same era.
Auto-encoding variational bayes, 2022
Diederik P Kingma and Max Welling · 2022
Cited alongside, same era.
Expanding large pre-trained unimodal models with multimodal information injection for image-text multimodal classification
Tao Liang, Guosheng Lin, Mingyang Wan, Tianrui Li, Guojun Ma, and Fengmao Lv · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Cited alongside, same era.
Audioldm: Text-to-audio generation with latent diffusion models, 2023
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D. Plumbley · 2023
Closest in time.
Mapl: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting, 2023
Oscar Mañas, Pau Rodriguez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, and Aishwarya Agrawal · 2023
Closest in time.
Automatic prompt optimization with ”gradient descent” and beam search, 2023
Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng · 2023
Closest in time.
Uniboost: Unsupervised unimodal pre-training for boosting zero-shot vision-language tasks, 2023
Yanan Sun, Zihan Zhong, Qi Fan, Chi-Keung Tang, and Yu-Wing Tai · 2023
Closest in time.
Any-to-any generation via composable diffusion, 2023
Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation, 2023
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models, 2023
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Closest in time.
Discrete contrastive diffusion for cross-modal music and image generation, 2023
Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan · 2023
Closest in time.
Unis-mmc: Multimodal classification via unimodality-supervised multimodal contrastive learning, 2023
Heqing Zou, Meng Shen, Chen Chen, Yuchen Hu, Deepu Rajan, and Eng Siong Chng · 2023
Closest in time.