Fetching the paper…
Reading the bibliography…
We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes.
Determining optical flow
Horn, B. K.; and Schunck, B. G. 1981 · 1981
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
Maximum Filter Vibrato Suppression for Onset Detection
Böck, S.; and Widmer, G. 2013 · 2013
Earlier work this paper cites.
Learning Spatiotemporal Features with 3D Convolutional Networks
Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Chen, L.; Srivastava, S.; Duan, Z.; and Xu, C. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
To reverse the gradient or not: An empirical comparison of adversarial and multi-task learning in speech recognition
Adi, Y.; Zeghidour, N.; Collobert, R.; Usunier, N.; Liptchinsky, V.; and Synnaeve, G. 2019 · 2019
Earlier work this paper cites.
Towards Accurate Generative Models of Video: A New Metric Challenges
Unterthiner, T.; van Steenkiste, S.; Kurach, K.; Marinier, R.; Michalski, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Towards audio to scene image synthesis using generative adversarial network
Wan, C.-H.; Chuang, S.-P.; and Lee, H.-Y. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2020
Earlier work this paper cites.
Sound2sight: Generating visual dynamics from sound and context
Chatterjee, M.; and Cherian, A. 2020 · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Chen, H.; Xie, W.; Vedaldi, A.; and Zisserman, A. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Robust one shot audio to video generation
Kumar, N.; Goel, S.; Narang, A.; and Hasan, M. 2020 · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Earlier work this paper cites.
NWT: towards natural audio-to-video generation with representation learning
Mama, R.; Tyndel, M. S.; Kadhim, H.; Clifford, C.; and Thurairatnam, R. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Godiva: Generating open-domain videos from natural descriptions
Wu, C.; Huang, L.; Zhang, Q.; Li, B.; Ji, L.; Yang, F.; Sapiro, G.; and Duan, N. 2021 · 2021
Cited alongside, same era.
Video and text matching with conditioned embeddings
Ali, A.; Schwartz, I.; Hazan, T.; and Wolf, L. 2022 · 2022
Cited alongside, same era.
Beats: Audio pre-training with acoustic tokenizers
Chen, S.; Wu, Y.; Wang, C.; Liu, S.; Tompkins, D.; Chen, Z.; and Wei, F. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.; Ghasemipour, S. K. S.; Gontijo-Lopes, R.; Ayan, B. K.; Salimans, T.; et al. 2022 · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; et al. 2022 · 2022
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual description
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vqgan-clip: Open domain image generation and editing with natural language guidance
Crowson, K.; Biderman, S.; Kornis, D.; Stander, D.; Hallahan, E.; Castricato, L.; and Raff, E. 2022 · 2022
Cited alongside, same era.
Make-a-scene: Scene-based text-to-image generation with human priors
Gafni, O.; Polyak, A.; Ashual, O.; Sheynin, S.; Parikh, D.; and Taigman, Y. 2022 · 2022
Cited alongside, same era.
Latent space explanation by intervention
Gat, I.; Lorberbom, G.; Schwartz, I.; and Hazan, T. 2022 · 2022
Cited alongside, same era.
Long video generation with time-agnostic vqgan and time-sensitive transformer
Ge, S.; Hayes, T.; Yang, H.; Yin, X.; Pang, G.; Jacobs, D.; Huang, J.-B.; and Parikh, D. 2022 · 2022
Cited alongside, same era.
Vag: A uniform model for cross-modal visual-audio mutual generation
Hao, W.; Guan, H.; and Zhang, Z. 2022 · 2022
Cited alongside, same era.
Latent video diffusion models for high-fidelity video generation with arbitrary lengths
He, Y.; Yang, T.; Zhang, Y.; Shan, Y.; and Chen, Q. 2022 · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D. P.; Poole, B.; Norouzi, M.; Fleet, D. J.; et al. 2022 · 2022
Cited alongside, same era.
Villegas, R.; Babaeizadeh, M.; Kindermans, P.-J.; Moraldo, H.; Zhang, H.; Saffar, M. T.; Castro, S.; Kunze, J.; and Erhan, D. 2022 · 2022
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip
Wu, H.-H.; Seetharaman, P.; Kumar, K.; and Bello, J. P. 2022b · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J.; Xu, Y.; Koh, J. Y.; Luong, T.; Baid, G.; Wang, Z.; Vasudevan, V.; Ku, A.; Yang, Y.; Ayan, B. K.; et al. 2022 · 2022
Later among the works it cites.
Audio-to-image cross-modal generation
Żelaszczyk, M.; and Mańdziuk, J. 2022 · 2022
Later among the works it cites.
Simple and Controllable Music Generation
Copet, J.; Kreuk, F.; Gat, I.; Remez, T.; Kant, D.; Synnaeve, G.; Adi, Y.; and Défossez, A. 2023 · 2023
Closest in time.
Textually Pretrained Speech Language Models
Hassid, M.; Remez, T.; Nguyen, T. A.; Gat, I.; Conneau, A.; Kreuk, F.; Copet, J.; Defossez, A.; Synnaeve, G.; Dupoux, E.; et al. 2023 · 2023
Closest in time.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
Ruan, L.; Ma, Y.; Yang, H.; He, H.; Liu, B.; Fu, J.; Yuan, N. J.; Jin, Q.; and Guo, B. 2023 · 2023
Closest in time.
I hear your true colors: Image guided audio generation
Sheffer, R.; and Adi, Y. 2023 · 2023
Closest in time.
AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
Yariv, G.; Gat, I.; Wolf, L.; Adi, Y.; and Schwartz, I. 2023 · 2023
Closest in time.
Magvit: Masked generative video transformer
Yu, L.; Cheng, Y.; Sohn, K.; Lezama, J.; Zhang, H.; Chang, H.; Hauptmann, A. G.; Yang, M.-H.; Hao, Y.; Essa, I.; et al. 2023 · 2023
Closest in time.
Factor graph attention
Schwartz, I.; Yu, S.; Hazan, T.; and Schwing, A. G. 2019 · 2048
Closest in time.
Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory
Park, S. J.; Kim, M.; Hong, J.; Choi, J.; and Ro, Y. M. 2022 · 2070
Closest in time.