Fetching the paper…
Reading the bibliography…
Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets.
“Visually indicated sounds,”
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H. Adelson, and William T. Freeman, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, L ukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“CNN architectures for large-scale audio classification,”
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al., · 2017
Earlier work this paper cites.
“Visually indicated sound generation by perceptually optimized classification,”
Kan Chen, Chuanxi Zhang, Chen Fang, Zhaowen Wang, Trung Bui, and Ram Nevatia, · 2018
Earlier work this paper cites.
“Visual to sound: Generating natural sound for videos in the wild,”
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg, · 2018
Earlier work this paper cites.
“Fréchet audio distance: A metric for evaluating music enhancement algorithms,”
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi, · 2018
Earlier work this paper cites.
“Generating visually aligned sound from videos,”
Peihao Chen, Yang Zhang, Mingkui Tan, Hongdong Xiao, Deng Huang, and Chuang Gan, · 2020
Earlier work this paper cites.
“VggSound: A large-scale audio-visual dataset,”
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman, · 2020
Earlier work this paper cites.
“Taming visually guided sound generation,”
Vladimir Iashin and Esa Rahtu, · 2021
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,”
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang, · 2022
Cited alongside, same era.
“AudioLDM: Text-to-audio generation with latent diffusion models,”
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley, · 2023
Closest in time.
“AudioGen: Textually guided audio generation,”
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi, · 2023
Closest in time.
“AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,”
Haohe Liu, Qiao Tian, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley, · 2023
Closest in time.
“Simple and controllable music generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2023
Closest in time.
“I hear your true colors: Image guided audio generation,”
Roy Sheffer and Yossi Adi, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“FlashAttention: Fast and memory-efficient exact attention with io-awareness,”
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré, · 2022
Cited alongside, same era.
“Classifier-free diffusion guidance,”
Jonathan Ho and Tim Salimans, · 2022
Cited alongside, same era.
“Efficient training of audio transformers with patchout,”
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer, · 2022
Cited alongside, same era.
“Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models,”
Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao, · 2023
Closest in time.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al., · 2023
Closest in time.
“ImageBind: One embedding space to bind them all,”
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra, · 2023
Closest in time.