Fetching the paper…
Reading the bibliography…
We propose Im2Wav, an image guided open-domain audio generation system.
“The influence of bilingualism on third language acquisition: Focus on multilingualism,”
Jasone Cenoz, · 2013
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord et al., · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani et al., · 2017
Earlier work this paper cites.
“Gans trained by a two time-scale update rule converge to a local nash equilibrium,”
Martin Heusel et al., · 2017
Earlier work this paper cites.
“Cnn architectures for large-scale audio classification,”
Shawn Hershey et al., · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke et al., · 2017
Earlier work this paper cites.
“Visually indicated sound generation by perceptually optimized classification,”
Kan Chen et al., · 2018
Earlier work this paper cites.
“Visual to sound: Generating natural sound for videos in the wild,”
Yipin Zhou et al., · 2018
Earlier work this paper cites.
“A style-based generator architecture for generative adversarial networks,”
Tero Karras, Samuli Laine, and Timo Aila, · 2019
Earlier work this paper cites.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar et al., · 2019
Cited alongside, same era.
“Generating diverse high-fidelity images with vq-vae-2,”
Ali Razavi, Aaron van den Oord, and Oriol Vinyals, · 2019
Cited alongside, same era.
“Generating long sequences with sparse transformers,”
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever, · 2019
Cited alongside, same era.
“Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.,”
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi, · 2019
Cited alongside, same era.
“Vggsound: A large-scale audio-visual dataset,”
Honglie Chen et al., · 2020
Cited alongside, same era.
“Generating visually aligned sound from videos,”
“Taming visually guided sound generation,”
Vladimir Iashin and Esa Rahtu, · 2021
Later among the works it cites.
“Classifier-free diffusion guidance,”
Jonathan Ho and Tim Salimans, · 2021
Later among the works it cites.
“Clipscore: A reference-free evaluation metric for image captioning,”
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi, · 2021
Later among the works it cites.
“Efficient training of audio transformers with patchout,”
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer, · 2021
Later among the works it cites.
“Hierarchical text-conditional image generation with clip latents,”
Aditya Ramesh et al., · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peihao Chen et al., · 2020
Cited alongside, same era.
“Jukebox: A generative model for music,”
Prafulla Dhariwal et al., · 2020
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford et al., · 2021
Cited alongside, same era.
“Glide: Towards photorealistic image generation and editing with text-guided diffusion models,”
Alex Nichol et al., · 2021
Cited alongside, same era.
Oran Gafni et al., · 2022
Closest in time.
“Laion-5b: laion-5b: A new era of open large-scale multi-modal datasets,” 2022
Christoph Schuhmann et al., · 2022
Closest in time.
“Audiogen: Textually guided audio generation,”
Felix Kreuk et al., · 2022
Closest in time.
“Wav2clip: Learning robust audio representations from clip,”
Ho-Hsiang Wu et al., · 2022
Closest in time.