Fetching the paper…
Reading the bibliography…
Recent advances in visually-induced audio generation are based on sampling short, low-fidelity, and one-class sounds.
Signal estimation from modified short-time Fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Towards good practices for very deep two-stream convnets
Limin Wang, Yuanjun Xiong, Zhe Wang, and Yu Qiao · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visually Indicated Sounds
Andrew Owens, Phillip Isola, Josh H. McDermott, Antonio Torralba, Edward H. Adelson, and William T. Freeman · 2016
Earlier work this paper cites.
Generative Adversarial Text to Image Synthesis
Scott E. Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Improved Techniques for Training GANs
Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Val Gool · 2016
Earlier work this paper cites.
Learning Dense Correspondence via 3D-Guided Cycle Consistency
Tinghui Zhou, Philipp Krähenbühl, Mathieu Aubry, Qi-Xing Huang, and Alexei A. Efros · 2016
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan, and Chenliang Xu · 2017
Earlier work this paper cites.
Audio Set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, R. Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron J. Weiss, and Kevin W. Wilson · 2017
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Image-to-Image Translation with Conditional Adversarial Networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros · 2017
Earlier work this paper cites.
SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron C. Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
Conditional Image Synthesis with Auxiliary Classifier GANs
Augustus Odena, Christopher Olah, and Jonathon Shlens · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, and Hongsheng Li · 2017
Cited alongside, same era.
Benchmark analysis of representative deep neural network architectures
Simone Bianco, Remi Cadene, Luigi Celona, and Paolo Napoletano · 2018
Cited alongside, same era.
Demystifying MMD GANs
Mikolaj Binkowski, Danica J. Sutherland, Michael Arbel, and Arthur Gretton · 2018
Cited alongside, same era.
Visually indicated sound generation by perceptually optimized classification
Kan Chen, Chuanxi Zhang, Chen Fang, Zhaowen Wang, Trung Bui, and Ram Nevatia · 2018
Cited alongside, same era.
StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation
Yunjey Choi, Min-Je Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo · 2018
Cited alongside, same era.
Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
High Fidelity Speech Synthesis with Adversarial Networks
Mikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C. Cobo, and Karen Simonyan · 2020
Later among the works it cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Later among the works it cites.
MaskGAN: Towards Diverse and Interactive Facial Image Manipulation
Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo · 2020
Later among the works it cites.
World-consistent video-to-video synthesis
Arun Mallya, Ting-Chun Wang, Karan Sapra, and Ming-Yu Liu · 2020
Later among the works it cites.
A differentiable perceptual audio metric learned from just noticeable differences
Pranay Manocha, Adam Finkelstein, Richard Zhang, Nicholas J. Bryan, Gautham J. Mysore, and Zeyu Jin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2018
Cited alongside, same era.
Stochastic adversarial video prediction
Alex X Lee, Richard Zhang, Frederik Ebert, Pieter Abbeel, Chelsea Finn, and Sergey Levine · 2018
Cited alongside, same era.
Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications
Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, and Gaurav Sharma · 2018
Cited alongside, same era.
cGANs with Projection Discriminator
Takeru Miyato and Masanori Koyama · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
MoCoGAN: Decomposing Motion and Content for Video Generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz · 2018
Cited alongside, same era.
Video-to-Video Synthesis
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Nikolai Yakovenko, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
Cited alongside, same era.
Drumgan: Synthesis of drum sounds with timbral feature conditioning using generative adversarial networks
Javier Nistal, Stefan Lattner, and Gael Richard · 2020
Later among the works it cites.
Audeo: Audio Generation for a Silent Performance Video
Kun Su, Xiulong Liu, and Eli Shlizerman · 2020
Later among the works it cites.
Spectrogram Analysis Via Self-Attention for Realizing Cross-Model Visual-Audio Generation
Huadong Tan, Guang Wu, Pengcheng Zhao, and Yanxiang Chen · 2020
Later among the works it cites.
Drum Synthesis and Rhythmic Transformation with Adversarial Autoencoders
Maciej Tomczak, Masataka Goto, and Jason Hockman · 2020
Later among the works it cites.
Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction
Yi Zhao, Haoyu Li, Cheng-I Lai, Jennifer Williams, Erica Cooper, and Junichi Yamagishi · 2020
Later among the works it cites.
Visually guided sound source separation using cascaded opponent filter network
Lingyu Zhu and Esa Rahtu · 2020
Later among the works it cites.
Spectrogram Inpainting for Interactive Generation of Instrument Sounds
Théis Bazin, Gaëtan Hadjeres, Philippe Esling, and Mikhail Malt · 2021
Closest in time.
CogView: Mastering Text-to-Image Generation via Transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al · 2021
Closest in time.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Closest in time.
Catch-A-Waveform: Learning to Generate Audio from a Single Short Example
Gal Greshler, Tamar Rott Shaham, and Tomer Michaeli · 2021
Closest in time.
Generative Speech Coding with Predictive Variance Regularization
W Bastiaan Kleijn, Andrew Storus, Michael Chinen, Tom Denton, Felicia SC Lim, Alejandro Luebs, Jan Skoglund, and Hengchin Yeh · 2021
Closest in time.
Collaborative Learning to Generate Audio-Video Jointly
Vinod K Kurmi, Vipul Bajaj, Badri N Patro, KS Venkatesh, Vinay P Namboodiri, and Preethi Jyothi · 2021
Closest in time.
Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning
Xubo Liu, Turab Iqbal, Jinzheng Zhao, Qiushi Huang, Mark D Plumbley, and Wenwu Wang · 2021
Closest in time.
Strumming to the Beat: Audio-Conditioned Contrastive Video Textures
Medhini Narasimhan, Shiry Ginosar, Andrew Owens, Alexei A Efros, and Trevor Darrell · 2021
Closest in time.
Styleclip: Text-driven Manipulation of Stylegan Imagery
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski · 2021
Closest in time.
Latent video transformer
Ruslan Rakhimov, Denis Volkhonskiy, Alexey Artemov, Denis Zorin, and Evgeny Burnaev · 2021
Closest in time.
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Closest in time.
FACEGAN: Facial Attribute Controllable rEenactment GAN
Soumya Tripathy, Juho Kannala, and Esa Rahtu · 2021
Closest in time.
VideoGPT: Video Generation using VQ-VAE and Transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas · 2021
Closest in time.
SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Closest in time.