Fetching the paper…
Reading the bibliography…
We propose a novel system that takes as an input body movements of a musician playing a musical instrument and generates music in an unsupervised setting.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Unsupervised learning of spoken language with visual context
David Harwath, Antonio Torralba, and James Glass · 2016
Earlier work this paper cites.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan, and Chenliang Xu · 2017
Earlier work this paper cites.
Latent constraints: Learning to generate conditionally from unconditional generative models
Jesse Engel, Matthew Hoffman, and Adam Roberts · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Openpose: realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2018
Earlier work this paper cites.
The challenge of realistic music generation: modelling raw audio at scale
Sander Dieleman, Aaron van den Oord, and Karen Simonyan · 2018
Earlier work this paper cites.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Earlier work this paper cites.
Enabling factorized piano music modeling and generation with the maestro dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2018
Cited alongside, same era.
Music transformer: Generating music with long-term structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck · 2018
Cited alongside, same era.
Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications
Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, and Gaurav Sharma · 2018
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al · 2018
Cited alongside, same era.
Clarinet: Parallel wave generation in end-to-end text-to-speech
Learning individual styles of conversational gesture
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik · 2019
Later among the works it cites.
You said that?: Synthesising talking faces from audio
Amir Jamaludin, Joon Son Chung, and Andrew Zisserman · 2019
Later among the works it cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville · 2019
Later among the works it cites.
Speech2face: Learning the face behind a voice
Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T Freeman, Michael Rubinstein, and Wojciech Matusik · 2019
Later among the works it cites.
Melnet: A generative model for audio in the frequency domain
Sean Vasquez and Mike Lewis · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Ping, Kainan Peng, and Jitong Chen · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
On gans and gmms
Eitan Richardson and Yair Weiss · 2018
Cited alongside, same era.
Audio to body dynamics
Eli Shlizerman, Lucio Dery, Hayden Schoen, and Ira Kemelmacher-Shlizerman · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
Spatial temporal graph convolutional networks for skeleton-based action recognition
Sijie Yan, Yuanjun Xiong, and Dahua Lin · 2018
Cited alongside, same era.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Cited alongside, same era.
Unsupervised representation learning with long-term dynamics for skeleton based action recognition
Nenggan Zheng, Jun Wen, Risheng Liu, Liangqu Long, Jianhua Dai, and Zhefeng Gong · 2018
Cited alongside, same era.
Performancenet: Score-to-audio music generation with multi-band convolutional residual network
Bryan Wang and Yi-Hsuan Yang · 2019
Later among the works it cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Later among the works it cites.
Generating visually aligned sound from videos
Peihao Chen, Yang Zhang, Mingkui Tan, Hongdong Xiao, Deng Huang, and Chuang Gan · 2020
Closest in time.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Closest in time.
Foley music: Learning to generate music from videos
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba · 2020
Closest in time.
Music gesture for visual sound separation
Chuang Gan, Deng Huang, Hang Zhao, Joshua B Tenenbaum, and Antonio Torralba · 2020
Closest in time.
Sight to sound: An end-to-end approach for visual piano transcription
A Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman · 2020
Closest in time.
Solos: A dataset for audio-visual music analysis
Juan F Montesinos, Olga Slizovskaia, and Gloria Haro · 2020
Closest in time.
Audeo: Audio generation for a silent performance video
Kun Su, Xiulong Liu, and Eli Shlizerman · 2020
Closest in time.
Predict & cluster: Unsupervised skeleton based action recognition
Kun Su, Xiulong Liu, and Eli Shlizerman · 2020
Closest in time.