Fetching the paper…
Reading the bibliography…
We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Least squares support vector machine classifiers
Johan AK Suykens and Joos Vandewalle · 1999
Earlier work this paper cites.
Evaluation of multiple-f0 estimation and tracking systems
Mert Bay, Andreas F Ehmann, and J Stephen Downie · 2009
Earlier work this paper cites.
Fluidsynth real-time and thread safety challenges
David Henningsson and FD Team · 2011
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Detection of piano keys pressed in video
Potcharapol Suteparuk · 2014
Earlier work this paper cites.
Automatic piano tutoring system using consumer-level depth camera
Seungmin Rho, Jae-In Hwang, and Junho Kim · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Fundamentals of music processing: Audio, analysis, algorithms, applications
Meinard Müller · 2015
Earlier work this paper cites.
Real-time piano music transcription based on computer vision
Mohammad Akbari and Howard Cheng · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Unsupervised learning of spoken language with visual context
David Harwath, Antonio Torralba, and James Glass · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Earlier work this paper cites.
An image analysis approach for transcription of music played on keyboard-like instruments
Souvik Sinha Deb and Ajit Rajwade · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Cited alongside, same era.
Piano music transcription based on computer vision
Robert McCaffrey · 2017
Cited alongside, same era.
A real-time system for online learning-based visual transcription of piano music
Mohammad Akbari, Jie Liang, and Howard Cheng · 2018
Later among the works it cites.
Clarinet: Parallel wave generation in end-to-end text-to-speech
Wei Ping, Kainan Peng, and Jitong Chen · 2018
Later among the works it cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Later among the works it cites.
Timbretron: A wavenet (cyclegan (cqt (audio))) pipeline for musical timbre transfer
Sicong Huang, Qiyang Li, Cem Anil, Xuchan Bao, Sageev Oore, and Roger B Grosse · 2018
Later among the works it cites.
Enabling factorized piano music modeling and generation with the maestro dataset
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts, Ian Simon, Colin Raffel, Jesse Engel, Sageev Oore, and Douglas Eck · 2017
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C Cobo, Florian Stimberg, et al · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Cited alongside, same era.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Cited alongside, same era.
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2018
Later among the works it cites.
Dancing to music
Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz · 2019
Later among the works it cites.
Learning individual styles of conversational gesture
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik · 2019
Later among the works it cites.
Speech2face: Learning the face behind a voice
Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T Freeman, Michael Rubinstein, and Wojciech Matusik · 2019
Later among the works it cites.
Virtual piano using computer vision
Seongjae Kang, Jaeyoon Kim, and Sung-eui Yoon · 2019
Later among the works it cites.
Observing pianist accuracy and form with computer vision
Jangwon Lee, Bardia Doosti, Yupeng Gu, David Cartledge, David Crandall, and Christopher Raphael · 2019
Later among the works it cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville · 2019
Later among the works it cites.
Performancenet: Score-to-audio music generation with multi-band convolutional residual network
Bryan Wang and Yi-Hsuan Yang · 2019
Later among the works it cites.
Score-cam: Improved visual explanations via score-weighted class activation mapping
Haofan Wang, Mengnan Du, Fan Yang, and Zijian Zhang · 2019
Later among the works it cites.
Multi-label image classification by feature attention network
Zheng Yan, Weiwei Liu, Shiping Wen, and Yin Yang · 2019
Later among the works it cites.
Sight to sound: An end-to-end approach for visual piano transcription
A Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman · 2020
Closest in time.