Fetching the paper…
Reading the bibliography…
We introduce a dataset for facilitating audio-visual analysis of music performances.
S. Shimoda, M. Hayashi, and Y. Kanatsugu, “New chroma-key imagining technique with hi-vision background,” IEEE Trans. Broadcasting , vol. 35, no. 4, pp. 357–361, 1989
1989
Earlier work this paper cites.
C. Tomasi and T. Kanade, “Detection and tracking of point features,” School of Computer Science, Carnegie Mellon University, Tech. Rep. CMU-CS-91-132, Apr. 1991
1991
Earlier work this paper cites.
T. Fujishima, “Realtime chord recognition of musical sound: A system using common lisp music,” in Proc. Intl. Comput. Music Conf. (ICMC) , 1999, pp. 464–467
1999
Earlier work this paper cites.
M. Goto, H. Hashiguchi, T. Nishimura, and R. Oka, “RWC music database: Popular, classical and jazz music databases.” in Proc. Intl. Soc. for Music Info. Retrieval (ISMIR) , vol. 2, 2002, pp. 287–288
2002
Earlier work this paper cites.
D. Murphy, “Tracking a conductor’s baton,” in Proc. Danish Conf. Pattern Recognition and Image Anal. , vol. 2003, 2003, p. 05
2003
Earlier work this paper cites.
D. Radicioni, L. Anselma, and V. Lombardo, “A segmentation-based prototype to compute string instruments fingering,” in Proc. Conf. Interdisciplinary Musicology (CIM) , vol. 17, 2004, p. 97
2004
Earlier work this paper cites.
O. Gillet and G. Richard, “Automatic transcription of drum sequences using audiovisual features,” in Proc. IEEE Intl. Conf. Acoust., Speech, and Sig. Process. (ICASSP) , vol. 3, 2005, pp. iii–205
2005
Earlier work this paper cites.
A.-M. Burns and M. M. Wanderley, “Visual methods for the retrieval of guitarist fingering,” in Proc. Intl. Conf. New Interfaces for Musical Expression (NIME) , 2006, pp. 196–199
2006
Earlier work this paper cites.
D. Gorodnichy and A. Yogeswaran, “Detection and tracking of pianist hands and fingers,” in Proc. Canadian Conf. Comput. and Robot Vision , 2006, pp. 63–63
2006
Earlier work this paper cites.
O. Gillet and G. Richard, “ENST-Drums: an extensive audio-visual database for drum signals processing.” in Proc. Intl. Soc. for Music Info. Retrieval (ISMIR) , 2006, pp. 156–159
2006
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Trans. Audio, Speech, Language Process. , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
B. Zhang, J. Zhu, Y. Wang, and W. K. Leow, “Visual analysis of fingering for pedagogical violin transcription,” in Proc. ACM Intl. Conf. Multimedia , 2007, pp. 521–524
2007
Earlier work this paper cites.
C. Kerdvibulvech and H. Saito, “Vision-based guitarist fingering tracking using a bayesian classifier and particle filters,” in Advances in Image and Video Tech. Springer, 2007, pp. 625–638
2007
Earlier work this paper cites.
G. E. Poliner and D. P. Ellis, “A discriminative model for polyphonic piano transcription,” EURASIP J. on Applied Sig. Process. , vol. 2007, no. 1, pp. 154–154, 2007
2007
Earlier work this paper cites.
M. Paleari, B. Huet, A. Schutz, and D. Slock, “A multimodal approach to music transcription,” in Proc. IEEE Intl. Conf. Image Process. (ICIP) , 2008, pp. 93–96
2008
Earlier work this paper cites.
M. Vinyes, “MTG MASS database,” http://www.mtg.upf.edu/static/mass/resources
2008
Earlier work this paper cites.
J. Pätynen, V. Pulkki, and T. Lokki, “Anechoic recording system for symphony orchestra,” Acta Acust. united Ac. , vol. 94, no. 6, pp. 856–865, 2008
2008
Earlier work this paper cites.
M. Bay, A. F. Ehmann, and J. S. Downie, “Evaluation of multiple-F0 estimation and tracking systems.” in Proc. Intl. Soc. for Music Info. Retrieval (ISMIR) , 2009, pp. 315–320
2009
Earlier work this paper cites.
J. Scarr and R. Green, “Retrieval of guitarist fingering information using computer vision,” in Proc. Intl. Conf. Image and Vision Computing New Zealand (IVCNZ) , 2010, pp. 1–7
2010
Earlier work this paper cites.
V. Emiya, R. Badeau, and B. David, “Multipitch estimation of piano sounds using a new probabilistic spectral smoothness principle,” IEEE/ACM Trans. Audio, Speech, Language Process. , vol. 18, no. 6, pp. 1643–1654, 2010
2010
Earlier work this paper cites.
Z. Duan, B. Pardo, and C. Zhang, “Multiple fundamental frequency estimation by modeling spectral peaks and non-peak regions,” IEEE/ACM Trans. Audio, Speech, Language Process. , vol. 18, no. 8, pp. 2121–2133, 2010
2010
Earlier work this paper cites.
D. Sun, S. Roth, and M. J. Black, “Secrets of optical flow estimation and their principles,” in Proc. IEEE Conf. Comput. Vision and Pattern Recognition (CVPR) , June 2010, pp. 2432–2439
2010
Cited alongside, same era.
J. Abeßer, O. Lartillot, C. Dittmar, T. Eerola, and G. Schuller, “Modeling musical attributes to characterize ensemble recordings using rhythmic audio features,” in Proc. IEEE Intl. Conf. Acoustics, Speech and Sig. Process. (ICASSP) , 2011, pp. 189–192
2011
Cited alongside, same era.
Z. Duan and B. Pardo, “Soundprism: An online system for score-informed source separation of music audio,” IEEE J. of Selected Topics in Sig. Process. , vol. 5, no. 6, pp. 1205–1215, 2011
2011
Cited alongside, same era.
F. Platz and R. Kopiez, “When the eye listens: A meta-analysis of how audio-visual presentation enhances the appreciation of music performance,” Music Perception: An Interdisciplinary J. , vol. 30, no. 1, pp. 71–83, 2012
2012
Cited alongside, same era.
L. Su and Y.-H. Yang, “Escaping from the abyss of manual annotation: New methodology of building polyphonic datasets for automatic music transcription,” in Proc. Intl. Symp. Comput. Music Multidisciplinary Res. , 2015, pp. 309–321
2015
Later among the works it cites.
T.-S. Chan, T.-C. Yeh, Z.-C. Fan, H.-W. Chen, L. Su, Y.-H. Yang, and R. Jang, “Vocal activity informed singing voice separation with the iKala dataset,” in Proc. IEEE Intl. Conf. Acoust., Speech and Sig. Process (ICASSP) , 2015, pp. 718–722
2015
Later among the works it cites.
A. Perez-Carrillo, J.-L. Arcos, and M. Wanderley, “Estimation of guitar fingering and plucking controls based on multimodal analysis of motion, audio and musical score.” in Proc. Intl. Symp. Comput. Music Multidisciplinary Res. (CMMR) , 2015, pp. 71–87
2015
Later among the works it cites.
“Audacity,” http://www.audacityteam.org
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Caramiaux, M. M. Wanderley, and F. Bevilacqua, “Segmenting and parsing instrumentalists’ gestures,” J. of New Music Res. , vol. 41, no. 1, pp. 13–29, 2012
2012
Cited alongside, same era.
E. Benetos, A. Klapuri, and S. Dixon, “Score-informed transcription for automatic piano tutoring,” in Proc. IEEE European Sig. Process. Conf. (EUSIPCO) , 2012, pp. 2153–2157
2012
Cited alongside, same era.
S. Hargreaves, A. Klapuri, and M. Sandler, “Structural segmentation of multitrack audio,” IEEE/ACM Trans. Audio, Speech, Language Process. , vol. 20, no. 10, pp. 2637–2647, 2012
2012
Cited alongside, same era.
C.-J. Tsay, “Sight over sound in the judgment of music performance,” Nat. Academy of Sci. , vol. 110, no. 36, pp. 14 580–14 585, 2013
2013
Cited alongside, same era.
A. Oka and M. Hashimoto, “Marker-less piano fingering recognition using sequential depth images,” in Proc. Korea-Japan Joint Workshop on Frontiers of Comput. Vision (FCV) , 2013, pp. 1–4
2013
Cited alongside, same era.
B. Yao, J. Ma, and L. Fei-Fei, “Discovering object functionality,” in Proc. IEEE Intl. Conf. Comput. Vision (ICCV) , 2013, pp. 2512–2519
2013
Cited alongside, same era.
J. Fritsch and M. D. Plumbley, “Score informed audio source separation using constrained nonnegative matrix factorization and score synthesis,” in Proc. IEEE Intl. Conf. Acoust., Speech and Sig. Process. (ICASSP) , 2013, pp. 888–891
2013
Cited alongside, same era.
A. Bazzica, C. C. Liem, and A. Hanjalic, “Exploiting instrument-wise playing/non-playing labels for score synchronization of symphonic music.” in Proc. Intl. Soc. for Music Info. Retrieval (ISMIR) , 2014, pp. 201–206
2014
Cited alongside, same era.
M. Mauch, C. Cannam, R. Bittner, G. Fazekas, J. Salamon, J. Dai, J. Bello, and S. Dixon, “Computer-aided melody note transcription using the tony software: Accuracy and efficiency,” in Proc. Intl. Conf. Tech. for Music Notation and Representation , 2015
2015
Later among the works it cites.
M. Miron, J. J. Carabias-Orti, J. J. Bosch, E. Gómez, and J. Janer, “Score-informed source separation for multichannel orchestral recordings,” J. of Elect. and Comput. Eng. , vol. 2016, 2016
2016
Closest in time.
2016
Closest in time.
“Final cut pro,” http://www.apple.com/final-cut-pro/
2016
Closest in time.
A. Bazzica, C. C. Liem, and A. Hanjalic, “On detecting the playing/non-playing activity of musicians in symphonic music videos,” Comput. Vision and Image Understanding , vol. 144, pp. 188–204, 2016
2016
Closest in time.
C. Smith, http://expandedramblings.com/index.php/youtube-statistics/4/
2017
Closest in time.
S. Parekh, S. Essid, A. Ozerov, N. Duong, P. Perez, and G. Richard, “Motion informed audio source separation,” in Proc. IEEE Intl. Conf. Acoust., Speech and Sig. Process. (ICASSP) , 2017, pp. 6–10
2017
Closest in time.
2017
Closest in time.
B. Li, K. Dinesh, G. Sharma, and Z. Duan, “Video-based vibrato detection and analysis for polyphonic string music,” in Proc. Intl. Soc. for Music Inform. Retrieval (ISMIR) , 2017
2017
Closest in time.
K. Dinesh, B. Li, X. Liu, Z. Duan, and G. Sharma, “Visually informed multi-pitch analysis of string ensembles,” in Proc. IEEE Intl. Conf. Acoust., Speech, and Sig. Process. (ICASSP) , 2017, pp. 3021–3025
2017
Closest in time.
B. Li, K. Dinesh, Z. Duan, and G. Sharma, “See and listen: Score-informed association of sound tracks to players in chamber music performance videos,” in Proc. IEEE Intl. Conf. Acoust., Speech, and Sig. Process. (ICASSP) , 2017, pp. 2906–2910
2017
Closest in time.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in Proc. IEEE Intl. Conf. Acoust., Speech and Sig. Process (ICASSP) , 2017, pp. 776–780
2017
Closest in time.
B. Li, C. Xu, and Z. Duan, “Audiovisual source association for string ensembles through multi-modal vibrato analysis,” in Proc. Sound and Music Computing (SMC) , 2017
2017
Closest in time.
L. Chen, S. Srivastava, Z. Duan, and C. Xu, “Deep cross-modal audio-visual generation,” in Proc. ACM Thematic Workshops of Multimedia , 2017, pp. 349–357
2017
Closest in time.
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, (2018) Data from: “Creating A Multi-track Classical Music Performance Dataset for Multi-modal Music Analysis: Challenges, Insights, and Applications.” Dryad Digital Repository. http://doi.org/10.5061/dryad.ng3r749
2018
Closest in time.