Fetching the paper…
Reading the bibliography…
In music and speech, meaning is derived at multiple levels of context.
“Communication of emotions in vocal expression and music performance: Different channels, same code?,”
Patrik Juslin and Petri Laukka, · 2003
Earlier work this paper cites.
“Statistical characterisation of melodic pitch contours and its application for melody extraction,”
Justin Salamon, Geoffroy Peeters, and Axel Roebel, · 2012
Earlier work this paper cites.
“Melody extraction from polyphonic music signals using pitch contour characteristics,”
Justin Salamon and Emilia Gomez, · 2012
Earlier work this paper cites.
“Efficient estimation of word representations in vector space,”
Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean, · 2013
Earlier work this paper cites.
“Chord2vec: Learning musical chord embeddings,”
Sephora Madjiheurem, Lizhen Qu, and Christian Walder, · 2016
Earlier work this paper cites.
“Frequency estimation from waveforms using multi-layered neural networks,”
Prateek Verma and Ronald W Schafer, · 2016
Earlier work this paper cites.
“Pitch contours as a mid-level representation for music informatics,”
Rachel Bittner, Justin Salamon, Juan Bosch, and Juan Bello, · 2017
Earlier work this paper cites.
“Towards the characterization of singing styles in world music,”
M. Panteli, R. Bittner, J. P. Bello, and S. Dixon, · 2017
Cited alongside, same era.
“The expression of emotion in the singing voice: Acoustic patterns in vocal performance,”
Klaus R. Scherer, Johan Sundberg, Bernardino Fantini, Stéphanie Trznadel, and Florian Eyben, · 2017
Cited alongside, same era.
“Modeling musical context with word2vec,”
Dorien Herremans and Ching-Hua Chuan, · 2017
Cited alongside, same era.
“Cnn architectures for large-scale audio classification,”
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, · 2017
Cited alongside, same era.
“Crepe: A convolutional representation for pitch estimation,”
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello, · 2018
Cited alongside, same era.
“Vocalset: A singing voice dataset,”
Julia Wilkins, Prem Seetharaman, Alison Wahl, and Bryan Pardo, · 2018
Later among the works it cites.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R. Livingstone and Frank A. Russo, · 2018
Later among the works it cites.
“Fundamental frequency contour classification: A comparison between hand-crafted and cnn-based features,”
J. Abeßer and M. Müller, · 2019
Later among the works it cites.
“Intonation trajectories within tones in unaccompanied soprano, alto, tenor, bass quartet singing.,”
Jiajie Dai and Simon Dixon, · 2019
Later among the works it cites.
“A simple framework for contrastive learning of visual representations,”
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu-An Chung and James R. Glass, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Disentangled speech embeddings using cross-modal self-supervision,”
Arsha Nagrani, Joon Son Chung, Samuel Albanie, and Andrew Zisserman, · 2020
Closest in time.
“Multi-task self-supervised visual learning,”
Carl Doersch and Andrew Zisserman, · 2060
Closest in time.