Fetching the paper…
Reading the bibliography…
The ultimate goal of transfer learning is to reduce labeled data requirements by exploiting a pre-existing embedding model trained for different datasets or tasks.
1905
Earlier work this paper cites.
F. Boller and J. Becker, “Dementiabank database guide,”
2005
Earlier work this paper cites.
B. Schuller, S. Steidl, A. Batliner, J. Hirschberg, J. Burgoon, A. Baird, A. Elkins, Y. Zhang, E. Coutinho, and K. Evanini, “The INTERSPEECH 2016 computational paralinguistics challenge: Deception, sincerity and native language,” in
2005
Earlier work this paper cites.
S. Haq, P. J. Jackson, and J. Edge, “Speaker-dependent audio-visual emotion recognition.” in
2009
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,”
2009
Earlier work this paper cites.
F. Eyben, M. Wöllmer, and B. Schuller, “openSMILE: The Munich versatile and fast open-source audio feature extractor,” in
2010
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,”
2011
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “CREMA-D: Crowd-sourced emotional multimodal actors dataset,”
2014
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in
2014
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in
2015
Earlier work this paper cites.
R. Arandjelović, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, “Netvlad: CNN architecture for weakly supervised place recognition,” 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,”
2017
Cited alongside, same era.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,”
2017
Cited alongside, same era.
2018
Later among the works it cites.
S. Latif, R. Rana, J. Qadir, and J. Epps, “Variational autoencoders for learning latent representations of speech emotion: a preliminary study,”
2018
Later among the works it cites.
X. Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, C. Riquelme, M. Lucic, J. Djolonga, A. S. Pinto, M. Neumann, A. Dosovitskiy, L. Beyer, O. Bachem, M. Tschannen, M. Michalski, O. Bousquet, S. Gelly, and N. Houlsby, “The visual task adaptation benchmark,” 2019
2019
Later among the works it cites.
V. Cheplygina, M. de Bruijne, and J. P. Pluim, “Not-so-supervised: a survey of semi-supervised, multi-instance, and transfer learning in medical image analysis,”
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
K. MacLean, “Voxforge,”
2018
Cited alongside, same era.
P. Warden, “Speech commands: A dataset for limited-vocabulary speech recognition,” 2018
2018
Cited alongside, same era.
C. Tan, F. Sun, T. Kong, W. Zhang, C. Yang, and C. Liu, “A survey on deep transfer learning,” in
2018
Cited alongside, same era.
S. Parthasarathy and C. Busso, “Ladder networks for emotion recognition: Using unsupervised auxiliary tasks to improve predictions of emotional attributes,”
2018
Cited alongside, same era.
2019
Later among the works it cites.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning Problem-Agnostic Speech Representations from Multiple Self-Supervised Tasks,” in
2019
Later among the works it cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An unsupervised autoregressive model for speech representation learning,”
2019
Later among the works it cites.
E. Ghaleb, M. Popa, and S. Asteriadis, “Multimodal and temporal perception of audio-visual cues for emotion recognition,” in
2019
Later among the works it cites.
Y. Tian, C. Xu, and D. Li, “Deep audio prior,” 2019
2019
Later among the works it cites.
S. Latif, R. Rana, S. Khalifa, R. Jurdak, J. Qadir, and B. W. Schuller, “Deep representation learning in speech processing: Challenges, recent advances, and future trends,” 2020
2020
Closest in time.
Kaggle, “Tensorflow speech recognition challenge,”
2020
Closest in time.
M. Plakal and D. Ellis, “Yamnet,” Jan 2020. [Online]. Available:
2020
Closest in time.