“Very deep convolutional networks for large-scale image recognition,”
K. Simonyan and A. Zisserman, · 2015
Cited alongside, same era.
“Youtube-8M: A large-scale video classification benchmark,”
Original
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan, · 2016
Cited alongside, same era.
“openXBOW: Introducing the Passau open-source crossmodal bag-of-words toolkit,”
M. Schmitt and B. Schuller, · 2017
Cited alongside, same era.
“Audio set: An ontology and human-labeled dataset for audio events,”
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, · 2017
Cited alongside, same era.
“CNN architectures for large-scale audio classification,”
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, et al., · 2017
Cited alongside, same era.
“MobileNets: Efficient convolutional neural networks for mobile vision applications,”
Original
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, · 2017
Cited alongside, same era.
“Voice stress analysis: A new framework for voice and effort in human performance,”
M. Van Puyvelde, X. Neyt, F. McGlone, and N. Pattyn, · 2018
Cited alongside, same era.
“GLUE: A multi-task benchmark and analysis platform for natural language understanding,”
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, · 2018
Cited alongside, same era.
“Speech emotion recognition for performance interaction,”
N. Vryzas, R. Kotsakis, A. Liatsou, C. A. Dimoulas, and G. Kalliris, · 2018
Cited alongside, same era.
“A Canadian French emotional speech dataset,”
P. Gournay, O. Lahaie, and R. Lefebvre, · 2018
Cited alongside, same era.
“The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,”
S. R. Livingstone and F. A. Russo, · 2018
Cited alongside, same era.
“ShEMO: a large-scale validated database for persian speech emotion detection,”
O. M. Nezami, P. J. Lou, and M. Karami, · 2019
Cited alongside, same era.