Fetching the paper…
Reading the bibliography…
Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques.
“Robust speech recognition in noise: an evaluation using the spine corpus.,”
J. Hansen, R. Sarikaya, U. Yapanel, and B.L. Pellom, · 2001
Earlier work this paper cites.
“High performance digit recognition in real car environments.,”
U. Yapanel, X. Zhang, and J. Hansen, · 2002
Earlier work this paper cites.
“Pixels that sound,”
E. Kidron, Y. Schechner, and M. Elad, · 2005
Earlier work this paper cites.
“Classification of acoustic events using svm-based clustering schemes,”
A. Temko and C. Nadeu, · 2006
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, · 2009
Earlier work this paper cites.
“The PASCAL Visual Object Classes (VOC) challenge,”
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, · 2010
Earlier work this paper cites.
“Real-world acoustic event detection,”
X. Zhuang, X. Zhou, M. Hasegawa-Johnson, and T. Huang, · 2010
Earlier work this paper cites.
“A blind segmentation approach to acoustic event detection based on i-vector,”
Z. Huang, Y. Cheng, K. Li, V. Hautamäki, and C. Lee, · 2013
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality,”
T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean, · 2013
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
J. Salamon, C. Jacoby, and J. P. Bello, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
K. Simonyan and A. Zisserman, · 2015
Earlier work this paper cites.
“Reliable detection of audio events in highly noisy environments,”
P. Foggia, N. Petkov, A. Saggese, N. Strisciuglio, and M. Vento, · 2015
Cited alongside, same era.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Cited alongside, same era.
“TUT database for acoustic scene classification and sound event detection,”
A. Mesaros, T. Heittola, and T. Virtanen, · 2016
Cited alongside, same era.
“NetVLAD: CNN architecture for weakly supervised place recognition,”
R. Arandjelović, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, · 2016
Cited alongside, same era.
“Visually indicated sounds.,”
A. Owens, P. Isola, J.H. McDermott, A. Torralba, E.H. Adelson, and W.T. Freeman, · 2016
Cited alongside, same era.
“Deep convolutional neural networks and data augmentation for acoustic event detection,”
T. Naoya, G. Michael, P. Beat, and V. Luc, · 2016
“VoxCeleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Later among the works it cites.
“Look, listen and learn,”
R. Arandjelović and A. Zisserman, · 2017
Later among the works it cites.
“Convolutional recurrent neural networks for music classification,”
K. Choi, G. Fazekas, M. Sandler, and K. Cho, · 2017
Later among the works it cites.
“VoxCeleb2: Deep speaker recognition,”
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Later among the works it cites.
“The sound of pixels,”
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, · 2018
Later among the works it cites.
“Audio-visual event localization in unconstrained videos,”
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Recurrent neural networks for polyphonic sound event detection in real life recordings,”
G. Parascandolo, H. Huttunen, and T. Virtanen, · 2016
Cited alongside, same era.
“Openimages: A public dataset for large-scale multi-label and multi-class image classification.,”
I. Krasin, T. Duerig, N. Alldrin, A. Veit, S. Abu-El-Haija, S. Belongie, D. Cai, Z. Feng, V. Ferrari, V. Gomes, A. Gupta, D. Narayanan, C. Sun, G. Chechik, and K. Murphy, · 2016
Cited alongside, same era.
“CNN architectures for large-scale audio classification,”
S. Hershey, S. Chaudhuri, D Ellis, J Gemmeke, A. Jansen, C Moore, M Plakal, D Platt, R Saurous, B Seybold, M Slaney, R Weiss, and K Wilson, · 2017
Cited alongside, same era.
“Freesound datasets: a platform for the creation of open audio datasets,”
F. Eduardo, P. Jordi, F. Xavier, F. Frederic, B. Dmitry, F. Andrés, O. Sergio, P. Alastair, and S. Xavier, · 2017
Cited alongside, same era.
“Audio set: An ontology and human-labeled dataset for audio events,”
J Gemmeke, D Ellis, D Freedman, A Jansen, W Lawrence, C Moore, M Plakal, and M Ritter, · 2017
Cited alongside, same era.
“Large-scale weakly supervised audio classification using gated convolutional neural network,”
Y. Xu, Q. Kong, W. Wang, and M. D. Plumbley, · 2018
Later among the works it cites.
“A short note about kinetics-600.,”
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman, · 2018
Later among the works it cites.
“Utterance-level aggregation for speaker recognition in the wild,”
W. Xie, A. Nagrani, J.S. Chung, and A. Zisserman, · 2019
Later among the works it cites.
“Sound event detection in the DCASE 2017 challenge,”
A. Mesaros, A. Diment, B. Elizalde, T. Heittola, E. Vincent, B. Raj, and T. Virtanen, · 2019
Later among the works it cites.
“VoxCeleb: Large-scale speaker verification in the wild,”
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, · 2020
Closest in time.