Fetching the paper…
Reading the bibliography…
Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds.
“Learning a metric for music similarity,”
Malcolm Slaney, Kilian Weinberger, and William White, · 2008
Earlier work this paper cites.
“Distance metric learning for large margin nearest neighbor classification,”
Kilian Q Weinberger and Lawrence K Saul, · 2009
Earlier work this paper cites.
“Unsupervised feature learning for audio classification using convolutional deep belief networks,”
Honglak Lee, Peter Pham, Yan Largman, and Andrew Y Ng, · 2009
Earlier work this paper cites.
“Stacked convolutional auto-encoders for hierarchical feature extraction,”
Jonathan Masci, Ueli Meier, Dan Cireşan, and Jürgen Schmidhuber, · 2011
Earlier work this paper cites.
“Weak top-down constraints for unsupervised acoustic model training.,”
Aren Jansen, Samuel Thomas, and Hynek Hermansky, · 2013
Earlier work this paper cites.
“Learning fine-grained image similarity with deep ranking,”
Jiang Wang, Yang Song, Thomas Leung, Chuck Rosenberg, Jingbin Wang, James Philbin, Bo Chen, and Ying Wu, · 2014
Earlier work this paper cites.
“Weak semantic context helps phonetic learning in a model of infant language acquisition.,”
Stella Frank, Naomi Feldman, and Sharon Goldwater, · 2014
Earlier work this paper cites.
“Phonetics embedding learning with side information,”
Gabriel Synnaeve, Thomas Schatz, and Emmanuel Dupoux, · 2014
Earlier work this paper cites.
“Unsupervised learning of visual representations using videos,”
Xiaolong Wang and Abhinav Gupta, · 2015
Earlier work this paper cites.
“Deep metric learning using triplet network,”
Elad Hoffer and Nir Ailon, · 2015
Earlier work this paper cites.
“The Zero Resource Speech Challenge 2015,”
Maarten Versteegh, Roland Thiolliere, Thomas Schatz, Xuan-Nga Cao, Xavier Anguera, Aren Jansen, and Emmanuel Dupoux, · 2015
Cited alongside, same era.
“Unsupervised neural network based feature extraction using weak top-down constraints,”
Herman Kamper, Micha Elsner, Aren Jansen, and Sharon Goldwater, · 2015
Cited alongside, same era.
“Learning to see by moving,”
Pulkit Agrawal, Joao Carreira, and Jitendra Malik, · 2015
Cited alongside, same era.
“Unsupervised visual representation learning by context prediction,”
Carl Doersch, Abhinav Gupta, and Alexei A Efros, · 2015
Cited alongside, same era.
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2015
Cited alongside, same era.
“Deep convolutional neural networks and data augmentation for acoustic event detection,”
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Later among the works it cites.
“Unsupervised learning of spoken language with visual context,”
David Harwath, Antonio Torralba, and James Glass, · 2016
Later among the works it cites.
“CNN architectures for large-scale audio classification,”
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al., · 2017
Closest in time.
“Convolutional recurrent neural networks for polyphonic sound event detection,”
Emre Cakır, Giambattista Parascandolo, Toni Heittola, Heikki Huttunen, and Tuomas Virtanen, · 2017
Closest in time.
“A first attempt at polyphonic sound event detection using connectionist temporal classification,”
Yun Wang and Florian Metze, · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naoya Takahashi, Michael Gygli, Beat Pfister, and Luc Van Gool, · 2016
Cited alongside, same era.
“Colorful image colorization,”
Richard Zhang, Phillip Isola, and Alexei A Efros, · 2016
Cited alongside, same era.
“Joint learning of speaker and phonetic similarities with Siamese networks.,”
Neil Zeghidour, Gabriel Synnaeve, Nicolas Usunier, and Emmanuel Dupoux, · 2016
Cited alongside, same era.
“Deep convolutional acoustic word embeddings using word-pair side information,”
Herman Kamper, Weiran Wang, and Karen Livescu, · 2016
Cited alongside, same era.
“Context encoders: Feature learning by inpainting,”
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros, · 2016
Cited alongside, same era.
“Audio set: A strongly labeled dataset of audio events,”
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Closest in time.
“Unsupervised feature learning based on deep models for environmental audio tagging,”
Yong Xu, Qiang Huang, Wenwu Wang, Peter Foster, Siddharth Sigtia, Philip JB Jackson, and Mark D Plumbley, · 2017
Closest in time.
Relja Arandjelović and Andrew Zisserman, · 2017
Closest in time.
“Visually grounded learning of keyword prediction from untranscribed speech,”
Herman Kamper, Shane Settle, Gregory Shakhnarovich, and Karen Livescu, · 2017
Closest in time.