Fetching the paper…
Reading the bibliography…
We explore self-supervised models that can be potentially deployed on mobile devices to learn general purpose audio representations.
Towards Federated Learning at Scale: System Design
Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloe Kiddon, Jakub Konečný, Stefano Mazzocchi, H. Brendan McMahan, Timon Van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander · 1902
Earlier work this paper cites.
Unsupervised feature learning for audio classification using convolutional deep belief networks
Honglak Lee, Yan Largman, Peter Pham, and Andrew Y Ng · 2009
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Unsupervised Visual Representation Learning by Context Prediction
Carl Doersch, Abhinav Gupta, and Alexei A. Efros · 2015
Earlier work this paper cites.
LibriSpeech: An ASR Corpus Based on Public Domain Adio Books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
MUSAN: A Music, Speech, and Noise Corpus
David Snyder, Guoguo Chen, and Daniel Povey · 2015
Earlier work this paper cites.
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee · 2016
Earlier work this paper cites.
Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks
Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox · 2016
Earlier work this paper cites.
Analysis of DNN approaches to speaker identification
Pavel Matejka, Ondrej Glembek, Ondrej Novotny, Oldrich Plchot, Frantisek Grezl, Lukas Burget, and Jan Honza Cernocky · 2016
Earlier work this paper cites.
Shuffle and Learn: Unsupervised Learning using Temporal Order Verification
Ishan Misra, C. Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Context Encoders: Feature Learning by Inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros · 2016
Earlier work this paper cites.
Richard Zhang, Phillip Isola, and Alexei A. Efros · 2016
Earlier work this paper cites.
Now Playing: Continuous low-power music recognition
Blaise Agüera y Arcas, Beat Gfeller, Ruiqi Guo, Kevin Kilgour, Sanjiv Kumar, James Lyon, Julian Odell, Marvin Ritter, Dominik Roblek, Matthew Sharifi, and Mihajlo Velimirović · 2017
Cited alongside, same era.
Learning and Evaluating Musical Features with Deep Autoencoders
Mason Bretan, Sageev Oore, Doug Eck, and Larry Heck · 2017
Cited alongside, same era.
Multi-task Self-Supervised Visual Learning
Carl Doersch and Andrew Zisserman · 2017
Cited alongside, same era.
Self-Supervised Video Representation Learning With Odd-One-Out Networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Cited alongside, same era.
Audio Set: An ontology and human-labeled dataset for audio events
Unsupervised Learning of Semantic Audio Representations
Aren Jansen, Manoj Plakal, Ratheet Pandya, Daniel P. W. Ellis, Shawn Hershey, Jiayang Liu, R. Channing Moore, and Rif A. Saurous · 2018
Later among the works it cites.
Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Later among the works it cites.
Detection and Classification of Acoustic Scenes and Events
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen · 2018
Later among the works it cites.
Spoken Language Identification, 2018
Tomasz Oponowicz · 2018
Later among the works it cites.
Audio-Visual Scene Analysis with Self-Supervised Multisensory Features
Andrew Owens and Alexei A Efros · 2018
Later among the works it cites.
Learning Sight from Sound: Ambient Sound Provides Supervision for Visual Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
CNN architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel P.W. Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, Malcolm Slaney, Ron J Weiss, and Kevin Wilson · 2017
Cited alongside, same era.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Unsupervised Feature Learning for Audio Analysis
Matthias Meyer, Jan Beutel, and Lothar Thiele · 2017
Cited alongside, same era.
Learning Features by Watching Objects Move
Deepak Pathak, Ross Girshick, Piotr Dollár, Trevor Darrell, and Bharath Hariharan · 2017
Cited alongside, same era.
Unsupervised Feature Learning Based on Deep Models for Environmental Audio Tagging
Yong Xu, Qiang Huang, Wenwu Wang, Peter Foster, Siddharth Sigtia, Philip J B Jackson, and Mark D Plumbley · 2017
Cited alongside, same era.
Relja Arandjelović and Andrew Zisserman · 2018
Cited alongside, same era.
Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
Yu-An Chung and James Glass · 2018
Cited alongside, same era.
Andrew Owens, Jiajun Wu, Josh H Mcdermott, William T Freeman, and Antonio Torralba · 2018
Later among the works it cites.
MobileNetV2: Inverted Residuals and Linear Bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Later among the works it cites.
Dan Stowell, | Mike Wood, | Hanna Pamuła, Yannis Stylianou, and Hervé Glotin · 2018
Later among the works it cites.
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
Pete Warden · 2018
Later among the works it cites.
Learning and Using the Arrow of Time
Donglai Wei, Jospeh Lim, Andrew Zisserman, and William T Freeman · 2018
Later among the works it cites.
Look, Listen and Learn More: Design Choices for Deep Audio Embeddings
Jason Cramer, Ho-Hsiang Wu, Justin Salamon, and Juan Pablo Bello · 2019
Closest in time.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle and Michael Carbin · 2019
Closest in time.
Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio · 2019
Closest in time.
Representation Learning with Contrastive Predictive Coding
Oriol Vinyals van den Oord, Yazhe Li · 2019
Closest in time.