Fetching the paper…
Reading the bibliography…
Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings.
L. Wiskott and T. J. Sejnowski, “Slow feature analysis: Unsupervised learning of invariances,”
2002
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,”
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in
2017
Earlier work this paper cites.
A. van den Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Earlier work this paper cites.
A. Jansen, M. Plakal, R. Pandya, D. P. W. Ellis, S. Hershey, J. Liu, R. C. Moore, and R. A. Saurous, “Unsupervised learning of semantic audio representations,” in
2018
Earlier work this paper cites.
Z. Wu, Y. Xiong, S. Yu, and D. Lin, “Unsupervised Feature Learning via Non-Parametric Instance Discrimination,” in
2018
Earlier work this paper cites.
E. Fonseca, M. Plakal, D. P. W. Ellis, F. Font, X. Favory, and X. Serra, “Learning Sound Event Classifiers from Web Audio with Noisy Labels,” in
2019
Earlier work this paper cites.
E. Fonseca, F. Font, and X. Serra, “Model-agnostic Approaches to Handling Noisy Labels When Training Sound Event Classifiers,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, “Universal sound separation,” in
2019
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
2019
Earlier work this paper cites.
S. Wisdom, J. R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, and R. A. Saurous, “Differentiable consistency constraints for improved deep speech enhancement,” in
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,”
2019
Cited alongside, same era.
E. Fonseca, S. Hershey, M. Plakal, D. P. Ellis, A. Jansen, and R. C. Moore, “Addressing missing labels in large-scale sound event recognition using a teacher-student framework with loss masking,”
2020
Cited alongside, same era.
B. Zhu, K. Xu, Q. Kong, H. Wang, and Y. Peng, “Audio tagging by cross filtering noisy labels,”
2020
Cited alongside, same era.
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. van den Oord, “Learning robust and multilingual speech representations,” in
P. H. Le-Khac, G. Healy, and A. F. Smeaton, “Contrastive representation learning: A framework and review,”
2020
Later among the works it cites.
N. Turpault, S. Wisdom, H. Erdogan, J. R. Hershey, R. Serizel, E. Fonseca, P. Seetharaman, and J. Salamon, “Improving sound event detection in domestic environments using sound separation,” in
2020
Later among the works it cites.
Y. Tian, C. Sun, B. Poole, D. Krishnan, C. Schmid, and P. Isola, “What Makes for Good Views for Contrastive Learning?”
2020
Later among the works it cites.
S. Wisdom, E. Tzinis, H. Erdogan, R. J. Weiss, K. Wilson, and J. R. Hershey, “Unsupervised sound separation using mixture invariant training,” in
2020
Later among the works it cites.
A. Nandan and J. Vepa, “Language Agnostic Speech Embeddings for Emotion Classification,” in
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. de Chaumont Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv, “Towards learning a universal non-semantic representation of speech,” in
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in
2020
Cited alongside, same era.
X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with momentum contrastive learning,”
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Jansen, D. P. Ellis, S. Hershey, R. C. Moore, M. Plakal, A. C. Popat, and R. A. Saurous, “Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision,” in
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,”
2020
Later among the works it cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive learning of general-purpose audio representations,” in
2021
Closest in time.
E. Fonseca, D. Ortego, K. McGuinness, N. E. O’Connor, and X. Serra, “Unsupervised contrastive learning of sound event representations,” in
2021
Closest in time.
N. Turpault, R. Serizel, S. Wisdom, H. Erdogan, J. Hershey, E. Fonseca, P. Seetharaman, and J. Salamon, “Sound event detection and separation: a benchmark on Desed synthetic soundscapes,” in
2021
Closest in time.
S. Wisdom, H. Erdogan, D. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. Hershey, “What’s all the FUSS about free universal sound separation data?” in
2021
Closest in time.