Fetching the paper…
Reading the bibliography…
After its sweeping success in vision and language tasks, pure attention-based neural architectures (e.g.
J. Li, W. Dai, F. Metze, S. Qu, and S. Das, “A comparison of deep learning methods for environmental sound detection,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 126–130
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , ser. NIPS’17. USA: Curran Associates Inc., 2017, pp. 6000–6010. [Online]. Available: http://dl.acm.org/citation.cfm?id=3295222.3295349
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Wang, “Polyphonic sound event detection with weak labeling,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le, “Don’t decay the learning rate, increase the batch size,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
L. Ford, H. Tang, F. Grondin, and J. Glass, “A Deep Residual Network for Large-Scale Acoustic Scene Analysis,” in Proc. Interspeech 2019 , 2019, pp. 2568–2572. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-2731
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
Y. Wang, J. Li, and F. Metze, “A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 31–35
2019
Cited alongside, same era.
2019
Cited alongside, same era.
L. N. Smith and N. Topin, “Super-convergence: Very fast training of neural networks using large learning rates,” in Artificial intelligence and machine learning for multi-domain operations applications , vol. 11006. International Society for Optics and Photonics, 2019, p. 1100612
Y. Gong, Y.-A. Chung, and J. Glass, “Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3292–3306, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Koutini, J. Schlüter, H. Eghbal-zadeh, and G. Widmer, “Efficient training of audio transformers with patchout,” 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
K. He, R. Girshick, and P. Dollár, “Rethinking imagenet pre-training,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 4918–4927
2019
Cited alongside, same era.
2020
Cited alongside, same era.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 347–10 357
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. Hershey, D. P. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal, “The benefit of temporally-strong labels in audio event classification,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 366–370
2021
Later among the works it cites.
2021
Later among the works it cites.
J. B. Li, K. Ma, S. Qu, P.-Y. Huang, and F. Metze, “Audio-visual event recognition through the lens of adversary,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 616–620
2021
Later among the works it cites.
X. Chen, C.-J. Hsieh, and B. Gong, “When vision transformers outperform resnets without pre-training or strong data augmentations,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=LtKcMgGOeLt
2022
Closest in time.