Fetching the paper…
Reading the bibliography…
The auditory system plays a substantial role in shaping the overall human perceptual experience.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proceedings of the 23rd ACM International Conference on Multimedia . Brisbane Australia: ACM, Oct. 2015, pp. 1015–1018
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , Y. Bengio and Y. LeCun, Eds., 2015
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic propositional image caption evaluation,” in Computer Vision – ECCV 2016 . Cham: Springer International Publishing, 2016, pp. 382–398
2016
Earlier work this paper cites.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. P. Murphy, “Improved image captioning via policy gradient optimization of SPIDEr,” 2017 IEEE International Conference on Computer Vision (ICCV) , pp. 873–881, 2016
2016
Earlier work this paper cites.
O. Vinyals, C. Blundell, T. Lillicrap, K. Kavukcuoglu, and D. Wierstra, “Matching networks for one shot learning,” in Advances in Neural Information Processing Systems , vol. 29, 2016
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . New Orleans, LA: IEEE, Mar. 2017, pp. 776–780
2017
Earlier work this paper cites.
J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” Advances in Neural Information Processing Systems , vol. 30, pp. 4077–4087, 2017
2017
Earlier work this paper cites.
A. Suhr, S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi, “A corpus for reasoning about natural language grounded in photographs,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , A. Korhonen, D. Traum, and L. Màrquez, Eds. Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 6418–6428
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 119–132
2019
Earlier work this paper cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 736–740
2020
Earlier work this paper cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Earlier work this paper cites.
S. Hershey, D. P. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal, “The benefit of temporally-strong labels in audio event classification,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 366–370
2021
Earlier work this paper cites.
X. Xu, H. Dinkel, M. Wu, Z. Xie, and K. Yu, “Investigating local and global information for automated audio captioning with transfer learning,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Jun. 2021, pp. 905–909
2021
Earlier work this paper cites.
P. Yang, X. Wang, X. Duan, H. Chen, R. Hou, C. Jin, and W. Zhu, “AVQA: A dataset for audio-visual question answering on videos,” in Proceedings of the 30th ACM International Conference on Multimedia , ser. MM ’22. New York, NY, USA: Association for Computing Machinery, 2022, pp. 3480–3491, event-place: Lisboa, Portugal
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan, “Flamingo: A visual language model for few-shot learning,” in Advances in Neural Information Processing Systems , 2022
2022
Earlier work this paper cites.
X. Xu, Z. Xie, M. Wu, and K. Yu, “The SJTU system for DCASE2022 challenge task 6: Audio captioning with audio-text retrieval pre-training,” DCASE2022 Challenge, Tech. Rep., 2022
2022
Earlier work this paper cites.
P.-Y. Huang, H. Xu, J. B. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked autoencoders that listen,” in Advances in Neural Information Processing Systems , A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022
2022
Earlier work this paper cites.
S. Lipping, P. Sudarsanam, K. Drossos, and T. Virtanen, “Clotho-AQA: A crowdsourced dataset for audio question answering,” in 30th European Signal Processing Conference (EUSIPCO) , 2022, pp. 1140–1144
2022
Earlier work this paper cites.
A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, and T. Virtanen, “STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events,” in Proceedings of the 7th Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022) , Nancy, France, November 2022
2022
Earlier work this paper cites.
A. Guzhov, F. Raue, J. Hees, and A. Dengel, “Audioclip: Extending clip to image, text and audio,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2022, pp. 976–980, iSSN: 2379-190X
2022
Cited alongside, same era.
J. Liang, H. Phan, and E. Benetos, “Leveraging label hierachies for few-shot everyday sound recognition,” in Proceedings of the 7th Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022) , Nancy, France, Nov. 2022
2022
Cited alongside, same era.
R. Zhang, W. Zhang, R. Fang, P. Gao, K. Li, J. Dai, Y. Qiao, and H. Li, “Tip-Adapter: Training-free adaption of CLIP for few-shot classification,” in Computer Vision – ECCV 2022 , ser. Lecture Notes in Computer Science, S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner, Eds. Cham: Springer Nature Switzerland, 2022, pp. 493–510
2022
Cited alongside, same era.
J. Von Oswald, E. Niklasson, E. Randazzo, J. a. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov, “Transformers learn in-context by gradient descent,” in Proceedings of the 40th International Conference on Machine Learning (ICML’23) . JMLR.org, 2023
2023
Closest in time.
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang, “CLAP: Learning audio concepts from natural language supervision,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.
M. Kim, K. Sung-Bin, and T.-H. Oh, “Prefix tuning for automated audio captioning,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Jun. 2023, pp. 1–5
2023
Closest in time.
T. Kouzelis and V. Katsouros, “Weakly-supervised automated audio captioning via text only training,” in Proceedings of the 8th Detection and Classification of Acoustic Scenes and Events 2023 Workshop (DCASE2023) , Tampere, Finland, September 2023, pp. 81–85
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
W. Dai, J. Li, D. Li, A. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “InstructBLIP: Towards general-purpose vision-language models with instruction tuning,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
M. Shukor, C. Dancette, A. Rame, and M. Cord, “UnIVAL: Unified model for image, video, audio and language tasks,” Transactions on Machine Learning Research , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proceedings of the 40th International Conference on Machine Learning , ser. ICML’23. JMLR.org, 2023
2023
Cited alongside, same era.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio pre-training with acoustic tokenizers,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 23–29 Jul 2023, pp. 5178–5193
2023
Cited alongside, same era.
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
2023
Closest in time.
D. Dai, Y. Sun, L. Dong, Y. Hao, S. Ma, Z. Sui, and F. Wei, “Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers,” in Findings of the Association for Computational Linguistics: ACL 2023 , A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 4005–4019
2023
Closest in time.
2024
Closest in time.
S. Hu, L. Zhou, S. Liu, S. Chen, L. Meng, H. Hao, J. Pan, X. Liu, J. Li, S. Sivasankaran, L. Liu, and F. Wei, “WavLLM: Towards robust and adaptive speech large language model,” in Findings of the Association for Computational Linguistics: EMNLP 2024 , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 4552–4572
2024
Closest in time.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
Z. Kong, A. Goel, R. Badlani, W. Ping, R. Valle, and B. Catanzaro, “Audio Flamingo: A novel audio language model with few-shot learning and dialogue abilities,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
R. Zhang, J. Han, C. Liu, A. Zhou, P. Lu, Y. Qiao, H. Li, and P. Gao, “LLaMA-Adapter: Efficient fine-tuning of large language models with zero-initialized attention,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
J. Liang, H. Zhang, H. Liu, Y. Cao, Q. Kong, X. Liu, W. Wang, M. D. Plumbley, H. Phan, and E. Benetos, “WavCraft: Audio editing and generation with large language models,” in ICLR 2024 Workshop on Large Language Model (LLM) Agents , 2024
2024
Closest in time.
B. Ding, T. Zhang, C. Wang, G. Liu, J. Liang, R. Hu, Y. Wu, and D. Guo, “Acoustic scene classification: A comprehensive survey,” Expert Syst. Appl. , vol. 238, Feb. 2024
2024
Closest in time.
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” IEEE/ACM Trans. Audio, Speech and Lang. Proc. , vol. 32, p. 3339–3354, Jun. 2024
2024
Closest in time.
J. Liang, H. Phan, and E. Benetos, “Learning from taxonomy: Multi-label few-shot classification for everyday sound recognition,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 771–775
2024
Closest in time.
J. Liang, I. Nolasco, B. Ghani, H. Phan, E. Benetos, and D. Stowell, “Mind the domain gap: A systematic analysis on bioacoustic sound event detection,” in 2024 32nd European Signal Processing Conference (EUSIPCO) , 2024, pp. 1257–1261
2024
Closest in time.
2024
Closest in time.
B. Jin, J. Yoon, J. Han, and S. O. Arik, “Long-context LLMs meet RAG: Overcoming challenges for long inputs in RAG,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.
H. Zhang, V. Cheung, H. Nishioka, S. Dixon, and S. Furuya, “LLaQo: Towards a query-based coach in expressive performance assessment,” Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2025 , 2025
2025
Closest in time.