Fetching the paper…
Reading the bibliography…
The ability of artificial intelligence (AI) systems to perceive and comprehend audio signals is crucial for many applications.
Freesound technical demo
Frederic Font, Gerard Roma, and Xavier Serra · 2013
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
A study of instrument-wise onset detection in beijing opera percussion ensembles
Mi Tian, Ajay Srinivasamurthy, Mark Sandler, and Xavier Serra · 2014
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Tut database for acoustic scene classification and sound event detection
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Clear: A dataset for compositional language and elementary acoustic reasoning
Jerome Abdelnour, Giampiero Salvi, and Jean Rouat · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Audio set classification with attention model: A probabilistic perspective
Qiuqiang Kong, Yong Xu, Wenwu Wang, and Mark D Plumbley · 2018
Earlier work this paper cites.
Knowledge transfer from weakly labeled audio using convolutional neural network for sound events and scenes
Anurag Kumar, Maksim Khadkevich, and Christian Fügen · 2018
Earlier work this paper cites.
Reducing model complexity for dnn based large-scale audio classification
Yuzhong Wu and Tan Lee · 2018
Earlier work this paper cites.
Multi-level attention model for weakly supervised audio classification
Changsong Yu, Karim Said Barsim, Qiuqiang Kong, and Bin Yang · 2018
Earlier work this paper cites.
Look, listen, and learn more: Design choices for deep audio embeddings
Jason Cramer, Ho-Hsiang Wu, Justin Salamon, and Juan Pablo Bello · 2019
Earlier work this paper cites.
A deep residual network for large-scale acoustic scene analysis
Logan Ford, Hao Tang, François Grondin, and James R Glass · 2019
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Weakly labelled audioset tagging with attention neural networks
Qiuqiang Kong, Changsong Yu, Yong Xu, Turab Iqbal, Wenwu Wang, and Mark D Plumbley · 2019
Earlier work this paper cites.
Crowdsourcing a dataset of audio captions
Samuel Lipping, Konstantinos Drossos, and Tuoams Virtanen · 2019
Earlier work this paper cites.
A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling
Yun Wang, Juncheng Li, and Florian Metze · 2019
Earlier work this paper cites.
Zero-shot audio classification based on class label embeddings
Huang Xie and Tuomas Virtanen · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Cited alongside, same era.
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi, Daiki Takeuchi, and Masahiro Yasuda · 2020
Cited alongside, same era.
Sound event detection of weakly labelled data with cnn-transformer and automatic threshold optimization
Qiuqiang Kong, Yong Xu, Wenwu Wang, and Mark D Plumbley · 2020
Cited alongside, same era.
Wavprompt: Towards few-shot spoken language understanding with frozen language models
Heting Gao, Junrui Ni, Kaizhi Qian, Yang Zhang, Shiyu Chang, and Mark Hasegawa-Johnson · 2022
Later among the works it cites.
Vocalsound: A dataset for improving human vocal sounds recognition
Yuan Gong, Jin Yu, and James Glass · 2022
Later among the works it cites.
Zero-shot audio classification using synthesised classifiers and pre-trained models
Zheng Gu, Xinzhou Xu, Shuo Liu, and Björn Schuller · 2022
Later among the works it cites.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Later among the works it cites.
Masked autoencoders that listen
Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A crnn-gru based reinforcement learning approach to audio captioning
Xuenan Xu, Heinrich Dinkel, Mengyue Wu, and Kai Yu · 2020
Cited alongside, same era.
Fsd50k: an open dataset of human-labeled sound events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra · 2021
Cited alongside, same era.
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Cited alongside, same era.
The benefit of temporally-strong labels in audio event classification
Shawn Hershey, Daniel PW Ellis, Eduardo Fonseca, Aren Jansen, Caroline Liu, R Channing Moore, and Manoj Plakal · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al · 2021
Cited alongside, same era.
Zero-shot single-microphone sound classification and localization in a building via the synthesis of unseen features
Seungjun Lee, Haesang Yang, Hwiyong Choi, and Woojae Seong · 2021
Cited alongside, same era.
Eungbeom Kim, Jinhee Kim, Yoori Oh, Kyungsu Kim, Minju Park, Jaeheon Sim, Jinwoo Lee, and Kyogu Lee · 2022
Later among the works it cites.
Automated audio captioning with keywords guidance
Xinhao Mei, Xubo Liu, Haohe Liu, Jianyuan Sun, Mark D Plumbley, and Wenwu Wang · 2022
Later among the works it cites.
Unsupervised audio-caption aligning learns correspondences between individual sound events and textual phrases
Huang Xie, Okko Räsänen, Konstantinos Drossos, and Tuomas Virtanen · 2022
Later among the works it cites.
Speechprompt v2: Prompt tuning for speech classification tasks
Kai-Wei Chang, Yu-Kai Wang, Hua Shen, Iu-thing Kang, Wei-Cheng Tseng, Shang-Wen Li, and Hung-yi Lee · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
Pengi: An audio language model for audio tasks
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
Clap learning audio concepts from natural language supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D Plumbley, Yuexian Zou, and Wenwu Wang · 2023
Closest in time.
Dataset balancing can hurt model performance
R Channing Moore, Daniel PW Ellis, Eduardo Fonseca, Shawn Hershey, Aren Jansen, and Manoj Plakal · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.