Fetching the paper…
Reading the bibliography…
Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data.
Freesound technical demo. In Proc. ACM MM . 411–412
Frederic Font, Gerard Roma, and Xavier Serra. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context. In Proc. ECCV . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In Proc. ICCV . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Polyphonic sound event detection using multi label deep neural networks. In Proc. IJCNN . 1–7
Emre Cakir, Toni Heittola, Heikki Huttunen, and Tuomas Virtanen. 2015 · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. 2016 · 2016
Earlier work this paper cites.
Automated audio captioning with recurrent neural networks. In Proc. IEEE WASPAA . 374–378
Konstantinos Drossos, Sharath Adavanne, and Tuomas Virtanen. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events. In Proc. IEEE ICASSP . 776–780
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Van den Oord et al · 2018
Earlier work this paper cites.
A multi-device dataset for urban acoustic scene classification. In Proc. DCASE . 9–13
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proc. ACL . 2556–2565
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proc. NAACL . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristin Toutanova. 2019 · 2019
Earlier work this paper cites.
AudioCaps: Generating Captions for Audios in The Wild. In Proc. NAACL . 119–132
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019 · 2019
Earlier work this paper cites.
ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proc. NIPS . 13–23
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proc. ICCV . 2630–2640
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
Earlier work this paper cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In Proc. ICLR . 1–16
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019 · 2019
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning. In Proc. CVPR . 6720–6731
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Self-supervised multimodal versatile networks. In Proc. NIPS , Vol. 33. 25–37
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. 2020 · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proc. NIPS , Vol. 33. 12449–12460
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Clotho: An audio captioning dataset. In Proc. IEEE ICASSP . 736–740
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2020 · 2020
Cited alongside, same era.
Audio Captioning Based on Combined Audio and Semantic Embeddings. In Proc. ISM . 41–48
Ayşegül Özkaya Eren and Mustafa Sert. 2020 · 2020
Cited alongside, same era.
Audio Retrieval with Natural Language Queries. In Proc. ISCA Interspeech . 2411–2415
Andreea-Maria Oncescu, A. Sophia Koepke, João F. Henriques, Zeynep Akata, and Samuel Albanie. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In Proc. ICML . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Contrastive learning of general-purpose audio representations. In Proc. IEEE ICASSP . 3875–3879
Aaqib Saeed, David Grangier, and Neil Zeghidour. 2021 · 2021
Later among the works it cites.
SimVLM: Simple Visual Language Model Pretraining with Weak Supervision. In Proc. ICLR . 1–17
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021 · 2021
Later among the works it cites.
Investigating local and global information for automated audio captioning with transfer learning. In Proc. IEEE ICASSP . 905–909
Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Zeyu Xie, and Kai Yu. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi, Daiki Takeuchi, and Masahiro Yasuda. 2020 · 2020
Cited alongside, same era.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley. 2020 · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks. In Proc. ECCV . Springer, 121–137
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Cited alongside, same era.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. In Proc. NIPS , Vol. 34. 24206–24221
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021 · 2021
Cited alongside, same era.
Clar: Contrastive learning of auditory representations. In Proc. AISTATS . PMLR, 2530–2538
Haider Al-Tahan and Yalda Mohsenzadeh. 2021 · 2021
Cited alongside, same era.
PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation
Yuan Gong, Yu-An Chung, and James Glass. 2021b · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Enriching Ontology with Temporal Commonsense for Low-Resource Audio Tagging. In Proc. CIKM . 3652–3656
Zhiling Zhang, Zelin Zhou, Haifeng Tang, Guangwei Li, Mengyue Wu, and Kenny Q Zhu. 2021 · 2021
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Later among the works it cites.
Fsd50k: an open dataset of human-labeled sound events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. 2022 · 2022
Later among the works it cites.
Audioclip: Extending clip to image, text and audio. In Proc. IEEE ICASSP . 976–980
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. 2022 · 2022
Later among the works it cites.
Efficient Training of Audio Transformers with Patchout. In Proc. ISCA Interspeech . 2753–2757
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer. 2022 · 2022
Later among the works it cites.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In Proc. ICML . PMLR, 23318–23340
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022 · 2022
Later among the works it cites.
Wav2CLIP: Learning Robust Audio Representations From CLIP. In Proc. IEEE ICASSP
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. 2022 · 2022
Later among the works it cites.
The SJTU System for DCASE2022 Challenge Task 6: Audio Captioning with Audio-Text Retrieval Pre-training
Xuenan Xu, Zeyu Xie, Mengyue Wu, and Kai Yu. 2022 · 2022
Later among the works it cites.
Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer. In Proc. NAACL . 4492–4507
Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Can audio captions be evaluated with image caption metrics?. In Proc. IEEE ICASSP . 981–985
Zelin Zhou, Zhiling Zhang, Xuenan Xu, Zeyu Xie, Mengyue Wu, and Kenny Q Zhu. 2022 · 2022
Later among the works it cites.
Clap learning audio concepts from natural language supervision. In Proc. IEEE ICASSP . IEEE, 1–5
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. 2023 · 2023
Closest in time.