Fetching the paper…
Reading the bibliography…
Multilingual audio-text retrieval (ML-ATR) is a challenging task that aims to retrieve audio clips or multilingual texts from databases.
Language-agnostic bert sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020 · 2007
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019 · 2019
Earlier work this paper cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2020 · 2020
Earlier work this paper cites.
Opus-mt–building open translation services for the world
Jörg Tiedemann and Santhosh Thottingal. 2020 · 2020
Earlier work this paper cites.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. 2022 · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Earlier work this paper cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou. 2022 · 2022
Earlier work this paper cites.
Audio-text retrieval in context
Siyu Lou, Xuenan Xu, Mengyue Wu, and Kai Yu. 2022 · 2022
Earlier work this paper cites.
No language left behind: Scaling human-centered machine translation (2022)
NLLB Team, Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, et al. 2022 · 2022
Earlier work this paper cites.
Wav2clip: Learning robust audio representations from clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. 2022 · 2022
Earlier work this paper cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022 · 2022
Cited alongside, same era.
Dp-sgd without clipping: The lipschitz neural network way
Louis Béthune, Thomas Masséna, Thibaut Boissin, Yannick Prudent, Corentin Friedrich, Franck Mamalet, Aurélien Bellet, Mathieu Serrurier, and David Vigouroux. 2023 · 2023
Cited alongside, same era.
Multilingual audio captioning using machine translated data
Matéo Cousin, Etienne Labbé, and Thomas Pellegrini. 2023 · 2023
Cited alongside, same era.
Sonar: sentence-level multimodal and language-agnostic representations
Paul-Ambroise Duquenne, Holger Schwenk, and Benoît Sagot. 2023 · 2023
Cited alongside, same era.
Clap learning audio concepts from natural language supervision
Ced: Consistent ensemble distillation for audio tagging
Heinrich Dinkel, Yongqing Wang, Zhiyong Yan, Junbo Zhang, and Yujun Wang. 2024 · 2024
Later among the works it cites.
Improving the consistency in cross-lingual cross-modal retrieval with 1-to-k contrastive learning
Zhijie Nie, Richong Zhang, Zhangchi Feng, Hailang Huang, and Xudong Liu. 2024 · 2024
Later among the works it cites.
Gpa: global and prototype alignment for audio-text retrieval
Yuxin Xie, Zhihong Zhu, Xianwei Zhuang, Liming Liang, Zhichang Wang, and Yuexian Zou. 2024 · 2024
Later among the works it cites.
Diffatr: Diffusion-based generative modeling for audio-text retrieval
Yifei Xin, Xuxin Cheng, Zhihong Zhu, Xusheng Yang, and Yuexian Zou. 2024 · 2024
Later among the works it cites.
Bridging language gaps in audio-text retrieval
Zhiyong Yan, Heinrich Dinkel, Yongqing Wang, Jizhong Liu, Junbo Zhang, Yujun Wang, and Bin Wang. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. 2023 · 2023
Cited alongside, same era.
Compa: Addressing the gap in compositional reasoning in audio-language models
Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Evuru, S Ramaneswaran, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha. 2023 · 2023
Cited alongside, same era.
Imbalanced open set domain adaptation via moving-threshold estimation and gradual alignment
Jinghan Ru, Jun Tian, Chengwei Xiao, Jingjing Li, and Heng Tao Shen. 2023 · 2023
Cited alongside, same era.
Collat: on adding fine-grained audio understanding to language models using token-level locked-language tuning
Dadallage AR Silva, Spencer Whitehead, Christopher Lengerich, and Hugh Leather. 2023 · 2023
Cited alongside, same era.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023 · 2023
Cited alongside, same era.
Deepseek llm: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al. 2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Cacophony: An improved contrastive audio-text model
Ge Zhu, Jordan Darefsky, and Zhiyao Duan. 2024 · 2024
Later among the works it cites.
Kdpror: A knowledge-decoupling probabilistic framework for video-text retrieval
Xianwei Zhuang, Hongxiang Li, Xuxin Cheng, Zhihong Zhu, Yuxin Xie, and Yuexian Zou. 2024 · 2024
Later among the works it cites.
Reclap: Improving zero shot audio classification by describing sounds
Sreyan Ghosh, Sonal Kumar, Chandra Kiran Reddy Evuru, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha. 2025 · 2025
Closest in time.
Do we really have to filter out random noise in pre-training data for language models?
Jinghan Ru, Yuxin Xie, Xianwei Zhuang, Yuguo Yin, and Yuexian Zou. 2025 · 2025
Closest in time.
Xianwei Zhuang, Yuxin Xie, Yufan Deng, Liming Liang, Jinghan Ru, Yuguo Yin, and Yuexian Zou. 2025 · 2025
Closest in time.