Fetching the paper…
Reading the bibliography…
Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets.
Musical genre classification of audio signals
George Tzanetakis and Perry Cook. 2002 · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics . 150–157
Chin-Yew Lin and Eduard Hovy. 2003 · 2003
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
A dataset and taxonomy for urban sound research. In Proceedings of the 22nd ACM International Conference on Multimedia . 1041–1044
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello. 2014 · 2014
Earlier work this paper cites.
Look, listen and learn. In Proceedings of the IEEE International Conference on Computer Vision . 609–617
Relja Arandjelovic and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
An overview of audio event detection methods from feature extraction to classification
Elham Babaee, Nor Badrul Anuar, Ainuddin Wahid Abdul Wahab, Shahaboddin Shamshirband, and Anthony T Chronopoulos. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 776–780
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 131–135
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider. In Proceedings of the IEEE International Conference on Computer Vision . 873–881
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy. 2017 · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European Conference on Computer Vision . 631–648
Andrew Owens and Alexei A Efros. 2018 · 2018
Earlier work this paper cites.
Video representation learning by dense predictive coding. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops . 1–10
Tengda Han, Weidi Xie, and Andrew Zisserman. 2019 · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . 119–132
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2630–2640
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. 2020 · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 721–725
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. 2020 · 2020
Cited alongside, same era.
Clotho: An audio captioning dataset. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 736–740
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2020 · 2020
Cited alongside, same era.
TAU Urban Acoustic Scenes 2020 Mobile, Development dataset
Toni Heittola, Annamaria Mesaros, and Tuomas Virtanen. 2020 · 2020
Cited alongside, same era.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley. 2020 · 2020
Cited alongside, same era.
Rethinking CNN models for audio classification
Kamalesh Palanisamy, Dipika Singhania, and Angela Yao. 2020 · 2020
Wav2clip: Learning robust audio representations from clip. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 4563–4567
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. 2022 · 2022
Later among the works it cites.
A comprehensive survey of automated audio captioning
Xuenan Xu, Mengyue Wu, and Kai Yu. 2022 · 2022
Later among the works it cites.
Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics . 4492–4507
Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Audio–visual segmentation. In Proceedings of the European Conference on Computer Vision . 386–403
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Localizing visual sounds the hard way. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16867–16876
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman. 2021 · 2021
Cited alongside, same era.
Diversity and bias in audio captioning datasets. In Detection and Classication of Acoustic Scenes and Events . 90–94
Irene Martin Morato and Annamaria Mesaros. 2021 · 2021
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano. 2021 · 2021
Cited alongside, same era.
Audio retrieval with natural language queries
Andreea-Maria Oncescu, A Koepke, Joao F Henriques, Zeynep Akata, and Samuel Albanie. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 646–650
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022 · 2022
Cited alongside, same era.
Audioclip: Extending clip to image, text and audio. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 976–980
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. 2022 · 2022
Cited alongside, same era.
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
Max Bain, Jaesung Huh, Tengda Han, and Andrew Zisserman. 2023 · 2023
Closest in time.
Acoustic scene classification: a comprehensive survey
Biyun Ding, Tao Zhang, Chao Wang, Ganjun Liu, Jinhua Liang, Ruimin Hu, Yulin Wu, and Difei Guo. 2023 · 2023
Closest in time.
Multi-scale attention for audio question answering
Di Hu Guangyao li, Yixin Xu. 2023 · 2023
Closest in time.
Audio-visual contrastive learning with temporal self-supervision. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 7996–8004
Simon Jenni, Alexander Black, and John Collomosse. 2023 · 2023
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Closest in time.
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D Plumbley, Yuexian Zou, and Wenwu Wang. 2023 · 2023
Closest in time.
AV-SAM: Segment anything model meets audio-visual localization and segmentation
Shentong Mo and Yapeng Tian. 2023b · 2023
Closest in time.
SARdBScene: Dataset and Resnet Baseline for Audio Scene Source Counting and Analysis. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 1–5
Michael Nigro and Sridhar Krishnan. 2023 · 2023
Closest in time.
Learning audio-visual source localization via false negative aware contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6420–6429
Weixuan Sun, Jiayi Zhang, Jianyuan Wang, Zheyuan Liu, Yiran Zhong, Tianpeng Feng, Yandong Guo, Yanhao Zhang, and Nick Barnes. 2023 · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 1–5
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023 · 2023
Closest in time.
Blat: Bootstrapping language-audio pre-training based on audioset tag-guided synthetic data. In Proceedings of the 31st ACM International Conference on Multimedia . 2756–2764
Xuenan Xu, Zhiling Zhang, Zelin Zhou, Pingyue Zhang, Zeyu Xie, Mengyue Wu, and Kenny Q Zhu. 2023 · 2023
Closest in time.
Adaptive sparse pairwise loss for object re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 19691–19701
Xiao Zhou, Yujie Zhong, Zhen Cheng, Fan Liang, and Lin Ma. 2023 · 2023
Closest in time.
Avsegformer: Audio-visual segmentation with transformer. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 12155–12163
Shengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang, and Tong Lu. 2024 · 2024
Closest in time.
Synchformer: Efficient synchronization from sparse cues. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing . 5325–5329
Vladimir Iashin, Weidi Xie, Esa Rahtu, and Andrew Zisserman. 2024 · 2024
Closest in time.
Annotation-free audio-visual segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 5604–5614
Jinxiang Liu, Yu Wang, Chen Ju, Chaofan Ma, Ya Zhang, and Weidi Xie. 2024 · 2024
Closest in time.
Multimodal c4: An open, billion-scale corpus of images interleaved with text
Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi. 2024 · 2024
Closest in time.