Fetching the paper…
Reading the bibliography…
Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale video datasets.
“Speech Discrimination by Dynamic Programming,”
T. K. Vintsyuk, · 1968
Earlier work this paper cites.
“Learning a Similarity Metric Discriminatively, with Application to Face Verification,”
S. Chopra, R. Hadsell, and Y. LeCun, · 2005
Earlier work this paper cites.
“Learning with a Wasserstein Loss,”
C. Frogner, C. Zhang, H. Mobahi, M. Araya, and T. A. Poggio, · 2015
Earlier work this paper cites.
“Look, Listen and Learn,”
R. Arandjelovic and A. Zisserman, · 2017
Earlier work this paper cites.
“Attention is All you Need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Soft-DTW: A Differentiable Loss Function for Time-series,”
M. Cuturi and M. Blondel, · 2017
Earlier work this paper cites.
“Audio Set: An Ontology and Human-labeled Dataset for Audio Events,”
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, · 2017
Earlier work this paper cites.
“Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization,”
B. Korbar, D. Tran, and L. Torresani, · 2018
Earlier work this paper cites.
“The Sound of Pixels,”
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, · 2018
Earlier work this paper cites.
“Objects that Sound,”
R. Arandjelovic and A. Zisserman, · 2018
Earlier work this paper cites.
“Audio-Visual Scene Analysis with Self-Supervised Multisensory Features,”
A. Owens and A. A. Efros, · 2018
Earlier work this paper cites.
“Representation Learning with Contrastive Predictive Coding,” 2019
A. van den Oord, Y. Li, and O. Vinyals, · 2019
Earlier work this paper cites.
“Computational Optimal Transport: With Applications to Data Science,”
G. Peyré and M. Cuturi, · 2019
Earlier work this paper cites.
“The Sound of Motions,”
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba, · 2019
Earlier work this paper cites.
“Decoupled Weight Decay Regularization,”
I. Loshchilov and F. Hutter, · 2019
Earlier work this paper cites.
“VGGSound: A Large-scale Audio-Visual Dataset,”
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, · 2020
Cited alongside, same era.
“Self-supervised Learning of Audio-Visual Objects from Video,”
T. Afouras, A. Owens, J. S. Chung, and A. Zisserman, · 2020
Cited alongside, same era.
“Learning Transferable Visual Models From Natural Language Supervision,”
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, · 2021
Cited alongside, same era.
“Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision,”
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, · 2021
Cited alongside, same era.
“Audio-Visual Instance Discrimination with Cross-Modal Agreement,”
P. Morgado, N. Vasconcelos, and I. Misra, · 2021
Cited alongside, same era.
“Active Contrastive Learning of Audio-Visual Video Representations,”
“The Power of Sound (TPoS): Audio Reactive Video Generation with Stable Diffusion,”
Y. Jeong, W. Ryoo, S. Lee, D. Seo, W. Byeon, S. Kim, and J. Kim, · 2023
Later among the works it cites.
“Any-to-Any Generation via Composable Diffusion,”
Z. Tang, Z. Yang, C. Zhu, M. Zeng, and M. Bansal, · 2023
Later among the works it cites.
“Contrastive Audio-Visual Masked Autoencoder,”
Y. Gong, A. Rouditchenko, A. H. Liu, D. Harwath, L. Karlinsky, H. Kuehne, and J. R. Glass, · 2023
Later among the works it cites.
“ImageBind: One Embedding Space To Bind Them All,”
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra, · 2023
Later among the works it cites.
“Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,”
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, · 2023
Later among the works it cites.
“Learning Audio-Visual Source Localization via False Negative Aware Contrastive Learning,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ma, Z. Zeng, D. McDuff, and Y. Song, · 2021
Cited alongside, same era.
“Robust Audio-Visual Instance Discrimination,”
P. Morgado, I. Misra, and N. Vasconcelos, · 2021
Cited alongside, same era.
“SeqMatchNet: Contrastive Learning with Sequence Matching for Place Recognition & Relocalization,”
S. Garg, M. Vankadari, and M. Milford, · 2021
Cited alongside, same era.
“Sequence-to-sequence contrastive learning for text recognition,”
A. Aberdam, R. Litman, S. Tsiper, O. Anschel, R. Slossberg, S. Mazor, R. Manmatha, and P. Perona, · 2021
Cited alongside, same era.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,”
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, · 2021
Cited alongside, same era.
“Audioclip: Extending Clip to Image, Text and Audio,”
A. Guzhov, F. Raue, J. Hees, and A. Dengel, · 2022
Cited alongside, same era.
“On Negative Sampling for Audio-Visual Contrastive Learning from Movies,” 2022
M. M. Kalayeh, S. Ardeshir, L. Liu, N. Kamath, and A. Chandrashekar, · 2022
Cited alongside, same era.
W. Sun, J. Zhang, J. Wang, Z. Liu, Y. Zhong, T. Feng, Y. Guo, Y. Zhang, and N. Barnes, · 2023
Later among the works it cites.
“MAViL: Masked Audio-Video Learners,”
P.-Y. Huang, V. Sharma, H. Xu, C. Ryali, H. Fan, Y. Li, S.-W. Li, G. Ghosh, J. Malik, and C. Feichtenhofer, · 2023
Later among the works it cites.
“Audio-Visual Contrastive Learning with Temporal Self-Supervision,”
S. Jenni, A. Black, and J. Collomosse, · 2023
Later among the works it cites.
“TempCLR: Temporal Alignment Representation with Contrastive Learning,”
Y. Yang, J. Ma, S. Huang, L. Chen, X. Lin, G. Han, and S.-F. Chang, · 2023
Later among the works it cites.
“BEATs: Audio Pre-Training with Acoustic Tokenizers,”
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, · 2023
Later among the works it cites.
“The Llama 3 Herd of Models,” 2024
Llama-Team, · 2024
Closest in time.
“DIFF-FOLEY: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models,”
S. Luo, C. Yan, C. Hu, and H. Zhao, · 2024
Closest in time.
“Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity,” 2024
S. Pascual, C. Yeh, I. Tsiamas, and J. Serrà, · 2024
Closest in time.
“Natural Language Supervision For General-Purpose Audio Representations,”
B. Elizalde, S. Deshmukh, and H. Wang, · 2024
Closest in time.
“Pushing the Limits of Zero-shot End-to-End Speech Translation,”
I. Tsiamas, G. Gállego, J. Fonollosa, and M. Costa-jussà, · 2024
Closest in time.