Fetching the paper…
Reading the bibliography…
Building artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research.
AutoFoley: Artificial synthesis of synchronized sound tracks for silent videos with deep learning
Ghose, S.; and Prevost, J. J. 2020 · 1907
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M.; and Cohen, N. J. 1989 · 1989
Earlier work this paper cites.
Analysis of ordinal paired comparison data
Agresti, A. 1992 · 1992
Earlier work this paper cites.
ITU-T Recommendation P.800: Methods for Subjective Determination of Transmission Quality
International Telecommunication Union. 1996 · 1996
Earlier work this paper cites.
Visually indicated sounds
Owens, A.; Isola, P.; McDermott, J.; Torralba, A.; Adelson, E. H.; and Freeman, W. T. 2016 · 2016
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Chen, L.; Srivastava, S.; Duan, Z.; and Xu, C. 2017 · 2017
Earlier work this paper cites.
AudioSet: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Visually indicated sound generation by perceptually optimized classification
Chen, K.; Zhang, C.; Fang, C.; Wang, Z.; Bui, T.; and Nevatia, R. 2018 · 2018
Earlier work this paper cites.
CMCGAN: A uniform framework for cross-modal visual-audio mutual generation
Hao, W.; Zhang, Z.; and Guan, H. 2018 · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
Owens, A.; and Efros, A. A. 2018 · 2018
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild
Zhou, Y.; Wang, Z.; Fang, C.; Bui, T.; and Berg, T. L. 2018 · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kilgour, K.; Zuluaga, M.; Roblek, D.; and Sharifi, M. 2019 · 2019
Earlier work this paper cites.
AudioCaps: Generating captions for audios in the wild
Kim, C. D.; Kim, B.; Lee, H.; and Kim, G. 2019 · 2019
Earlier work this paper cites.
Clotho: An audio captioning dataset
Drossos, K.; Lipping, S.; and Virtanen, T. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
PANNs: Large-scale pretrained audio neural networks for audio pattern recognition
Kong, Q.; Cao, Y.; Iqbal, T.; Wang, Y.; Wang, W.; and Plumbley, M. D. 2020 · 2020
Earlier work this paper cites.
Learning individual speaking styles for accurate lip to speech synthesis
Prajwal, K.; Mukhopadhyay, R.; Namboodiri, V. P.; and Jawahar, C. 2020 · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. 2021 · 2021
Cited alongside, same era.
Classifier-Free Diffusion Guidance
Ho, J.; and Salimans, T. 2021 · 2021
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2021 · 2021
Cited alongside, same era.
Taming visually guided sound generation
Iashin, V.; and Rahtu, E. 2021 · 2021
Cited alongside, same era.
How many data points is a prompt worth?
Wav2CLIP: Learning robust audio representations from CLIP
Wu, H.-H.; Seetharaman, P.; Kumar, K.; and Bello, J. P. 2022 · 2022
Later among the works it cites.
An empirical study of GPT-3 for few-shot knowledge-based VQA
Yang, Z.; Gan, Z.; Wang, J.; Hu, X.; Lu, Y.; Liu, Z.; and Wang, L. 2022 · 2022
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B.; Goldberg, Y.; and Ravfogel, S. 2022 · 2022
Later among the works it cites.
Foundational models defining a new era in vision: A survey and outlook
Awais, M.; Naseer, M.; Khan, S.; Anwer, R. M.; Cholakkal, H.; Shah, M.; Yang, M.-H.; and Khan, F. S. 2023 · 2023
Closest in time.
Cao, Y.; Li, S.; Liu, Y.; Yan, Z.; Dai, Y.; Yu, P. S.; and Sun, L. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Le Scao, T.; and Rush, A. M. 2021 · 2021
Cited alongside, same era.
Pretrained transformers as universal computation engines
Lu, K.; Grover, A.; Abbeel, P.; and Mordatch, I. 2021 · 2021
Cited alongside, same era.
ClipCap: Clip prefix for image captioning
Mokady, R.; Hertz, A.; and Bermano, A. H. 2021 · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Nichol, A. Q.; and Dhariwal, P. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Parametric UMAP embeddings for representation and semisupervised learning
Sainburg, T.; McInnes, L.; and Gentner, T. Q. 2021 · 2021
Cited alongside, same era.
ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection
Yamagishi, J.; Wang, X.; Todisco, M.; Sahidullah, M.; Patino, J.; Nautsch, A.; Liu, X.; Lee, K. A.; Kinnunen, T.; Evans, N.; and Delgado, H. 2021 · 2021
Cited alongside, same era.
CLIPSonic: Text-to-audio synthesis with unlabeled videos and pretrained language-vision models
Dong, H.-W.; Liu, X.; Pons, J.; Bhattacharya, G.; Pascual, S.; Serrà, J.; Berg-Kirkpatrick, T.; and McAuley, J. 2023 · 2023
Closest in time.
Text-to-Audio Generation using Instruction Guided Latent Diffusion Model
Ghosal, D.; Majumder, N.; Mehrish, A.; and Poria, S. 2023 · 2023
Closest in time.
Make-An-Audio: Text-to-audio generation with prompt-enhanced diffusion models
Huang, R.; Huang, J.; Yang, D.; Ren, Y.; Liu, L.; Li, M.; Ye, Z.; Liu, J.; Yin, X.; and Zhao, Z. 2023 · 2023
Closest in time.
AudioGen: Textually guided audio generation
Kreuk, F.; Synnaeve, G.; Polyak, A.; Singer, U.; Défossez, A.; Copet, J.; Parikh, D.; Taigman, Y.; and Adi, Y. 2023 · 2023
Closest in time.
Efficient domain adaptation for speech foundation models
Li, B.; Hwang, D.; Huo, Z.; Bai, J.; Prakash, G.; Sainath, T. N.; Sim, K. C.; Zhang, Y.; Han, W.; Strohman, T.; et al. 2023 · 2023
Closest in time.
Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models
Luo, S.; Yan, C.; Hu, C.; and Zhao, H. 2023 · 2023
Closest in time.
Foundation models for natural language processing: Pre-trained language models integrating media
Paaß, G.; and Giesselbach, S. 2023 · 2023
Closest in time.
I hear your true colors: Image guided audio generation
Sheffer, R.; and Adi, Y. 2023 · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
Closest in time.
DiffSound: Discrete diffusion model for text-to-sound generation
Yang, D.; Yu, J.; Wang, H.; Wang, W.; Weng, C.; Zou, Y.; and Yu, D. 2023 · 2023
Closest in time.
A survey on multimodal large language models
Yin, S.; Fu, C.; Zhao, S.; Li, K.; Sun, X.; Xu, T.; and Chen, E. 2023 · 2023
Closest in time.
Leveraging pre-trained AudioLDM for sound generation: A benchmark study
Yuan, Y.; Liu, H.; Liang, J.; Liu, X.; Plumbley, M. D.; and Wang, W. 2023 · 2023
Closest in time.
Meta-Transformer: A Unified Framework for Multimodal Learning
Zhang, Y.; Gong, K.; Zhang, K.; Li, H.; Qiao, Y.; Ouyang, W.; and Yue, X. 2023 · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT
Zhou, C.; Li, Q.; Li, C.; Yu, J.; Liu, Y.; Wang, G.; Zhang, K.; Ji, C.; Yan, Q.; He, L.; et al. 2023 · 2023
Closest in time.
Video background music generation with controllable music transformer
Di, S.; Jiang, Z.; Liu, S.; Wang, Z.; Zhu, L.; He, Z.; Liu, H.; and Yan, S. 2021 · 2045
Closest in time.