Fetching the paper…
Reading the bibliography…
In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
K. J. Piczak · 2015
Earlier work this paper cites.
Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos
A. Zadeh, R. Zellers, E. Pincus, and L.-P. Morency · 2016
Earlier work this paper cites.
FMA: A dataset for music analysis
M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson · 2017
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan · 2017
Earlier work this paper cites.
Freesound datasets: a platform for the creation of open audio datasets
E. Fonseca, J. Pons Puig, X. Favory, F. Font Corbera, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings
R. Lotfian and C. Busso · 2017
Earlier work this paper cites.
Tackling toxic online communication with recurrent capsule networks
S. Deshmukh and R. Rade · 2018
Earlier work this paper cites.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency · 2018
Earlier work this paper cites.
Cross modal audio search and retrieval with joint embeddings based on text and audio
B. Elizalde, S. Zarar, and B. Raj · 2019
Earlier work this paper cites.
AudioCaps: Generating Captions for Audios in The Wild
C. D. Kim, B. Kim, H. Lee, and G. Kim · 2019
Earlier work this paper cites.
Visualbert: Asimple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Earlier work this paper cites.
Meld: A multimodal multi-party dataset for emotion recognition in conversations
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Clotho: an audio captioning dataset
K. Drossos, S. Lipping, and T. Virtanen · 2020
Earlier work this paper cites.
Never-ending learning of sounds
B. M. Elizalde · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Towards learning a universal non-semantic representation of speech
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. de Chaumont Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv · 2020
Earlier work this paper cites.
Unsupervised contrastive learning of sound event representations
E. Fonseca, D. Ortego, K. McGuinness, N. E. O’Connor, and X. Serra · 2021
Cited alongside, same era.
Fsd50k: An open dataset of human-labeled sound events
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Audiolm: a language modeling approach to audio generation
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour · 2022
Later among the works it cites.
Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Later among the works it cites.
Describing emotions with acoustic property prompts for speech emotion recognition
H. Dhamyal, B. Elizalde, S. Deshmukh, H. Wang, B. Raj, and R. Singh · 2022
Later among the works it cites.
Clap: Learning audio concepts from natural language supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. L. Li and P. Liang · 2021
Cited alongside, same era.
What is the ground truth? reliability of multi-annotator data for audio tagging
I. Martín-Morató and A. Mesaros · 2021
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
R. Mokady, A. Hertz, and A. H. Bermano · 2021
Cited alongside, same era.
Byol for audio: Self-supervised learning for general-purpose audio representation
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino · 2021
Cited alongside, same era.
Byol for audio: Self-supervised learning for general-purpose audio representation
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, et al · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Cited alongside, same era.
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang · 2022
Later among the works it cites.
Ssast: Self-supervised audio spectrogram transformer
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass · 2022
Later among the works it cites.
Audioclip: Extending clip to image, text and audio
A. Guzhov, F. Raue, J. Hees, and A. Dengel · 2022
Later among the works it cites.
Cochlscene: Acquisition of acoustic scene data using crowdsourcing
I.-Y. Jeong and J. Park · 2022
Later among the works it cites.
Audio retrieval with natural language queries: A benchmark study
A. S. Koepke, A.-M. Oncescu, J. Henriques, Z. Akata, and S. Albanie · 2022
Later among the works it cites.
Clotho-aqa: A crowdsourced dataset for audio question answering
S. Lipping, P. Sudarsanam, K. Drossos, and T. Virtanen · 2022
Later among the works it cites.
HEAR: Holistic Evaluation of Audio Representations
J. Turian, J. Shier, et al · 2022
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip
H.-H. Wu, P. Seetharaman, K. Kumar, et al · 2022
Later among the works it cites.
Musiclm: Generating music from text
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, et al · 2023
Closest in time.
Audio Retrieval with WavText5K and CLAP Training
S. Deshmukh, B. Elizalde, and H. Wang · 2023
Closest in time.
Prompting audios using acoustic properties for emotion representation
H. Dhamyal, B. Elizalde, S. Deshmukh, H. Wang, B. Raj, and R. Singh · 2023
Closest in time.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Closest in time.
Prefix tuning for automated audio captioning
M. Kim, K. Sung-Bin, and T.-H. Oh · 2023
Closest in time.
MAPL: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting
O. Mañas, P. Rodriguez Lopez, S. Ahmadi, A. Nematzadeh, Y. Goyal, and A. Agrawal · 2023
Closest in time.
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang · 2023
Closest in time.
Improving weakly supervised sound event detection with self-supervised auxiliary tasks
S. Deshmukh, B. Raj, and R. Singh · 2079
Closest in time.