Fetching the paper…
Reading the bibliography…
This survey paper provides a comprehensive overview of the recent advancements and challenges in applying large language models to the field of audio signal processing.
D. Griffin and J. Lim, “Signal estimation from modified short-time Fourier transform,” IEEE Transactions on acoustics, speech, and signal processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
F. Jelinek, Statistical methods for speech recognition . MIT press, 1998
1998
Earlier work this paper cites.
V. W. Zue and J. R. Glass, “Conversational interfaces: Advances and challenges,” Proceedings of the IEEE , vol. 88, no. 8, pp. 1166–1180, 2000
2000
Earlier work this paper cites.
E. Matusov, S. Kanthak, and H. Ney, “On the integration of speech recognition and statistical machine translation,” in Ninth European Conference on Speech Communication and Technology , 2005
2005
Earlier work this paper cites.
M. Wester, “The emime bilingual database,” The University of Edinburgh, Tech. Rep., 2010
2010
Earlier work this paper cites.
B. Gold, N. Morgan, and D. Ellis, Speech and audio signal processing: processing and perception of speech and music . John Wiley & Sons, 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems , vol. 25, 2012
2012
Earlier work this paper cites.
J. K. Hansen and I. Fraunhofer, “Recognition of phonemes in a-cappella recordings using temporal patterns and mel frequency cepstral coefficients,” in 9th Sound and Music Computing Conference (SMC) , 2012, pp. 494–499
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Chachada and C.-C. J. Kuo, “Environmental sound recognition: A survey,” APSIPA Transactions on Signal and Information Processing , vol. 3, p. e14, 2014
2014
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE/ACM Transactions on audio, speech, and language processing , vol. 22, no. 10, pp. 1533–1545, 2014
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT press, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Improving neural machine translation models with monolingual data,” in 54th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics (ACL), 2016, pp. 86–96
2016
Earlier work this paper cites.
A. Bérard, O. Pietquin, L. Besacier, and C. Servan, “Listen and translate: A proof of concept for end-to-end speech-to-text translation,” in NIPS Workshop on end-to-end learning for speech and audio processing , 2016
2016
Earlier work this paper cites.
B. Li, T. N. Sainath, A. Narayanan, J. Caroselli, M. Bacchiani, A. Misra, I. Shafran, H. Sak, G. Pundak, K. K. Chin et al. , “Acoustic modeling for google home.” in Interspeech , 2017, pp. 399–403
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “Must-c: a multilingual speech translation corpus,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, 2019, pp. 2012–2017
2017
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” Proc. Interspeech 2017 , pp. 4006–4010, 2017
2017
Earlier work this paper cites.
S. Ö. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman et al. , “Deep voice: Real-time neural text-to-speech,” in International Conference on Machine Learning . PMLR, 2017, pp. 195–204
2017
Earlier work this paper cites.
M. J. Gales, K. M. Knill, and A. Ragni, “Low-resource speech recognition and keyword-spotting,” in Speech and Computer: 19th International Conference, SPECOM 2017, Hatfield, UK, September 12-16, 2017, Proceedings 19 . Springer, 2017, pp. 3–19
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “Cnn architectures for large-scale audio classification,” in 2017 ieee international conference on acoustics, speech and signal processing (icassp) . IEEE, 2017, pp. 131–135
2017
Earlier work this paper cites.
M. B. Hoy, “Alexa, siri, cortana, and more: an introduction to voice assistants,” Medical reference services quarterly , vol. 37, no. 1, pp. 81–88, 2018
2018
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Tasnet: time-domain audio separation network for real-time, single-channel speech separation,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 696–700
2018
Earlier work this paper cites.
W. Ping, K. Peng, and J. Chen, “Clarinet: Parallel wave generation in end-to-end text-to-speech,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
J. Gao, M. Galley, and L. Li, “Neural approaches to conversational ai,” in The 41st international ACM SIGIR conference on research & development in information retrieval , 2018, pp. 1371–1374
2018
Earlier work this paper cites.
I. V. Serban, R. Lowe, P. Henderson, L. Charlin, and J. Pineau, “A survey of available corpora for building data-driven dialogue systems: The journal version,” Dialogue & Discourse , vol. 9, no. 1, pp. 1–49, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
T. Yoshioka, I. Abramovski, C. Aksoylar, Z. Chen, M. David, D. Dimitriadis, Y. Gong, I. Gurvich, X. Huang, Y. Huang et al. , “Advances in online audio-visual meeting transcription,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 276–283
2019
Earlier work this paper cites.
H. Purwins, B. Li, T. Virtanen, J. Schlüter, S.-Y. Chang, and T. Sainath, “Deep learning for audio signal processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 2, pp. 206–219, 2019
2019
Earlier work this paper cites.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang et al. , “A comparative study on transformer vs rnn in speech applications,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 449–456
2019
Earlier work this paper cites.
K. Miyazaki, T. Toda, T. Hayashi, and K. Takeda, “Environmental sound processing and its applications,” IEEJ Transactions on Electrical and Electronic Engineering , vol. 14, no. 3, pp. 340–351, 2019
2019
Earlier work this paper cites.
Y. Yu, X. Si, C. Hu, and J. Zhang, “A review of recurrent neural networks: Lstm cells and network architectures,” Neural computation , vol. 31, no. 7, pp. 1235–1270, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 119–132
2019
Earlier work this paper cites.
D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra, “The mtg-jamendo dataset for automatic music tagging.” ICML, 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
D. Stoller, S. Durand, and S. Ewert, “End-to-end lyrics alignment for polyphonic music using an audio-to-character recognition model,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 181–185
2019
Earlier work this paper cites.
G. R. Dabike and J. Barker, “Automatic lyric transcription from karaoke vocal tracks: Resources and a baseline system.” in Interspeech , 2019, pp. 579–583
2019
Earlier work this paper cites.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3617–3621
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Turpault, R. Serizel, A. P. Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Workshop on Detection and Classification of Acoustic Scenes and Events , 2019
2019
Earlier work this paper cites.
N. Carlini, C. Liu, U. Erlingsson, J. Kos, and D. Song, “The secret sharer: Evaluating and testing unintended memorization in neural networks,” in 28th USENIX Security Symposium (USENIX Security 19) . USENIX Association, 2019, pp. 267–284
2019
Earlier work this paper cites.
S. Latif, J. Qadir, A. Qayyum, M. Usama, and S. Younis, “Speech technology for healthcare: Opportunities, challenges, and state of the art,” IEEE Reviews in Biomedical Engineering , vol. 14, pp. 342–356, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7669–7673
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
C. Wang, J. Pino, A. Wu, and J. Gu, “Covost: A diverse multilingual speech-to-text translation corpus,” in Proceedings of the Twelfth Language Resources and Evaluation Conference , 2020, pp. 4197–4203
2020
Earlier work this paper cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 736–740
2020
Earlier work this paper cites.
G. Meseguer-Brocal, A. Cohen-Hadria, and G. Peeters, “Creating dali, a large dataset of synchronized audio, lyrics, and notes,” Transactions of the International Society for Music Information Retrieval , vol. 3, no. 1, 2020
2020
Earlier work this paper cites.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “Vggsound: A large-scale audio-visual dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 721–725
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7871–7880
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Fan and X. Luo, “A survey of dialogue system evaluation,” in 32nd IEEE International Conference on Tools with Artificial Intelligence ICTAI . IEEE, 2020
2020
Earlier work this paper cites.
J.-P. Briot and F. Pachet, “Deep learning for music generation: challenges and directions,” Neural Computing and Applications , vol. 32, no. 4, pp. 981–993, 2020
2020
Earlier work this paper cites.
Y.-S. Huang and Y.-H. Yang, “Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions,” in Proceedings of the 28th ACM international conference on multimedia , 2020, pp. 1180–1188
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are RNNs: Fast autoregressive transformers with linear attention,” in Proceedings of the 37th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, H. D. III and A. Singh, Eds., vol. 119. PMLR, 13–18 Jul 2020, pp. 5156–5165. [Online]. Available: https://proceedings.mlr.press/v119/katharopoulos20a.html
2020
Earlier work this paper cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Latif, R. Rana, S. Khalifa, R. Jurdak, and B. W. Schuller, “Deep architecture enhancing robustness to noise, adversarial attacks, and cross-corpus setting for speech emotion recognition,” Proc. Interspeech 2020 , pp. 2327–2331, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. B. Cyphert, “A human being wrote this law review article: Gpt-3 and the practice of law,” UC Davis L. Rev. , vol. 55, p. 401, 2021
2021
Earlier work this paper cites.
A. Łańcucki, “Fastpitch: Parallel text-to-speech with pitch prediction,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6588–6592
2021
Earlier work this paper cites.
M. Zeng, X. Tan, R. Wang, Z. Ju, T. Qin, and T.-Y. Liu, “Musicbert: Symbolic music understanding with large-scale pre-training,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 791–800
2021
Earlier work this paper cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed et al. , “On generative spoken language modeling from raw audio,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1336–1354, 2021
2021
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Earlier work this paper cites.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu, “W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2021, pp. 244–250
2021
Earlier work this paper cites.
Y. Li, M. Tagliasacchi, O. Rybakov, V. Ungureanu, and D. Roblek, “Real-time speech frequency bandwidth extension,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 691–695
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Ens and P. Pasquier, “Building the metamidi dataset: Linking symbolic and audio musical data.” in ISMIR , 2021, pp. 182–188
2021
Earlier work this paper cites.
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 829–852, 2021
2021
Earlier work this paper cites.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “Vinvl: Revisiting visual representations in vision-language models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 5579–5588
2021
Earlier work this paper cites.
K. Schulze-Forster, C. S. Doire, G. Richard, and R. Badeau, “Phoneme level lyrics alignment and text-informed singing voice separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2382–2395, 2021
2021
Earlier work this paper cites.
Z. Li, Z. Li, J. Zhang, Y. Feng, and J. Zhou, “Bridging text and video: A universal multimodal transformer for audio-visual scene-aware dialog,” IEEE ACM Trans. Audio Speech Lang. Process. (TSLP) , vol. 29, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini, “Deduplicating training data makes language models better,” 2021
2021
Earlier work this paper cites.
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown et al. , “Extracting training data from large language models,” in 30th USENIX Security Symposium (USENIX Security 21) . USENIX Association, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
W.-C. Huang, C.-H. Wu, S.-B. Luo, K.-Y. Chen, H.-M. Wang, and T. Toda, “Speech recognition by simply fine-tuning bert,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7343–7347
2021
Cited alongside, same era.
S. Latif, R. Rana, S. Khalifa, R. Jurdak, J. Qadir, and B. W. Schuller, “Survey of deep representation learning for speech emotion recognition,” IEEE Transactions on Affective Computing , 2021
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys , vol. 55, no. 12, pp. 1–38, 2023
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” https://arxiv.org/pdf/2303.08774.pdf , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Cited alongside, same era.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022
2022
Cited alongside, same era.
Z. Gan, L. Li, C. Li, L. Wang, Z. Liu, J. Gao et al. , “Vision-language pre-training: Basics, recent advances, and future trends,” Foundations and Trends® in Computer Graphics and Vision , vol. 14, no. 3–4, pp. 163–352, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
G. Deshpande, A. Batliner, and B. W. Schuller, “Ai-based human audio processing for covid-19: A comprehensive overview,” Pattern recognition , vol. 122, p. 108289, 2022
2022
Cited alongside, same era.
S. Liu, A. Mallol-Ragolta, E. Parada-Cabaleiro, K. Qian, X. Jing, A. Kathan, B. Hu, and B. W. Schuller, “Audio self-supervised learning: A survey,” Patterns , vol. 3, no. 12, 2022
2022
Cited alongside, same era.
F. F. Xu, U. Alon, G. Neubig, and V. J. Hellendoorn, “A systematic evaluation of large language models of code,” in Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming , 2022, pp. 1–10
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour, “Audiolm: A language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2523–2533, 2023
2023
Closest in time.
2023
Closest in time.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” See https://vicuna. lmsys. org (accessed 14 April 2023) , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Cheng, Y. Zhang, M. Johnson, W. Macherey, and A. Bapna, “Mu 2
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua, “Next-gpt: Any-to-any multimodal llm,” 2023
2023
Closest in time.
2023
Closest in time.
Y. Higuchi, T. Ogawa, T. Kobayashi, and S. Watanabe, “Bectra: Transducer-based end-to-end asr with bert-enhanced encoder,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. Maiti, Y. Peng, T. Saeki, and S. Watanabe, “Speechlmscore: Evaluating speech generation using speech language model,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Zhang, T. Wu, X. Chen, S. Wen, S. Nepal, C. Paris, and Y. Xiang, “Dynalogue: A transformer-based dialogue system with dynamic attention,” in Proceedings of the ACM Web Conference 2023 , 2023, pp. 1604–1615
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. López-Zorrilla, M. I. Torres, and H. Cuayáhuitl, “Audio embedding-aware dialogue policy learning,” IEEE ACM Trans. Audio Speech Lang. Process. (TASLP) , vol. 3, 2023
2023
Closest in time.
2023
Closest in time.
T. A. Nguyen, E. Kharitonov, J. Copet, Y. Adi, W.-N. Hsu, A. Elkahky, P. Tomasello, R. Algayres, B. Sagot, A. Mohamed, and E. Dupoux, “Generative Spoken Dialogue Language Modeling,” Transactions of the Association for Computational Linguistics , vol. 11, 2023. [Online]. Available: https://doi.org/10.1162/tacl_a_00545
2023
Closest in time.
T. Gong, C. Lyu, S. Zhang, Y. Wang, M. Zheng, Q. Zhao, K. Liu, W. Zhang, P. Luo, and K. Chen, “Multimodal-gpt: A vision and language model for dialogue with humans,” 2023
2023
Closest in time.
C. Li, “Large multimodal models: Notes on cvpr 2023 tutorial,” 2023
2023
Closest in time.
2023
Closest in time.
S. Ji and X. Yang, “Emomusictv: Emotion-conditioned symbolic music generation with hierarchical transformer vae,” IEEE Transactions on Multimedia , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y.-J. Shih, H.-F. Wang, H.-J. Chang, L. Berry, H.-y. Lee, and D. Harwath, “Speechclip: Integrating speech with pre-trained vision and language model,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 715–722
2023
Closest in time.
L. Xu, L. Wang, S. Bi, H. Liu, and J. Wang, “Semi-supervised sound event detection with pre-trained model,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
M. C. Rillig, M. Ågerstrand, M. Bi, K. A. Gould, and U. Sauerland, “Risks and benefits of large language models for the environment,” Environmental Science & Technology , vol. 57, no. 9, pp. 3464–3466, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Geiping and T. Goldstein, “Cramming: Training a language model on a single gpu in one day.” in International Conference on Machine Learning . PMLR, 2023, pp. 11 117–11 143
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
M. Cascella, J. Montomoli, V. Bellini, and E. Bignami, “Evaluating the feasibility of chatgpt in healthcare: an analysis of multiple clinical and research scenarios,” Journal of Medical Systems , vol. 47, no. 1, p. 33, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
W. Sun, Z. Shi, S. Gao, P. Ren, M. de Rijke, and Z. Ren, “Contrastive learning reduces hallucination in conversations,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 11, 2023, pp. 13 618–13 626
2023
Closest in time.
P. P. Ray, “Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope,” Internet of Things and Cyber-Physical Systems , 2023
2023
Closest in time.