Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition.
On integrating a language model into neural machine translation
Gulcehre, C., Firat, O., Xu, K., Cho, K., and Bengio, Y · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Neuralwarp: Time-series similarity with warping networks
Grabocka, J., Schmidt Thieme, L., and etc · 2018
Earlier work this paper cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Luo, Y. and Mesgarani, N · 2019
Earlier work this paper cites.
Pythia: Ai-assisted code completion system
Svyatkovskiy, A., Zhao, Y., Fu, S., and Sundaresan, N · 2019
Earlier work this paper cites.
Language model prior for low-resource neural machine translation
Baziotis, C., Haddow, B., and Birch, A · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Discretalk: Text-to-speech as a machine translation problem
Hayashi, T. and Watanabe, S · 2020
Earlier work this paper cites.
A simple language model for task-oriented dialogue
Hosseini-Asl, E., McCann, B., Wu, C.-S., Yavuz, S., and Socher, R · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Multi-task learning based pre-trained language model for code completion
Liu, F., Li, G., Zhao, Y., and Jin, Z · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al · 2020
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2020
Earlier work this paper cites.
Searchable hidden intermediates for end-to-end models of decomposable sequence tasks
Dalmia, S., Yan, B., Raunak, V., Metze, F., and Watanabe, S · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Cited alongside, same era.
Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus
Huang, R., Chen, F., Ren, Y., Liu, J., Cui, C., and Zhao, Z · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
Imitating arbitrary talking style for realistic audio-driven talking face synthesis
Wu, H., Jia, J., Wang, H., Dou, Y., Duan, C., and Deng, Q · 2021
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Tf-gridnet: Making time-frequency domain models great again for monaural speaker separation
Wang, Z.-Q., Cornell, S., Choi, S., Lee, Y., Kim, B.-Y., and Watanabe, S · 2022
Later among the works it cites.
Audio pyramid transformer with domain adaption for weakly supervised sound event detection and audio classification
Xin, Y., Yang, D., and Zou, Y · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ye, Z., Wang, H., Yang, D., and Zou, Y · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Cited alongside, same era.
LangChain, 10 2022
Chase, H · 2022
Cited alongside, same era.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Cited alongside, same era.
textless-lib: A library for textless spoken language processing
Kharitonov, E., Copet, J., Lakhotia, K., Nguyen, T. A., Tomasello, P., Lee, A., Elkahky, A., Hsu, W.-N., Mohamed, A., Dupoux, E., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D · 2022
Later among the works it cites.
Visinger: Variational inference with adversarial learning for end-to-end singing voice synthesis
Zhang, Y., Cong, J., Xue, H., Xie, L., Zhu, P., and Bi, M · 2022
Later among the works it cites.
Musiclm: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Closest in time.
Generative spoken dialogue language modeling
Nguyen, T. A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W.-N., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., et al · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N · 2023
Closest in time.
Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis
Ye, Z., Jiang, Z., Ren, Y., Liu, J., He, J., and Zhao, Z · 2023
Closest in time.