Lamda: Language models for dialog applications
Original
Cohen, A. D., Roberts, A., Molina, A., Butryna, A., Jin, A., Kulshreshtha, A., Hutchinson, B., Zevenbergen, B., Aguera-Arcas, B. H., ching Chang, C., Cui, C., Du, C., Adiwardana, D. D. F., Chen, D., Lepikhin, D. D., Chi, E. H., Hoffman-John, E., Cheng, H.-T., Lee, H., Krivokon, I., Qin, J., Hall, J., Fenton, J., Soraker, J., Meier-Hellstern, K., Olson, K., Aroyo, L. M., Bosma, M. P., Pickett, M. J., Menegali, M. A., Croak, M., Díaz, M., Lamm, M., Krikun, M., Morris, M. R., Shazeer, N., Le, Q. V., Bernstein, R., Rajakumar, R., Kurzweil, R., Thoppilan, R., Zheng, S., Bos, T., Duke, T., Doshi, T., Zhao, V. Y., Prabhakaran, V., Rusch, W., Li, Y., Huang, Y., Zhou, Y., Xu, Y., and Chen, Z · 2022
Later among the works it cites.
High fidelity neural audio compression
Original
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Later among the works it cites.
Riffusion - Stable diffusion for real-time music generation, 2022
Forsgren, S. and Martiros, H · 2022
Later among the works it cites.
Video diffusion models
Original
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Later among the works it cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Original
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Later among the works it cites.
Mulan: A joint embedding of music audio and natural language
Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P. W · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation, 2022
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2022
Later among the works it cites.
GLIDE: towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Original
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Original
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M · 2022
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual description
Original
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Original
Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation, 2022
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y · 2022
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2022
Later among the works it cites.