Fetching the paper…
Reading the bibliography…
In recent years, image generation has shown a great leap in performance, where diffusion models play a central role.
2016
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” NeurIPS , 2017
2017
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 5329–5333
2018
Earlier work this paper cites.
C.-H. Wan, S.-P. Chuang, and H.-Y. Lee, “Towards audio to scene image synthesis using generative adversarial network,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 496–500
2019
Earlier work this paper cites.
I. Schwartz, S. Yu, T. Hazan, and A. G. Schwing, “Factor graph attention,” in CVPR , 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “Vggsound: A large-scale audio-visual dataset,” in ICASSP , 2020
2020
Earlier work this paper cites.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed et al. , “On generative spoken language modeling from raw audio,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1336–1354, 2021
2021
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” Advances in Neural Information Processing Systems , vol. 34, pp. 8780–8794, 2021
2021
Earlier work this paper cites.
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in International Conference on Machine Learning . PMLR, 2021, pp. 8162–8171
2021
Earlier work this paper cites.
V. Iashin and E. Rahtu, “Taming visually guided sound generation,” in British Machine Vision Conference (BMVC) , 2021
2021
Earlier work this paper cites.
R. Gao and K. Grauman, “Visualvoice: Audio-visual speech separation with cross-modal consistency,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2021, pp. 15 490–15 500
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Cited alongside, same era.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 12 873–12 883
2021
Cited alongside, same era.
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman, “Make-a-scene: Scene-based text-to-image generation with human priors,” in ECCV , 2022
2022
Cited alongside, same era.
Y. Tewel, Y. Shalev, I. Schwartz, and L. Wolf, “Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,” in CVPR , 2022
2022
Later among the works it cites.
F. Kreuk, A. Polyak, J. Copet, E. Kharitonov, T. A. Nguyen, M. Rivière, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y. Adi, “Textless speech emotion conversion using discrete & decomposed representations,” in EMNLP , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. Mcgrew, I. Sutskever, and M. Chen, “Glide: Towards photorealistic image generation and editing with text-guided diffusion models,” in International Conference on Machine Learning . PMLR, 2022, pp. 16 784–16 804
2022
Cited alongside, same era.
2022
Cited alongside, same era.
M. Żelaszczyk and J. Mańdziuk, “Audio-to-image cross-modal generation,” in IJCNN , 2022
2022
Later among the works it cites.
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello, “Wav2clip: Learning robust audio representations from clip,” in ICASSP , 2022
2022
Later among the works it cites.
K. Crowson, S. Biderman, D. Kornis, D. Stander, E. Hallahan, L. Castricato, and E. Raff, “Vqgan-clip: Open domain image generation and editing with natural language guidance,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVII . Springer, 2022, pp. 88–105
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in CVPR , 2023
2023
Closest in time.