Fetching the paper…
Reading the bibliography…
In this paper, we explore audio-editing with non-rigid text edits.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention (MICCAI) , N. Navab, J. Hornegger, W. M. Wells, and A. F. Frangi, Eds., 2015, pp. 234–241
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of GANs for improved quality, stability, and variation,” in International Conference on Learning Representations (ICLR) , 2018
2018
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis,” in International Conference on Neural Information Processing Systems (NeurIPS) , 2020
2020
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International Conference on Machine Learning (ICML) , vol. 139, 2021, pp. 8821–8831
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 10 674–10 685
2021
Earlier work this paper cites.
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon, “SDEdit: Guided image synthesis and editing with stochastic differential equations,” in International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
T. Brooks, A. Holynski, and A. A. Efros, “InstructPix2Pix: Learning to follow image editing instructions,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 18 392–18 402
2022
Earlier work this paper cites.
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations (ICLR) , 2022
2022
Earlier work this paper cites.
L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M.-H. Yang, “Diffusion models: A comprehensive survey of methods and applications,” ACM Computing Surveys , vol. 56, no. 4, pp. 1–39, 2023
2023
Cited alongside, same era.
G. Couairon, J. Verbeek, H. Schwenk, and M. Cord, “DiffEdit: Diffusion-based semantic image editing with mask guidance,” in International Conference on Learning Representations (ICLR) , 2023
2023
Cited alongside, same era.
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-An-Audio: Text-to-audio generation with prompt-enhanced diffusion models,” in International Conference on Machine Learning (ICML) , 2023, pp. 13 916–13 932
2023
Cited alongside, same era.
2023
Y. Wang, Z. Ju, X. Tan, L. He, Z. Wu, J. Bian, and sheng zhao, “AUDIT: Audio editing by following instructions with latent diffusion models,” in International Conference on Neural Information Processing Systems (NeurIPS) , 2023
2023
Closest in time.
B. Kawar, S. Zada, O. Lang, O. Tov, H. Chang, T. Dekel, I. Mosseri, and M. Irani, “Imagic: Text-based real image editing with diffusion models,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023, pp. 6007–6017
2023
Closest in time.
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, and A. Roberts, “The Flan Collection: Designing data and methods for effective instruction tuning,” in International Conference on Machine Learning (ICML) , vol. 202, 2023, pp. 22 631–22 648
2023
Closest in time.
N. Ding, Y. Qin, G. Yang, F. Wei, Z. Yang, Y. Su, S. Hu, Y. Chen, C.-M. Chan, W. Chen et al. , “Parameter-efficient fine-tuning of large-scale pre-trained language models,” Nature Machine Intelligence , vol. 5, no. 3, pp. 220–235, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” International Conference on Machine Learning (ICML) , 2023
2023
Cited alongside, same era.
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “CLAP: learning audio concepts from natural language supervision,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Cited alongside, same era.
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “AudioGen: Textually guided audio generation,” in International Conference on Learning Representations (ICLR) , 2023
2023
Cited alongside, same era.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” in International Conference on Neural Information Processing Systems (NeurIPS) , 2023
2023
Cited alongside, same era.
D. Ghosal, N. Majumder, A. Mehrish, and S. Poria, “Text-to-audio generation using instruction guided latent diffusion model,” in International Conference on Multimedia , 2023, pp. 3590––3598
2023
Cited alongside, same era.
2023
Closest in time.
Z. Fu, H. Yang, A. M.-C. So, W. Lam, L. Bing, and N. Collier, “On the effectiveness of parameter-efficient fine-tuning,” in AAAI Conference on Artificial Intelligence , vol. 37, no. 11, 2023, pp. 12 799–12 807
2023
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in International Conference on Neural Information Processing Systems (NeurIPS) , 2023
2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in International Conference on Neural Information Processing Systems (NeurIPS) , vol. 36, 2024
2024
Closest in time.
Y. Li, H. Wang, Q. Jin, J. Hu, P. Chemerys, Y. Fu, Y. Wang, S. Tulyakov, and J. Ren, “SnapFusion: Text-to-image diffusion model on mobile devices within two seconds,” in International Conference on Neural Information Processing Systems (NeurIPS) , vol. 36, 2024
2024
Closest in time.
Y. Li, Y. Yu, C. Liang, N. Karampatziakis, P. He, W. Chen, and T. Zhao, “LoftQ: LoRA-fine-tuning-aware quantization for large language models,” in International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.