Fetching the paper…
Reading the bibliography…
We propose Diffusion Inference-Time T-Optimization (DITTO), a general-purpose frame-work for controlling pre-trained text-to-music diffusion models at inference-time via optimizing initial noise latents.
A scale for the measurement of the psychological magnitude pitch
Stevens, S. S., Volkmann, J., and Newman, E. B · 1937
Earlier work this paper cites.
Sound structure in music
Erickson, R · 1975
Earlier work this paper cites.
Learning similarity from collaborative filters
McFee, B., Barrington, L., and Lanckriet, G. R · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
NICE: Non-linear independent components estimation
Dinh, L., Krueger, D., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Analyzing song structure with spectral clustering
McFee, B. and Ellis, D · 2014
Earlier work this paper cites.
Fundamentals of music processing: Audio, analysis, algorithms, applications
Müller, M · 2015
Earlier work this paper cites.
U-Net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Density estimation using real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
FMA: A dataset for music analysis
Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X · 2017
Earlier work this paper cites.
MuseGAN: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment
Dong, H.-W., Hsiao, W.-Y., Yang, L.-C., and Yang, Y.-H · 2018
Earlier work this paper cites.
Frechet audio distance: A metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2018
Earlier work this paper cites.
Symbolic music similarity through a graph-based representation
Simonetta, F., Carnovalini, F., Orio, N., and Rodà, A · 2018
Earlier work this paper cites.
Memcnn: A python/pytorch package for creating memory-efficient invertible neural networks
Leemput, S. C. v., Teuwen, J., Ginneken, B. v., and Manniesing, R · 2019
Earlier work this paper cites.
Music SketchNet: Controllable music generation via factorized representations of pitch and rhythm
Chen, K., Wang, C.-i., Berg-Kirkpatrick, T., and Dubnov, S · 2020
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Musichifi: Fast high-fidelity stereo vocoding
Zhu, G., Caceres, J.-P., Duan, Z., and Bryan, N. J · 2020
Earlier work this paper cites.
Controllable deep melody generation via hierarchical music structure representation
Dai, S., Jin, Z., Gomes, C., and Dannenberg, R · 2021
Cited alongside, same era.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Cited alongside, same era.
Gan inversion: A survey
Xia, W., Zhang, Y., Yang, Y., Xue, J.-H., Zhou, B., and Yang, M.-H · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Cited alongside, same era.
Riffusion: Stable diffusion for real-time music generation, 2022
Forsgren, S. and Martiros, H · 2022
Cited alongside, same era.
Custom-Edit: Text-guided image editing with customized diffusion models
Choi, J., Choi, Y., Kim, Y., Kim, J., and Yoon, S · 2023
Later among the works it cites.
Diffusion posterior sampling for general noisy inverse problems
Chung, H., Kim, J., McCann, M. T., Klasky, M. L., and Ye, J. C · 2023
Later among the works it cites.
Directly fine-tuning diffusion models on differentiable rewards
Clark, K., Vicol, P., Swersky, K., and Fleet, D. J · 2023
Later among the works it cites.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2023
Later among the works it cites.
VampNet: Music generation via masked acoustic token modeling
Garcia, H. F., Seetharaman, P., Kumar, R., and Pardo, B · 2023
Later among the works it cites.
Adapting frechet audio distance for generative music evaluation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An image is worth one word: Personalizing text-to-image generation using textual inversion, 2022
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Cited alongside, same era.
Multi-instrument music synthesis with spectrogram diffusion
Hawthorne, C., Simon, I., Roberts, A., Zeghidour, N., Gardner, J., Manilow, E., and Engel, J · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Cited alongside, same era.
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Bigvgan: A universal neural vocoder with large-scale training
Lee, S.-g., Ping, W., Ginsburg, B., Catanzaro, B., and Yoon, S · 2022
Cited alongside, same era.
Gui, A., Gamper, H., Braun, S., and Emmanouilidou, D · 2023
Later among the works it cites.
Photorealistic video generation with diffusion models
Gupta, A., Yu, L., Sohn, K., Gu, X., Hahn, M., Li, F.-F., Essa, I., Jiang, L., and Lezama, J · 2023
Later among the works it cites.
Optimizing diffusion noise can serve as universal motion priors
Karunratanakul, K., Preechakul, K., Aksan, E., Beeler, T., Suwajanakorn, S., and Tang, S · 2023
Later among the works it cites.
Imagic: Text-based real image editing with diffusion models
Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H., Dekel, T., Mosseri, I., and Irani, M · 2023
Later among the works it cites.
Consistency trajectory models: Learning probability flow ode trajectory of diffusion
Kim, D., Lai, C.-H., Liao, W.-H., Murata, N., Takida, Y., Uesaka, T., He, Y., Mitsufuji, Y., and Ermon, S · 2023
Later among the works it cites.
High-fidelity audio compression with improved RVQGAN
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2023
Later among the works it cites.
Controllable music production with diffusion models and guidance gradients
Levy, M., Giorgi, B. D., Weers, F., Katharopoulos, A., and Nickson, T · 2023
Later among the works it cites.
Latent consistency models: Synthesizing high-resolution images with few-step inference
Luo, S., Tan, Y., Huang, L., Li, J., and Zhao, H · 2023
Later among the works it cites.
Null-text inversion for editing real images using guided diffusion models
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D · 2023
Later among the works it cites.
Effective real image editing with accelerated iterative diffusion inversion
Pan, Z., Gherardi, R., Xie, X., and Huang, S · 2023
Later among the works it cites.
Aligning text-to-image diffusion models with reward backpropagation
Prabhudesai, M., Goyal, A., Pathak, D., and Fragkiadaki, K · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2023
Later among the works it cites.
Mo \ \backslash ˆ usai: Text-to-music generation with long-context latent diffusion
Schneider, F., Jin, Z., and Schölkopf, B · 2023
Later among the works it cites.
Freeu: Free lunch in diffusion u-net
Si, C., Huang, Z., Jiang, Y., and Liu, Z · 2023
Later among the works it cites.
Freedom: Training-free energy-guided conditional diffusion model
Yu, J., Wang, Y., Zhao, C., Ghanem, B., and Zhang, J · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.
Uni-ControlNet: All-in-one control to text-to-image diffusion models
Zhao, S., Chen, D., Chen, Y.-C., Bao, J., Hao, S., Yuan, L., and Wong, K.-Y. K · 2023
Later among the works it cites.