Fetching the paper…
Reading the bibliography…
Editing signals using large pre-trained models, in a zero-shot manner, has recently seen rapid advancements in the image domain.
Numerical methods for large eigenvalue problems: revised edition
Saad, Y · 2011
Earlier work this paper cites.
MedleyDB: A multitrack dataset for annotation-intensive mir research
Bittner, R., Salamon, J., Tierney, M., Mauch, M., Cannam, C., and Bello, J · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
FMA: A dataset for music analysis
Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., et al · 2017
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2019
Earlier work this paper cites.
VGGSound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Earlier work this paper cites.
Interpreting the latent space of gans for semantic face editing
Shen, Y., Gu, J., Tang, X., and Zhou, B · 2020
Earlier work this paper cites.
GAN ”steerability” without optimization
Spingarn, N., Banner, R., and Michaeli, T · 2020
Earlier work this paper cites.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Earlier work this paper cites.
Taming visually guided sound generation
Iashin, V. and Rahtu, E · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Earlier work this paper cites.
SDEdit: Guided image synthesis and editing with stochastic differential equations
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
Closed-form factorization of latent semantics in GANs
Shen, Y. and Zhou, B · 2021
Cited alongside, same era.
StyleSpace analysis: Disentangled controls for stylegan image generation
Wu, Z., Lischinski, D., and Shechtman, E · 2021
Cited alongside, same era.
SINE: Single image editing with text-to-image diffusion models
Zhang, Z., Han, L., Ghosh, A., Metaxas, D. N., and Ren, J · 2021
Cited alongside, same era.
HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection
Chen, K., Du, X., Zhu, B., Ma, Z., Berg-Kirkpatrick, T., and Dubnov, S · 2022
Imagic: Text-based real image editing with diffusion models
Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H., Dekel, T., Mosseri, I., and Irani, M · 2023
Later among the works it cites.
Null-text inversion for editing real images using guided diffusion models
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D · 2023
Later among the works it cites.
Uncertainty quantification via neural posterior principal components
Nehme, E., Yair, O., and Michaeli, T · 2023
Later among the works it cites.
Audio editing with non-rigid text prompts
Paissan, F., Wang, Z., Ravanelli, M., Smaragdis, P., and Subakan, C · 2023
Later among the works it cites.
Unsupervised discovery of semantic latent directions in diffusion models
Park, Y.-H., Kwon, M., Jo, J., and Uh, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-or, D · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross-attention control
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-or, D · 2022
Cited alongside, same era.
DiffusionCLIP: Text-guided diffusion models for robust image manipulation
Kim, G., Kwon, T., and Ye, J. C · 2022
Cited alongside, same era.
Diffusion models already have a semantic latent space
Kwon, M., Jeong, J., and Uh, Y · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
MusicLM: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Cited alongside, same era.
DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2023
Later among the works it cites.
Plug-and-play diffusion features for text-driven image-to-image translation
Tumanyan, N., Geyer, M., Bagon, S., and Dekel, T · 2023
Later among the works it cites.
AUDIT: Audio editing by following instructions with latent diffusion models
Wang, Y., Ju, Z., Tan, X., He, L., Wu, Z., Bian, J., and sheng zhao · 2023
Later among the works it cites.
A latent space of stochastic diffusion models for zero-shot image editing and guidance
Wu, C. H. and De la Torre, F · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Later among the works it cites.
Adapting Frechet audio distance for generative music evaluation
Gui, A., Gamper, H., Braun, S., and Emmanouilidou, D · 2024
Closest in time.
An edit friendly DDPM noise space: Inversion and manipulations
Huberman-Spiegelglas, I., Kulikov, V., and Michaeli, T · 2024
Closest in time.
Training-free content injection using h-space in diffusion models
Jeong, J., Kwon, M., and Uh, Y · 2024
Closest in time.
On the posterior distribution in denoising: Application to uncertainty quantification
Manor, H. and Michaeli, T · 2024
Closest in time.
DITTO: Diffusion inference-time t-optimization for music generation
Novack, Z., McAuley, J., Berg-Kirkpatrick, T., and Bryan, N. J · 2024
Closest in time.
Investigating personalization methods in text to music generation
Plitsis, M., Kouzelis, T., Paraskevopoulos, G., Katsouros, V., and Panagakis, Y · 2024
Closest in time.
MusicMagus: Zero-shot text-to-music editing via diffusion models
Zhang, Y., Ikemiya, Y., Xia, G., Murata, N., Martínez, M., Liao, W.-H., Mitsufuji, Y., and Dixon, S · 2024
Closest in time.