Fetching the paper…
Reading the bibliography…
Generative models are transforming creative domains such as music generation, with inference-time strategies like Classifier-Free Guidance (CFG) playing a crucial role.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Weight averaging for neural networks and local resampling schemes
J. Utans · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Managing diversity in regression ensembles
G. Brown, J. L. Wyatt, and P. Tiňo · 2005
Earlier work this paper cites.
Bandit based monte-carlo planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Evolving a diversity of virtual creatures through novelty search and local competition
J. Lehman and K. O. Stanley · 2011
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
Robots that can adapt like animals
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Illuminating search spaces by mapping elites
J.-B. Mouret and J. Clune · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Beam search strategies for neural machine translation
M. Freitag and Y. Al-Onaizan · 2017
Earlier work this paper cites.
Objective-reinforced generative adversarial networks (organ) for sequence generation models
G. L. Guimaraes, B. Sanchez-Lengeling, C. Outeiral, P. L. C. Farias, and A. Aspuru-Guzik · 2017
Earlier work this paper cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
N. Jaques, S. Gu, D. Bahdanau, J. M. Hernández-Lobato, R. E. Turner, and D. Eck · 2017
Earlier work this paper cites.
Iterated Distillation and Amplification, 2018
A. Cotra · 2018
Earlier work this paper cites.
Hierarchical neural story generation
A. Fan, M. Lewis, and Y. Dauphin · 2018
Earlier work this paper cites.
Bach2bach: generating music using a deep reinforcement learning approach
N. Kotecha · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton · 2018
Earlier work this paper cites.
Generating informative and diverse conversational responses via adversarial information maximization
Y. Zhang, M. Galley, J. Gao, Z. Gan, X. Li, C. Brockett, and B. Dolan · 2018
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Earlier work this paper cites.
Qd-rl: Efficient mixing of quality and diversity in reinforcement learning
G. Cideron, T. Pierrot, N. Perrin, K. Beguir, and O. Sigaud · 2020
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis
J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Earlier work this paper cites.
Autoregressive knowledge distillation through imitation learning
A. Lin, J. Wohlwend, H. Chen, and T. Lei · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Cited alongside, same era.
Better aggregation in test-time augmentation
D. Shanmugam, D. Blalock, G. Balakrishnan, and J. Guttag · 2021
Cited alongside, same era.
Trading off diversity and quality in natural language generation
H. Zhang, D. Duckworth, D. Ippolito, and A. Neelakantan · 2021
Cited alongside, same era.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Cited alongside, same era.
Make-a-scene: Scene-based text-to-image generation with human priors
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Cited alongside, same era.
A survey on deep reinforcement learning for audio-based applications
S. Latif, H. Cuayáhuitl, F. Pervez, F. Shamshad, H. S. Ali, and E. Cambria · 2023
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, Z. Tu, and S. Shi · 2023
Later among the works it cites.
Audioldm: Text-to-audio generation with latent diffusion models
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. P. Mandic, W. Wang, and M. D. Plumbley · 2023
Later among the works it cites.
On distillation of guided diffusion models
C. Meng, R. Rombach, R. Gao, D. Kingma, S. Ermon, J. Ho, and T. Salimans · 2023
Later among the works it cites.
Tweet on the examples of negative prompting, 2023
E. Mostaque · 2023
Later among the works it cites.
Task arithmetic in the tangent space: Improved editing of pre-trained models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Cited alongside, same era.
MuLan: A joint embedding of music audio and natural language
Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. W. Ellis · 2022
Cited alongside, same era.
Reddit thread titled What is the point of the endless model merges?, 2022
Purplekeyboard · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
Diverse weight averaging for out-of-distribution generalization
A. Ramé, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, P. Gallinari, and M. Cord · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
G. Ortiz-Jimenez, A. Favero, and P. Frossard · 2023
Later among the works it cites.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
A. Ramé, G. Couairon, M. Shukor, C. Dancette, J.-B. Gaya, L. Soulier, and M. Cord · 2023
Later among the works it cites.
Stay on topic with classifier-free guidance
G. Sanchez, H. Fan, A. Spangher, E. Levi, P. S. Ammanamanchi, and S. Biderman · 2023
Later among the works it cites.
Moûsai: Text-to-music generation with long-context latent diffusion
F. Schneider, O. Kamal, Z. Jin, and B. Schölkopf · 2023
Later among the works it cites.
Scott: Self-consistent chain-of-thought distillation
P. Wang, Z. Wang, Z. Li, Y. Gao, B. Yin, and X. Ren · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem · 2024
Closest in time.
Back to basics: Revisiting REINFORCE style optimization for learning from human feedback in LLMs
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, A. Üstün, and S. Hooker · 2024
Closest in time.
Diffusion soup: Model merging for text-to-image diffusion models
B. Biggs, A. Seshadri, Y. Zou, A. Jain, A. Golatkar, Y. Xie, A. Achille, A. Swaminathan, and S. Soatto · 2024
Closest in time.
Rlhf deciphered: A critical analysis of reinforcement learning from human feedback for llms
S. Chaudhari, P. Aggarwal, V. Murahari, T. Rajpurohit, A. Kalyan, K. Narasimhan, A. Deshpande, and B. C. da Silva · 2024
Closest in time.
MusicRL: Aligning music generation to human preferences
G. Cideron, S. Girgin, M. Verzetti, D. Vincent, M. Kastelic, Z. Borsos, B. McWilliams, V. Ungureanu, O. Bachem, O. Pietquin, M. Geist, L. Hussenot, N. Zeghidour, and A. Agostinelli · 2024
Closest in time.
Quality diversity through human feedback: Towards open-ended diversity-driven optimization
L. Ding, J. Zhang, J. Clune, L. Spector, and J. Lehman · 2024
Closest in time.
Proving linear mode connectivity of neural networks via optimal transport
D. Ferbach, B. Goujaud, G. Gidel, and A. Dieuleveut · 2024
Closest in time.
Mad speech: Measures of acoustic diversity of speech
M. Futeral, A. Agostinelli, M. Tagliasacchi, N. Zeghidour, and E. Kharitonov · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al · 2024
Closest in time.
Detecting mode collapse in language models via narration
S. Hamilton · 2024
Closest in time.
Understanding the effects of RLHF on LLM generalisation and diversity
R. Kirk, I. Mediratta, C. Nalmpantis, J. Luketina, E. Hambro, E. Grefenstette, and R. Raileanu · 2024
Closest in time.
From distributional to overton pluralism: Investigating large language model alignment
T. Lake, E. Choi, and G. Durrett · 2024
Closest in time.
A guide to similarity measures
A. Levy, B. R. Shalom, and M. Chalamish · 2024
Closest in time.
Spurious feature diversification improves out-of-distribution generalization
Y. Lin, L. Tan, Y. Hao, H. Wong, H. Dong, W. Zhang, Y. Yang, and T. Zhang · 2024
Closest in time.
Creativity has left the chat: The price of debiasing language models
B. Mohammadi · 2024
Closest in time.
WARM: On the benefits of weight averaged reward models
A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret · 2024
Closest in time.
BOND: Aligning llms with best-of- n n distillation
P. G. Sessa, R. Dadashi, L. Hussenot, J. Ferret, N. Vieillard, A. Ramé, B. Shariari, S. Perrin, A. Friesen, G. Cideron, et al · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Closest in time.
Generalized people diversity: Learning a human perception-aligned diversity representation for people images
H. Srinivasan, C. Schumann, A. Sinha, D. Madras, G. O. Olanubi, A. Beutel, S. Ricco, and J. Chen · 2024
Closest in time.
Conditioned language policy: A general framework for steerable multi-objective finetuning
K. Wang, R. Kidambi, R. Sullivan, A. Agarwal, C. Dann, A. Michi, M. Gelmi, Y. Li, R. Gupta, A. Dubey, et al · 2024
Closest in time.