Fetching the paper…
Reading the bibliography…
AI-driven design problems, such as DNA/protein sequence design, are commonly tackled from two angles: generative modeling, which efficiently captures the feasible design space (e.g., natural images or biological sequences), and model-based optimization, which utilizes reward models for extrapolation.
Behavior regularized offline reinforcement learning
Wu, Y., G. Tucker, and O. Nachum (2019) · 1911
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., A. Kumar, G. Tucker, and J. Fu (2020) · 2005
Earlier work this paper cites.
Deployment-efficient reinforcement learning via model-based offline optimization
Matsushima, T., H. Furuta, Y. Matsuo, O. Nachum, and S. Gu (2020) · 2006
Earlier work this paper cites.
Hyperparameter selection for offline reinforcement learning
Paine, T. L., C. Paduraru, A. Michi, C. Gulcehre, K. Zolna, A. Novikov, Z. Wang, and N. de Freitas (2020) · 2007
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., A. Krause, S. M. Kakade, and M. Seeger (2009) · 2009
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., C. Meng, and S. Ermon (2020) · 2010
Earlier work this paper cites.
Ava: A large-scale database for aesthetic visual analysis
Murray, N., L. Marchesotti, and F. Perronnin (2012) · 2012
Earlier work this paper cites.
Finite-time analysis of kernelised contextual bandits
Valko, M., N. Korda, R. Munos, I. Flaounas, and N. Cristianini (2013) · 2013
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., E. Weiss, N. Maheswaranathan, and S. Ganguli (2015) · 2015
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Neu, G., A. Jonsson, and V. Gómez (2017) · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., R. Calandra, R. McAllister, and S. Levine (2018) · 2018
Earlier work this paper cites.
Automatic chemical design using a data-driven continuous representation of molecules
Gómez-Bombarelli, R., J. N. Wei, D. Duvenaud, J. M. Hernández-Lobato, B. Sánchez-Lengeling, D. Sheberla, J. Aguilera-Iparraguirre, T. D. Hirzel, R. P. Adams, and A. Aspuru-Guzik (2018) · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S. (2018) · 2018
Earlier work this paper cites.
Model-based reinforcement learning for biological sequence design
Angermueller, C., D. Dohan, D. Belanger, R. Deshpande, K. Murphy, and L. Colwell (2019) · 2019
Earlier work this paper cites.
Conditioning by adaptive sampling for robust design
Brookes, D., H. Park, and J. Listgarten (2019) · 2019
Earlier work this paper cites.
A theory of regularized markov decision processes
Geist, M., B. Scherrer, and O. Pietquin (2019) · 2019
Earlier work this paper cites.
Identification and massively parallel characterization of regulatory elements driving neural induction
Inoue, F., A. Kreimer, T. Ashuach, N. Ahituv, and N. Yosef (2019) · 2019
Earlier work this paper cites.
Garbage in, reward out: Bootstrapping exploration in multi-armed bandits
Kveton, B., C. Szepesvari, S. Vaswani, Z. Wen, T. Lattimore, and M. Ghavamzadeh (2019) · 2019
Earlier work this paper cites.
Human 5 utr design and variant effect prediction from a massively parallel translation assay
Sample, P. J., B. Wang, D. W. Reid, V. Presnyak, I. J. McFadyen, D. R. Morris, and G. Seelig (2019) · 2019
Earlier work this paper cites.
High-dimensional statistics: A non-asymptotic viewpoint
Wainwright, M. J. (2019) · 2019
Earlier work this paper cites.
Autofocused oracles for model-based design
Fannjiang, C. and J. Listgarten (2020) · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., A. Jain, and P. Abbeel (2020) · 2020
Earlier work this paper cites.
Morel: Model-based offline reinforcement learning
Kidambi, R., A. Rajeswaran, P. Netrapalli, and T. Joachims (2020) · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., A. Zhou, G. Tucker, and S. Levine (2020) · 2020
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Yu, T., G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma (2020) · 2020
Earlier work this paper cites.
Effective gene expression prediction from sequence by integrating long-range interactions
Avsec, Ž., V. Agarwal, D. Visentin, J. R. Ledsam, A. Grabska-Barwinska, K. R. Taylor, Y. Assael, J. Jumper, P. Kohli, and D. R. Kelley (2021) · 2021
Cited alongside, same era.
Machine learning for designing next-generation mrna therapeutics
Castillo-Hair, S. M. and G. Seelig (2021) · 2021
Cited alongside, same era.
Mitigating covariate shift in imitation learning via offline data without great coverage
Chang, J. D., M. Uehara, D. Sreenivas, R. Kidambi, and W. Sun (2021) · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and A. Nichol (2021) · 2021
Cited alongside, same era.
Continuous doubly constrained batch reinforcement learning
Fakoor, R., J. W. Mueller, K. Asadi, P. Chaudhari, and A. J. Smola (2021) · 2021
Cited alongside, same era.
Generative flow networks assisted biological sequence editing
Ghari, P. M., A. Tseng, G. Eraslan, R. Lopez, T. Biancalani, G. Scalia, and E. Hajiramezanali (2023) · 2023
Later among the works it cites.
Machine-guided design of synthetic cell type-specific cis-regulatory elements
Gosai, S. J., R. I. Castro, N. Fuentes, J. C. Butts, S. Kales, R. R. Noche, K. Mouri, P. C. Sabeti, S. K. Reilly, and R. Tewhey (2023) · 2023
Later among the works it cites.
Diffusion models for black-box optimization
Krishnamoorthy, S., S. M. Mashkaria, and A. Grover (2023) · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Lee, K., H. Liu, M. Ryu, O. Watkins, Y. Du, C. Boutilier, P. Abbeel, M. Ghavamzadeh, and S. S. Gu (2023) · 2023
Later among the works it cites.
Aligning text-to-image diffusion models with reward backpropagation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast activation maximization for molecular sequence design
Linder, J. and G. Seelig (2021) · 2021
Cited alongside, same era.
Improving black-box optimization in vae latent space using decoder uncertainty
Notin, P., J. M. Hernández-Lobato, and Y. Gal (2021) · 2021
Cited alongside, same era.
Conservative objective models for effective offline model-based optimization
Trabucco, B., A. Kumar, X. Geng, and S. Levine (2021) · 2021
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and W. Sun (2021) · 2021
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal (2021) · 2021
Cited alongside, same era.
Path integral sampler: a stochastic control approach for sampling
Zhang, Q. and Y. Chen (2021) · 2021
Cited alongside, same era.
When does return-conditioned supervised learning work for offline reinforcement learning?
Brandfonbrener, D., A. Bietti, J. Buckman, R. Laroche, and J. Bruna (2022) · 2022
Cited alongside, same era.
Prabhudesai, M., A. Goyal, D. Pathak, and K. Fragkiadaki (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. (2023) · 2023
Later among the works it cites.
Vargas, F., W. Grathwohl, and A. Doucet (2023) · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization
Wallace, B., M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik (2023) · 2023
Later among the works it cites.
Better aligning text-to-image models with human preference
Wu, X., K. Sun, F. Zhu, R. Zhao, and H. Li (2023) · 2023
Later among the works it cites.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Xu, J., X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong (2023) · 2023
Later among the works it cites.
Using human feedback to fine-tune diffusion models without any reward model
Yang, K., J. Tao, J. Lyu, C. Ge, J. Chen, Q. Li, W. Shen, X. Zhu, and X. Li (2023) · 2023
Later among the works it cites.
Reward-directed conditional diffusion: Provable distribution estimation and reward improvement
Yuan, H., K. Huang, C. Ni, M. Chen, and M. Wang (2023) · 2023
Later among the works it cites.
Provable offline preference-based reinforcement learning
Zhan, W., M. Uehara, N. Kallus, J. D. Lee, and W. Sun (2023) · 2023
Later among the works it cites.
Campbell, A., J. Yim, R. Barzilay, T. Rainforth, and T. Jaakkola (2024) · 2024
Closest in time.
Dna-diffusion: Leveraging generative models for controlling chromatin accessibility and gene expression via synthetic regulatory elements
Ferreira DaSilva, L., S. Senan, Z. M. Patel, A. J. Reddy, S. Gabbita, Z. Nussbaum, C. M. V. Cordova, A. Wenteler, N. Weber, T. M. Tunjic, et al. (2024) · 2024
Closest in time.
Diffusion models as constrained samplers for optimization with unknown constraints
Kong, L., Y. Du, W. Mu, K. Neklyudov, V. De Bortol, H. Wang, D. Wu, A. Ferber, Y.-A. Ma, C. P. Gomes, et al. (2024) · 2024
Closest in time.
reglm: Designing realistic regulatory dna with autoregressive language models
Lal, A., D. Garfield, T. Biancalani, and G. Eraslan (2024) · 2024
Closest in time.
Discdiff: Latent diffusion model for dna sequence generation
Li, Z., Y. Ni, W. A. Beardall, G. Xia, A. Das, G.-B. Stan, and Y. Zhao (2024) · 2024
Closest in time.
Visual instruction tuning
Liu, H., C. Li, Q. Wu, and Y. J. Lee (2024) · 2024
Closest in time.
Designing dna with tunable regulatory activity using discrete diffusion
Sarkar, A., Z. Tang, C. Zhao, and P. Koo (2024) · 2024
Closest in time.
Dirichlet flow matching with applications to dna sequence design
Stark, H., B. Jing, C. Wang, G. Corso, B. Berger, R. Barzilay, and T. Jaakkola (2024) · 2024
Closest in time.
Fine-tuning of continuous-time diffusion models as entropy-regularized control
Uehara, M., Y. Zhao, K. Black, E. Hajiramezanali, G. Scalia, N. L. Diamant, A. M. Tseng, T. Biancalani, and S. Levine (2024) · 2024
Closest in time.
Protein structure generation via folding diffusion
Wu, K. E., K. K. Yang, R. van den Berg, S. Alamdari, J. Y. Zou, A. X. Lu, and A. P. Amini (2024) · 2024
Closest in time.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Xiong, W., H. Dong, C. Ye, Z. Wang, H. Zhong, H. Ji, N. Jiang, and T. Zhang (2023) · 2024
Closest in time.