Fetching the paper…
Reading the bibliography…
Diffusion models have demonstrated their capabilities in modeling trajectories of multi-tasks.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A., and Munos, R · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Earlier work this paper cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2018
Earlier work this paper cites.
Open3D: A modern library for 3D data processing
Zhou, Q.-Y., Park, J., and Koltun, V · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L · 2019
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Reward-free exploration for reinforcement learning
Jin, C., Krishnamurthy, A., Simchowitz, M., and Yu, T · 2020
Earlier work this paper cites.
On reward-free reinforcement learning with linear function approximation
Wang, R., Du, S. S., Yang, L., and Salakhutdinov, R. R · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C · 2020
Earlier work this paper cites.
Diffwave: A versatile diffusion model for audio synthesis
Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B · 2021
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2021
Earlier work this paper cites.
Implicit behavioral cloning
Florence, P., Lynch, C., Zeng, A., Ramirez, O. A., Wahid, A., Downs, L., Wong, A., Lee, J., Mordatch, I., and Tompson, J · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Earlier work this paper cites.
Torsional diffusion for molecular conformer generation
Jing, B., Corso, G., Chang, J., Barzilay, R., and Jaakkola, T · 2022
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2022
Cited alongside, same era.
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
A generalist agent
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-maron, G., Giménez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Goal conditioned imitation learning using score-based diffusion policies
Reuss, M., Li, M., Jia, X., and Lioutikov, R · 2023
Later among the works it cites.
Masked world models for visual control
Seo, Y., Hafner, D., Liu, H., Liu, F., James, S., Lee, K., and Abbeel, P · 2023
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2023
Later among the works it cites.
Xskill: Cross embodiment skill discovery
Xu, M., Xu, Z., Chi, C., Veloso, M., and Song, S · 2023
Later among the works it cites.
Training diffusion models with reinforcement learning
Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S · 2024
Closest in time.
Edgi: Equivariant diffusion for planning with embodied agents
Brehmer, J., Bose, J., De Haan, P., and Cohen, T. S · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Paco: Parameter-compositional multi-task reinforcement learning
Sun, L., Zhang, H., Xu, W., and Tomizuka, M · 2022
Cited alongside, same era.
Vrl3: A data-driven framework for visual deep reinforcement learning
Wang, C., Luo, X., Ross, K., and Li, D · 2022
Cited alongside, same era.
Prompting decision transformer for few-shot policy generalization
Xu, M., Shen, Y., Zhang, S., Lu, Y., Zhao, D., Tenenbaum, J., and Gan, C · 2022
Cited alongside, same era.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A · 2022
Cited alongside, same era.
Is conditional generative modeling all you need for decision making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J. B., Jaakkola, T. S., and Agrawal, P · 2023
Cited alongside, same era.
Reinforcement learning for fine-tuning text-to-speech diffusion models
Chen, J., Byun, J.-S., Elsner, M., and Perrault, A · 2024
Closest in time.
Directly fine-tuning diffusion models on differentiable rewards
Clark, K., Vicol, P., Swersky, K., and Fleet, D. J · 2024
Closest in time.
Reinforcement learning for fine-tuning text-to-image diffusion models
Fan, Y., Watkins, O., Du, Y., Liu, H., Ryu, M., Boutilier, C., Abbeel, P., Ghavamzadeh, M., Lee, K., and Lee, K · 2024
Closest in time.
Learning an actionable discrete diffusion policy via large-scale actionless video pre-training
He, H., Bai, C., Pan, L., Zhang, W., Zhao, B., and Li, X · 2024
Closest in time.
Harmodt: Harmony multi-task decision transformer for offline reinforcement learning
Hu, S., Fan, Z., Shen, L., Zhang, Y., Wang, Y., and Tao, D · 2024
Closest in time.
Efficient diffusion policies for offline reinforcement learning
Kang, B., Ma, X., Du, C., Pang, T., and Yan, S · 2024
Closest in time.
Crossway diffusion: Improving diffusion-based visuomotor policy via self-supervised learning
Li, X., Belagali, V., Shang, J., and Ryoo, M. S · 2024
Closest in time.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Nakamoto, M., Zhai, S., Singh, A., Sobol Mark, M., Ma, Y., Finn, C., Kumar, A., and Levine, S · 2024
Closest in time.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2024
Closest in time.
Structure-based drug design with equivariant diffusion models
Schneuing, A., Harris, C., Du, Y., Didi, K., Jamasb, A., Igashov, I., Du, W., Gomes, C., Blundell, T. L., Lio, P., et al · 2024
Closest in time.
Finetuning text-to-image diffusion models for fairness
Shen, X., Du, C., Pang, T., Lin, M., Wong, Y., and Kankanhalli, M · 2024
Closest in time.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., and Dong, Y · 2024
Closest in time.
Regularized conditional diffusion model for multi-task preference alignment
Yu, X., Bai, C., He, H., Wang, C., and Li, X · 2024
Closest in time.
Preference aligned diffusion planner for quadrupedal locomotion control
Yuan, X., Shang, Z., Wang, Z., Wang, C., Shan, Z., Zhu, M., Bai, C., Li, X., Wan, W., and Harada, K · 2024
Closest in time.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Ze, Y., Zhang, G., Zhang, K., Hu, C., Wang, M., and Xu, H · 2024
Closest in time.
Online preference alignment for language models via count-based exploration
Bai, C., Zhang, Y., Qiu, S., Zhang, Q., Xu, K., and Li, X · 2025
Closest in time.
Diffusion policy policy optimization
Ren, A. Z., Lidard, J., Ankile, L. L., Simeonov, A., Agrawal, P., Majumdar, A., Burchfiel, B., Dai, H., and Simchowitz, M · 2025
Closest in time.