Fetching the paper…
Reading the bibliography…
Image-generation diffusion models have been fine-tuned to unlock new capabilities such as image-editing and novel view synthesis.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Coppeliasim (formerly v-rep): a versatile and scalable robot simulation framework
E. Rohmer, S. P. N. Singh, and M. Freese · 2013
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
3d simulation for robot arm control with deep q-learning
S. James and E. Johns · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Transferring end-to-end visuomotor control from simulation to real world for a multi-stage task
S. James, A. J. Davison, and E. Johns · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Pyrep: Bringing v-rep to deep robot learning
S. James, M. Freese, and A. J. Davison · 2019
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Camera-to-robot pose estimation from a single image
T. E. Lee, J. Tremblay, T. To, J. Cheng, T. Mosier, O. Kroemer, D. Fox, and S. Birchfield · 2020
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
A. Birhane, V. U. Prabhu, and E. Kahembwe · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Earlier work this paper cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Earlier work this paper cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Earlier work this paper cites.
Diffusers: State-of-the-art diffusion models
P. von Platen, S. Patil, A. Lozhkov, P. Cuenca, N. Lambert, K. Rasul, M. Davaadorj, D. Nair, S. Paul, W. Berman, Y. Xu, S. Liu, and T. Wolf · 2022
Earlier work this paper cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Z. Wang, J. J. Hunt, and M. Zhou · 2022
Earlier work this paper cites.
Is conditional generative modeling all you need for decision-making?
A. Ajay, Y. Du, A. Gupta, J. Tenenbaum, T. Jaakkola, and P. Agrawal · 2022
Earlier work this paper cites.
Structdiffusion: Language-guided creation of physically-valid structures using unseen objects
W. Liu, Y. Du, T. Hermans, S. Chernova, and C. Paxton · 2022
Earlier work this paper cites.
Planning with diffusion for flexible behavior synthesis
M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Earlier work this paper cites.
Photorealistic video generation with diffusion models
A. Gupta, L. Yu, K. Sohn, X. Gu, M. Hahn, L. Fei-Fei, I. Essa, L. Jiang, and J. Lezama · 2023
Earlier work this paper cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2023
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions
T. Brooks, A. Holynski, and A. A. Efros · 2023
Cited alongside, same era.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Cited alongside, same era.
Emergent correspondence from image diffusion
L. Tang, M. Jia, Q. Wang, C. P. Phoo, and B. Hariharan · 2023
Cited alongside, same era.
Zero-1-to-3: Zero-shot one image to 3d object
R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Cited alongside, same era.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, et al · 2023
Later among the works it cites.
Robotap: Tracking arbitrary points for few-shot visual imitation
M. Vecerik, C. Doersch, Y. Yang, T. Davchev, Y. Aytar, G. Zhou, R. Hadsell, L. Agapito, and J. Scholz · 2023
Later among the works it cites.
Any-point trajectory modeling for policy learning
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel · 2023
Later among the works it cites.
C.-F. Yang, H. Xu, T.-L. Wu, X. Gao, K.-W. Chang, and F. Gao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Shi, P. Wang, J. Ye, M. Long, K. Li, and X. Yang · 2023
Cited alongside, same era.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine · 2023
Cited alongside, same era.
Y. Du, M. Yang, P. Florence, F. Xia, A. Wahid, B. Ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, et al · 2023
Cited alongside, same era.
Dall-e-bot: Introducing web-scale diffusion models to robotics
I. Kapelyukh, V. Vosylius, and E. Johns · 2023
Cited alongside, same era.
Genaug: Retargeting behaviors to unseen situations via generative augmentation
Z. Chen, S. Kiami, A. Gupta, and V. Kumar · 2023
Cited alongside, same era.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Cited alongside, same era.
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Cited alongside, same era.
V. Saxena, Y. Koga, and D. Xu · 2023
Later among the works it cites.
Training diffusion models with reinforcement learning
K. Black, M. Janner, Y. Du, I. Kostrikov, and S. Levine · 2023
Later among the works it cites.
Streamdiffusion: A pipeline-level solution for real-time interactive generation
A. Kodaira, C. Xu, T. Hazama, T. Yoshimoto, K. Ohno, S. Mitsuhori, S. Sugano, H. Cho, Z. Liu, and K. Keutzer · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik · 2023
Later among the works it cites.
Stable bias: Analyzing societal representations in diffusion models
A. S. Luccioni, C. Akiki, M. Mitchell, and Y. Jernite · 2023
Later among the works it cites.
Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models
P. Schramowski, M. Brack, B. Deiseroth, and K. Kersting · 2023
Later among the works it cites.
Unsupervised semantic correspondence using stable diffusion
E. Hedlin, G. Sharma, S. Mahajan, H. Isack, A. Kar, A. Tagliasacchi, and K. M. Yi · 2024
Closest in time.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel · 2024
Closest in time.
Dnact: Diffusion guided multi-task 3d policy learning
G. Yan, Y.-H. Wu, and X. Wang · 2024
Closest in time.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan · 2024
Closest in time.
Pedipulate: Enabling manipulation skills using a quadruped robot’s leg
P. Arm, M. Mittal, H. Kolvenbach, and M. Hutter · 2024
Closest in time.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Closest in time.
The colosseum: A benchmark for evaluating generalization for robotic manipulation
W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox · 2024
Closest in time.
Render and diffuse: Aligning image and action spaces for diffusion-based behaviour cloning, 2024
V. Vosylius, Y. Seo, J. Uruç, and S. James · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Closest in time.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Closest in time.
Hierarchical diffusion policy for kinematics-aware multi-task robotic manipulation
X. Ma, S. Patidar, I. Haughton, and S. James · 2024
Closest in time.
Enabling stateful behaviors for diffusion-based policy learning
X. Liu, F. Weigend, Y. Zhou, and H. B. Amor · 2024
Closest in time.
Diffusion meets dagger: Supercharging eye-in-hand imitation learning
X. Zhang, M. Chang, P. Kumar, and S. Gupta · 2024
Closest in time.
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Closest in time.
Simple hierarchical planning with diffusion
C. Chen, F. Deng, K. Kawaguchi, C. Gulcehre, and S. Ahn · 2024
Closest in time.
Extracting reward functions from diffusion models
F. Nuti, T. Franzmeyer, and J. F. Henriques · 2024
Closest in time.
Aldm-grasping: Diffusion-aided zero-shot sim-to-real transfer for robot grasping
Y. Li, Z. Wu, H. Zhao, T. Yang, Z. Liu, P. Shu, J. Sun, R. Parasuraman, and T. Liu · 2024
Closest in time.
Large-scale actionless video pre-training via discrete diffusion for efficient policy learning
H. He, C. Bai, L. Pan, W. Zhang, B. Zhao, and X. Li · 2024
Closest in time.
Compositional foundation models for hierarchical planning
A. Ajay, S. Han, Y. Du, S. Li, A. Gupta, T. Jaakkola, J. Tenenbaum, L. Kaelbling, A. Srivastava, and P. Agrawal · 2024
Closest in time.
Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches, 2024
P. Sundaresan, Q. Vuong, J. Gu, P. Xu, T. Xiao, S. Kirmani, T. Yu, M. Stark, A. Jain, K. Hausman, D. Sadigh, J. Bohg, and S. Schaal · 2024
Closest in time.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, et al · 2024
Closest in time.
Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation
H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani · 2024
Closest in time.
Depth anything: Unleashing the power of large-scale unlabeled data
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao · 2024
Closest in time.