Fetching the paper…
Reading the bibliography…
We present PANDORA, a novel diffusion-based policy learning framework designed specifically for dexterous robotic piano performance.
Reinforcement learning with perturbed rewards
Wang Jingkang, Liu Yang, and Li Bo · 2018
Earlier work this paper cites.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel Van de Panne · 2018
Earlier work this paper cites.
Deep reinforcement learning for industrial insertion tasks with visual inputs and natural rewards
Schoettler Gerrit, Nair Ashvin, Luo Jianlan, Bahl Shikhar, Ojea Juan, Aparicio, Solowjow Eugen, and Levine Sergey · 2019
Earlier work this paper cites.
Mediapipe: A framework for building perception pipelines
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Learning robotic manipulation tasks via task progress based gaussian reward and loss adjusted exploration, 2021
Sulabh Kumra, Shirin Josh, and Ferat Sahin · 2021
Earlier work this paper cites.
Language control diffusion: Efficiently scaling through space, time, and tasks
Zhang Edwin, Lu Yujie, Wang William, and Zhang Amy · 2022
Earlier work this paper cites.
Compositional visual generation with composable diffusion models
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum · 2022
Earlier work this paper cites.
Learning reward functions for robotic manipulation by observing humans
Alakuijala Minttu, Dulac-Arnold Gabriel, Mairal Julien, and and Cordelia Schmid Jean, Ponce · 2022
Earlier work this paper cites.
Denoising diffusion implicit models, 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2022
Earlier work this paper cites.
Towards learning to play piano with dexterous hands and touch
Huazhe Xu, Yuping Luo, Shaoxiong Wang, Trevor Darrell, and Roberto Calandra · 2022
Earlier work this paper cites.
Training diffusion models with reinforcement learning
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine · 2023
Earlier work this paper cites.
Zero-shot robotic manipulation with pretrained image-editing diffusion models, 2023
Kevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Walke, Chelsea Finn, Aviral Kumar, and Sergey Levine · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song · 2023
Earlier work this paper cites.
Vision-language models are zero-shot reward models for reinforcement learning
Rocamonde Juan, Montesinos Victoriano, Nava Elvis, Perez Ethan, and Lindner David · 2023
Earlier work this paper cites.
Edmp: Ensemble-of-costs-guided diffusion for motion planning
Saha Kallol, Mandadi Vishal, Reddy Jayaram, Srikanth Ajit, Agarwal Aditya, Sen Bipasha, and and Madhava Krishna Arun, Singh · 2023
Cited alongside, same era.
Diffusion reward: Learning rewards via conditional video diffusion
Huang Tao, Jiang Guangqi, Ze Yanjie, and Xu Huazhe · 2023
Cited alongside, same era.
Manipulate by seeing: Creating manipulation controllers from pre-trained representations, 2023
Jianren Wang, Sudeep Dasari, Mohan Kumar Srirama, Shubham Tulsiani, and Abhinav Gupta · 2023
Cited alongside, same era.
Eureka: Human-level reward design via coding large language models
Ma Yecheng, Jason, Liang William, Wang Guanzhi, Huang De-An, Bastani Osbert, Jayaraman Dinesh, Zhu Yuke, Fan Linxi, and Anandkumar Anima · 2023
Cited alongside, same era.
Robopianist: Dexterous piano playing with deep reinforcement learning
Switch ema: A free lunch for better flatness and sharpness
Li Siyuan, Liu Zicheng, Tian Juanxi, Wang Ge, Wang Zedong, Jin Weiyang, Wu Di, Tan Cheng, Lin Tao, Liu Yang, Sun Baigui, and Z. Li and, Stan · 2024
Later among the works it cites.
Diff-dagger: Uncertainty estimation with diffusion policy for robotic manipulation
Lee Sung-Wook and Kuo Yen-Ling · 2024
Later among the works it cites.
Ldp: A local diffusion planner for efficient robot navigation and collision avoidance
Yu Wenhao, Peng Jie, Yang Huanyu, Zhang Junrui, Duan Yifan, and and Yanyong Zhang Jianmin, Ji · 2024
Later among the works it cites.
Dnact: Diffusion guided multi-task 3d policy learning
Ge Yan, Yueh-Hua Wu, and Xiaolong Wang · 2024
Later among the works it cites.
Depth anything v2, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Zakka, Philipp Wu, Laura Smith, Nimrod Gileadi, Taylor Howell, Xue Bin Peng, Sumeet Singh, Yuval Tassa, Pete Florence, Andy Zeng, et al · 2023
Cited alongside, same era.
Diffusion policy policy optimization
Ren Allen, Z., Lidard Justin, Ankile Lars, L., Simeonov Anthony, Agrawal Pulkit, Majumdar Anirudha, Burchfiel Benjamin, Dai Hongkai, and Simchowitz Max · 2024
Cited alongside, same era.
Visarl: Visual reinforcement learning guided by human saliency
Liang Anthony, Thomason Jesse, and Bıyık Erdem · 2024
Cited alongside, same era.
Exponential moving average of weights in deep learning: Dynamics and benefits
Morales-Brotons Daniel and and Hadrien Hendrikx Thijs, Vogels · 2024
Cited alongside, same era.
The ingredients for robotic diffusion transformers
Sudeep Dasari, Oier Mees, Sebastian Zhao, Mohan Kumar Srirama, and Sergey Levine · 2024
Cited alongside, same era.
Batch active learning of reward functions from human preferences
Bıyık Erdem, Anari Nima, and Sadigh Dorsa · 2024
Cited alongside, same era.
3d diffuser actor: Policy diffusion with 3d scene representations
Tsung-Wei Ke, Nikolaos Gkanatsios, and Katerina Fragkiadaki · 2024
Cited alongside, same era.
Adam with model exponential moving average is effective for nonconvex optimization
Ahn Kwangjun and Cutkosky Ashok · 2024
Cited alongside, same era.
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
Adapt2reward: Adapting video-language models to generalizable robotic rewards via failure prompts
Yang Yanting, Chen Minghao, Qiu Qibo, Wu Jiahao, Wang Wenxiao, Lin Binbin, Guan Ziyu, and He Xiaofei · 2024
Later among the works it cites.
Rl-vlm-f: Reinforcement learning from vision language foundation model feedback
Wang Yufei, Sun Zhanyi, Zhang Jesse, Xian Zhou, Biyik Erdem, Held David, and Erickson Zackory · 2024
Later among the works it cites.
Dare: Diffusion policy for autonomous robot exploration
Cao Yuhong, Lew Jeric, Liang Jingsong, Cheng Jin, and Sartoretti Guillaume · 2024
Later among the works it cites.
Gnfactor: Multi-task real robot learning with generalizable neural feature fields, 2024
Yanjie Ze, Ge Yan, Yueh-Hua Wu, Annabella Macaluso, Yuying Ge, Jianglong Ye, Nicklas Hansen, Li Erran Li, and Xiaolong Wang · 2024
Later among the works it cites.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu, Muhan Wang, and Huazhe Xu · 2024
Later among the works it cites.
https://en.wikipedia.org/wiki/Markov_decision_process
Markov decision process · 2025
Closest in time.
Diffusion forcing: Next-token prediction meets full-sequence diffusion
Boyuan Chen, Diego Martí Monsó, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann · 2025
Closest in time.
A real-to-sim-to-real approach to robotic manipulation with vlm-generated iterative keypoint rewards, 2025
Shivansh Patel, Xinchen Yin, Wenlong Huang, Shubham Garg, Hooshang Nayyeri, Li Fei-Fei, Svetlana Lazebnik, and Yunzhu Li · 2025
Closest in time.
Diffusion trajectory-guided policy for long-horizon robot manipulation
Fan Shichao, Yang Quantao, Liu Yajie, Wu Kun, Che Zhengping, Liu Qingjie, and Wan Min · 2025
Closest in time.