Fetching the paper…
Reading the bibliography…
In robot imitation learning, policy performance is tightly coupled with the quality and composition of the demonstration data.
Extensions of lipschitz mappings into a hilbert space
W. B. Johnson, J. Lindenstrauss, et al · 1984
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
A unified approach for motion and force control of robot manipulators: The operational space formulation
O. Khatib · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
E. Greensmith, P. L. Bartlett, and J. Baxter · 2004
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Active Learning
B. Settles · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Earlier work this paper cites.
Causal confusion in imitation learning
P. De Haan, D. Jayaraman, and S. Levine · 2019
Earlier work this paper cites.
On the accuracy of influence functions for measuring group effects
P. W. W. Koh, K.-S. Ang, H. Teo, and P. S. Liang · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
A. Ghorbani and J. Zou · 2019
Earlier work this paper cites.
Uncertainty-aware data aggregation for deep imitation learning
Y. Cui, D. Isele, S. Niekum, and K. Fujimura · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
J. Martens · 2020
Earlier work this paper cites.
Influence functions in deep learning are fragile
S. Basu, P. Pope, and S. Feizi · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2021
Earlier work this paper cites.
Deduplicating training data makes language models better
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini · 2022
Earlier work this paper cites.
Cacti: A framework for scalable multi-task multi-scene visual imitation learning
Z. Mandi, H. Bharadhwaj, V. Moens, S. Song, A. Rajeswaran, and V. Kumar · 2022
Earlier work this paper cites.
If influence functions are the answer, then what is the question?
J. Bae, N. Ng, A. Lo, M. Ghassemi, and R. B. Grosse · 2022
Earlier work this paper cites.
Datamodels: Understanding predictions with data and data with predictions
A. Ilyas, S. M. Park, L. Engstrom, G. Leclerc, and A. Madry · 2022
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2022
Cited alongside, same era.
Studying large language model generalization with influence functions
R. Grosse, J. Bae, C. Anil, N. Elhage, A. Tamkin, A. Tajdini, B. Steiner, D. Li, E. Durmus, E. Perez, et al · 2023
Cited alongside, same era.
Trak: Attributing model behavior at scale
S. M. Park, K. Georgiev, A. Ilyas, G. Leclerc, and A. Madry · 2023
Cited alongside, same era.
D4: Improving llm pretraining via document de-duplication and diversification
K. Tirumala, D. Simig, A. Aghajanyan, and A. Morcos · 2023
Cited alongside, same era.
Learning to discern: Imitating heterogeneous human demonstrations with preference and representation learning
S. Kuhar, S. Cheng, S. Chopra, M. Bronars, and D. Xu · 2023
Cited alongside, same era.
What is your data worth to gpt? llm-scale data valuation with influence functions
S. K. Choe, H. Ahn, J. Bae, K. Zhao, M. Kang, Y. Chung, A. Pratapa, W. Neiswanger, E. Strubell, T. Mitamura, et al · 2024
Later among the works it cites.
Data attribution at scale
A. Madry, A. Ilyas, L. Engstrom, S. M. Park, and K. Georgiev · 2024
Later among the works it cites.
Intriguing properties of data attribution on diffusion models
X. Zheng, T. Pang, C. Du, J. Jiang, and M. Lin · 2024
Later among the works it cites.
Less: Selecting influential data for targeted instruction tuning
M. Xia, S. Malladi, S. Gururangan, S. Arora, and D. Chen · 2024
Later among the works it cites.
Tsds: Data selection for task-specific model finetuning
Z. Liu, A. Karbasi, and T. Rekatsinas · 2024
Later among the works it cites.
How generalizable is my behavior cloning policy? a statistical approach to trustworthy performance evaluation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
RT-1: Robotics Transformer for Real-World Control at Scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. H. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T.-W. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han · 2023
Cited alongside, same era.
Mimicgen: A data generation system for scalable robot learning using human demonstrations
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y. Narang, L. Fan, Y. Zhu, and D. Fox · 2023
Cited alongside, same era.
Scaling Robot Learning with Semantically Imagined Experience
T. Yu, T. Xiao, J. Tompson, A. Stone, S. Wang, A. Brohan, J. Singh, C. Tan, D. M, J. Peralta, K. Hausman, B. Ichter, and F. Xia · 2023
Cited alongside, same era.
Modeldiff: A framework for comparing learning algorithms
H. Shah, S. M. Park, A. Ilyas, and A. Madry · 2023
Cited alongside, same era.
The journey, not the destination: How data guides diffusion models
K. Georgiev, J. Vendrow, H. Salman, S. M. Park, and A. Madry · 2023
Cited alongside, same era.
Data quality in imitation learning
S. Belkhale, Y. Cui, and D. Sadigh · 2023
Cited alongside, same era.
J. A. Vincent, H. Nishimura, M. Itkina, P. Shah, M. Schwager, and T. Kollar · 2024
Later among the works it cites.
Data attribution for diffusion models: Timestep-induced bias in influence estimation
T. Xie, H. Li, A. Bai, and C.-J. Hsieh · 2024
Later among the works it cites.
Racer: Rich language-guided failure recovery policies for imitation learning
Y. Dai, J. Lee, N. Fazeli, and J. Chai · 2024
Later among the works it cites.
Real-Time Anomaly Detection and Reactive Planning with Large Language Models
R. Sinha, A. Elhafsi, C. Agia, M. Foutter, E. Schmerling, and M. Pavone · 2024
Later among the works it cites.
Robot data curation with mutual information estimators
J. Hejna, S. Mirchandani, A. Balakrishna, A. Xie, A. Wahid, J. Tompson, P. Sanketi, D. Shah, C. Devin, and D. Sadigh · 2025
Closest in time.
Curating demonstrations using online experience
A. S. Chen, A. M. Lessing, Y. Liu, and C. Finn · 2025
Closest in time.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2025
Closest in time.
Robotic control via embodied chain-of-thought reasoning
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Remix: Optimizing data mixtures for large scale imitation learning
J. Hejna, C. A. Bhateja, Y. Jiang, K. Pertsch, and D. Sadigh · 2025
Closest in time.
Datamil: Selecting data for robot imitation learning with datamodels
S. Dass, A. Khaddaj, L. Engstrom, A. Madry, A. Ilyas, and R. Martín-Martín · 2025
Closest in time.
Attribute-to-delete: Machine unlearning via datamodel matching
K. Georgiev, R. Rinberg, S. M. Park, S. Garg, A. Ilyas, A. Madry, and S. Neel · 2025
Closest in time.
Magic: Near-optimal data attribution for deep learning
A. Ilyas and L. Engstrom · 2025
Closest in time.
Optimizing ml training with metagradient descent
L. Engstrom, A. Ilyas, B. Chen, A. Feldmann, W. Moses, and A. Madry · 2025
Closest in time.
Diffusion attribution score: Evaluating training data influence in diffusion model
J. Lin, L. Tao, M. Dong, and C. Xu · 2025
Closest in time.
Influence functions for scalable data attribution in diffusion models
B. K. Mlodozeniec, R. Eschenhagen, J. Bae, A. Immer, D. Krueger, and R. E. Turner · 2025
Closest in time.
Unpacking failure modes of generative policies: Runtime monitoring of consistency and progress
C. Agia, R. Sinha, J. Yang, Z. Cao, R. Antonova, M. Pavone, and J. Bohg · 2025
Closest in time.