Fetching the paper…
Reading the bibliography…
Evaluating learned robot control policies to determine their physical task-level capabilities costs experimenter time and effort.
Mixture density networks
Christopher M Bishop · 1994
Earlier work this paper cites.
Building surrogate models based on detailed and approximate simulations
Zhiguang Qian, Carolyn Conner Seepersad, V Roshan Joseph, Janet K Allen, and CF Jeff Wu · 2006
Earlier work this paper cites.
Probabilistic matrix factorization
Andriy Mnih and Russ R Salakhutdinov · 2007
Earlier work this paper cites.
Eric Brochu, Vlad M Cora, and Nando De Freitas · 2010
Earlier work this paper cites.
Active risk estimation
Christoph Sawade, Niels Landwehr, Steffen Bickel, and Tobias Scheffer · 2010
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel · 2011
Earlier work this paper cites.
Learning surrogate models for simulation-based optimization
Alison Cozad, Nikolaos V Sahinidis, and David C Miller · 2014
Earlier work this paper cites.
Efficient benchmarking of hyperparameter optimizers via surrogates
Katharina Eggensperger, Frank Hutter, Holger Hoos, and Kevin Leyton-Brown · 2015
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando De Freitas · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
World models
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Earlier work this paper cites.
RoboTHOR: An Open Simulation-to-Real Embodied AI Platform
Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi, Eric Kolve, Roozbeh Mottaghi, Jordi Salvador, Dustin Schwenk, Eli VanderBilt, Matthew Wallingford, Luca Weihs, Mark Yatskar, and Ali Farhadi · 2020
Earlier work this paper cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou · 2020
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song · 2020
Earlier work this paper cites.
Sim2Real Predictivity: Does Evaluation in Simulation Predict Real-World Performance?
Abhishek Kadian, Joanne Truong, Aaron Gokaslan, Alexander Clegg, Erik Wijmans, Stefan Lee, Manolis Savva, Sonia Chernova, and Dhruv Batra · 2020
Earlier work this paper cites.
Cost-aware bayesian optimization
Eric Hans Lee, Valerio Perrone, Cedric Archambeau, and Matthias Seeger · 2020
Earlier work this paper cites.
A general framework for uncertainty estimation in deep learning
Antonio Loquercio, Mattia Segu, and Davide Scaramuzza · 2020
Cited alongside, same era.
Cost-aware bayesian optimization via information directed sampling
Biswajit Paria, Willie Neiswanger, Ramina Ghods, Jeff Schneider, and Barnabás Póczos · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Cited alongside, same era.
Sim-to-real transfer for vision-and-language navigation
Peter Anderson, Ayush Shrivastava, Joanne Truong, Arjun Majumdar, Devi Parikh, Dhruv Batra, and Stefan Lee · 2021
Cited alongside, same era.
Active testing: Sample-efficient model evaluation
Jannik Kossen, Sebastian Farquhar, Yarin Gal, and Tom Rainforth · 2021
Cited alongside, same era.
Discovering user-interpretable capabilities of black-box planning agents
Contrast sets for evaluating language-guided robot policies
Abrar Anwar, Rohan Gupta, and Jesse Thomason · 2024
Later among the works it cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al · 2024
Later among the works it cites.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2024
Later among the works it cites.
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
Jensen Gao, Annie Xie, Ted Xiao, Chelsea Finn, and Dorsa Sadigh · 2024
Later among the works it cites.
Minillm: Knowledge distillation of large language models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pulkit Verma, Shashank Rao Marpally, and Siddharth Srivastava · 2021
Cited alongside, same era.
CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language Understanding
Dong Wang, Ning Ding, Piji Li, and Hai-Tao Zheng · 2021
Cited alongside, same era.
Sample efficient model evaluation
Emine Yilmaz, Peter Hayes, Raza Habib, Jordan Burgess, and David Barber · 2021
Cited alongside, same era.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Cited alongside, same era.
Differential assessment of black-box ai agents
Rashmeet Kaur Nayyar, Pulkit Verma, and Siddharth Srivastava · 2022
Cited alongside, same era.
Vint: A foundation model for visual navigation
Dhruv Shah, Ajay Sridhar, Nitish Dashora, Kyle Stachowicz, Kevin Black, Noriaki Hirose, and Sergey Levine · 2022
Cited alongside, same era.
Targeted active learning for probabilistic models
Christopher Tosh, Mauricio Tec, and Wesley Tansey · 2022
Cited alongside, same era.
Deploying and Evaluating LLMs to Program Service Mobile Robots
Zichao Hu, Francesca Lucchetti, Claire Schlesinger, Yash Saxena, Anders Freeman, Sadanand Modak, Arjun Guha, and Joydeep Biswas · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn · 2024
Later among the works it cites.
Robot learning as an empirical science: Best practices for policy evaluation
Hadas Kress-Gazit, Kunimatsu Hashimoto, Naveen Kuppuswamy, Paarth Shah, Phoebe Horgan, Gordon Richardson, Siyuan Feng, and Benjamin Burchfiel · 2024
Later among the works it cites.
Evaluating real-world robot manipulation policies in simulation
Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Charles Xu, Jianlan Luo, Tobias Kreiman, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine · 2024
Later among the works it cites.
Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks
Amit Parekh, Nikolas Vitsakis, Alessandro Suglia, and Ioannis Konstas · 2024
Later among the works it cites.
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
Wilbert Pumacay, Ishika Singh, Jiafei Duan, Ranjay Krishna, Jesse Thomason, and Dieter Fox · 2024
Later among the works it cites.
Modern bayesian experimental design
Tom Rainforth, Adam Foster, Desi R Ivanova, and Freddie Bickford Smith · 2024
Later among the works it cites.
How Generalizable Is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
Joseph A Vincent, Haruki Nishimura, Masha Itkina, Paarth Shah, Mac Schwager, and Thomas Kollar · 2024
Later among the works it cites.
Decomposing the generalization gap in imitation learning for visual robotic manipulation
Annie Xie, Lisa Lee, Ted Xiao, and Chelsea Finn · 2024
Later among the works it cites.
Remembr: Building and reasoning over long-horizon spatio-temporal memory for robot navigation
Abrar Anwar, John Welsh, Joydeep Biswas, Soha Pouya, and Yan Chang · 2025
Closest in time.
Hamster: Hierarchical action models for open-world robot manipulation
Yi Li, Yuquan Deng, Jesse Zhang, Joel Jang, Marius Memmel, Caelan Reed Garrett, Fabio Ramos, Dieter Fox, Anqi Li, Abhishek Gupta, and Ankit Goyal · 2025
Closest in time.