Fetching the paper…
Reading the bibliography…
When operating in service of people, robots need to optimize rewards aligned with end-user preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
On a space of totally additive functions
Leonid Vasilevich Kantorovich and SG Rubinshtein · 1958
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Cédric Villani et al · 2009
Earlier work this paper cites.
Low-dimensional embedding using adaptively selected ordinal data
Kevin G Jamieson and Robert D Nowak · 2011
Earlier work this paper cites.
Nonlinear inverse reinforcement learning with gaussian processes
Sergey Levine, Zoran Popovic, and Vladlen Koltun · 2011
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Unsupervised perceptual rewards for imitation learning
Pierre Sermanet, Kelvin Xu, and Sergey Levine · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al · 2017
Earlier work this paper cites.
Batch active preference-based learning of reward functions
Erdem Biyik and Dorsa Sadigh · 2018
Earlier work this paper cites.
Human-driven feature selection for a robotic agent learning classification tasks from demonstration
Kalesha Bullard, Sonia Chernova, and Andrea L Thomaz · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Earlier work this paper cites.
An incremental feature set refinement in a programming by demonstration scenario
Hoai Luu-Duc and Jun Miura · 2019
Earlier work this paper cites.
Computational optimal transport: With applications to data science
Gabriel Peyré, Marco Cuturi, et al · 2019
Earlier work this paper cites.
Provably efficient imitation learning from observation alone
Wen Sun, Anirudh Vemula, Byron Boots, and Drew Bagnell · 2019
Cited alongside, same era.
Wasserstein adversarial imitation learning
Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche, Jalal Etesami, and Thai Hong Linh · 2019
Cited alongside, same era.
Safe imitation learning via fast bayesian reward inference from preferences
Daniel Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum · 2020
Cited alongside, same era.
Concept2Robot: Learning manipulation concepts from instructions and human demonstrations
Lin Shao, Toki Migimatsu, Qiang Zhang, Karen Yang, and Jeannette Bohg · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Later among the works it cites.
Imitation learning with sinkhorn distances
Georgios Papagiannis and Yunpeng Li · 2022
Later among the works it cites.
Causal confusion and reward misidentification in preference-based reward learning
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D Dragan, and Daniel S Brown · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Motion2vec: Semi-supervised representation learning from surgical videos
Ajay Kumar Tanwani, Pierre Sermanet, Andy Yan, Raghav Anand, Mariano Phielipp, and Ken Goldberg · 2020
Cited alongside, same era.
Squeezesegv3: Spatially-adaptive convolution for efficient point-cloud segmentation
Chenfeng Xu, Bichen Wu, Zining Wang, Wei Zhan, Peter Vajda, Kurt Keutzer, and Masayoshi Tomizuka · 2020
Cited alongside, same era.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2020
Cited alongside, same era.
Feature expansive reward learning: Rethinking human input
Andreea Bobu, Marius Wiggert, Claire Tomlin, and Anca D Dragan · 2021
Cited alongside, same era.
When the ventral visual stream is not enough: A deep learning account of medial temporal lobe involvement in perception
Tyler Bonnen, Daniel LK Yamins, and Anthony D Wagner · 2021
Cited alongside, same era.
Fixation patterns in simple choice reflect optimal information sampling
Frederick Callaway, Antonio Rangel, and Thomas L Griffiths · 2021
Cited alongside, same era.
Learning generalizable robotic reward functions from” in-the-wild” human videos
Annie S Chen, Suraj Nair, and Chelsea Finn · 2021
Cited alongside, same era.
Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik · 2022
Later among the works it cites.
Image2point: 3d point-cloud understanding with 2d image pretrained models
Chenfeng Xu, Shijia Yang, Tomer Galanti, Bichen Wu, Xiangyu Yue, Bohan Zhai, Wei Zhan, Peter Vajda, Kurt Keutzer, and Masayoshi Tomizuka · 2022
Later among the works it cites.
Xirl: Cross-embodiment inverse reinforcement learning
Kevin Zakka, Andy Zeng, Pete Florence, Jonathan Tompson, Jeannette Bohg, and Debidatta Dwibedi · 2022
Later among the works it cites.
Time-efficient reward learning via visually assisted cluster ranking
David Zhang, Micah Carroll, Andreea Bobu, and Anca Dragan · 2022
Later among the works it cites.
See to touch: Learning tactile dexterity through visual incentives
Irmak Guzey, Yinlong Dai, Ben Evans, Soumith Chintala, and Lerrel Pinto · 2023
Closest in time.
Language-driven representation learning for robotics
Siddharth Karamcheti, Suraj Nair, Annie S Chen, Thomas Kollar, Chelsea Finn, Dorsa Sadigh, and Percy Liang · 2023
Closest in time.
Graph inverse reinforcement learning from diverse videos
Sateesh Kumar, Jonathan Zamora, Nicklas Hansen, Rishabh Jangir, and Xiaolong Wang · 2023
Closest in time.
Optimal transport for offline imitation learning
Yicheng Luo, zhengyao jiang, Samuel Cohen, Edward Grefenstette, and Marc Peter Deisenroth · 2023
Closest in time.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2023
Closest in time.
Benchmarks and algorithms for offline preference-based reward learning
Daniel Shin, Anca D Dragan, and Daniel S Brown · 2023
Closest in time.
Alignment with human representations supports robust few-shot learning
Ilia Sucholutsky and Thomas L Griffiths · 2023
Closest in time.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2023
Closest in time.