Fetching the paper…
Reading the bibliography…
Methods to learn under algorithmic triage have predominantly focused on supervised learning settings where each decision, or prediction, is independent of each other.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Instant messaging and interruption: Influence of task type on performance
Mary Czerwinski, Edward Cutrell, and Eric Horvitz · 2000
Earlier work this paper cites.
Temporal Abstraction in Reinforcement Learning
Doina Precup and Richard S. Sutton · 2000
Earlier work this paper cites.
Torcs, the open racing car simulator
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner · 2000
Earlier work this paper cites.
Behavioural impacts of advanced driver assistance systems–an overview
K. Brookhuis, D. De Waard, and W. Janssen · 2001
Earlier work this paper cites.
Learning and reasoning about interruption
Eric Horvitz and Johnson Apacible · 2003
Earlier work this paper cites.
Crawdad dataset cambridge/haggle (v. 2006-09-15)
James Scott, Richard Gass, Jon Crowcroft, Pan Hui, Christophe Diot, and Augustin Chaintreau · 2006
Earlier work this paper cites.
Understanding and developing models for detecting and differentiating breakpoints during interactive tasks
Shamsi T Iqbal and Brian P Bailey · 2007
Earlier work this paper cites.
Classification with a reject option using a hinge loss
P. Bartlett and M. Wegkamp · 2008
Earlier work this paper cites.
Feature construction for inverse reinforcement learning
Sergey Levine, Zoran Popovic, and Vladlen Koltun · 2010
Earlier work this paper cites.
Integrating reinforcement learning with human demonstrations of varying ability
Matthew E Taylor, Halit Bener Suay, and Sonia Chernova · 2011
Earlier work this paper cites.
Blending autonomous exploration and apprenticeship learning
Thomas J Walsh, Daniel K Hewlett, and Clayton T Morrison · 2011
Earlier work this paper cites.
Pomcop: Belief space planning for sidekicks in cooperative games
O. Macindoe, L. Kaelbling, and T. Lozano-Pérez · 2012
Earlier work this paper cites.
Teaching on a budget: Agents advising agents in reinforcement learning
Lisa Torrey and Matthew Taylor · 2013
Earlier work this paper cites.
The algorithmic anatomy of model-based evaluation
Nathaniel D Daw and Peter Dayan · 2014
Earlier work this paper cites.
Efficient model learning from joint-action demonstrations for human-robot collaborative tasks
S. Nikolaidis, R. Ramakrishnan, K. Gu, and J. Shah · 2015
Earlier work this paper cites.
On convergence of emphatic temporal-difference learning
H. Yu · 2015
Earlier work this paper cites.
Learning with rejection
C. Cortes, G. DeSalvo, and M. Mohri · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, S. Russell, P. Abbeel, and A. Dragan · 2016
Earlier work this paper cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S. Sutton, A. Rupam Mahmood, and Martha White · 2016
Cited alongside, same era.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Cited alongside, same era.
Carla: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Cited alongside, same era.
The atari grand challenge dataset
Vitaly Kurin, Sebastian Nowozin, Katja Hofmann, Lucas Beyer, and Bastian Leibe · 2017
Cited alongside, same era.
Mathematical models of adaptation in human-robot collaboration
S. Nikolaidis, J. Forlizzi, D. Hsu, J. Shah, and S. Srinivasa · 2017
Cited alongside, same era.
Learning to collaborate in markov decision processes
Goran Radanovic, Rati Devidze, David C. Parkes, and Adish Singla · 2019
Later among the works it cites.
The algorithmic automation problem: Prediction, triage, and human effort
M. Raghu, K. Blumer, G. Corrado, J. Kleinberg, Z. Obermeyer, and S. Mullainathan · 2019
Later among the works it cites.
Exploring applications of deep reinforcement learning for real-world autonomous driving systems
V. Talpaert et al · 2019
Later among the works it cites.
Combating label noise in deep learning using abstention
S. Thulasidasan, T. Bhattacharya, J. Bilmes, G. Chennupati, and J. Mohd-Yusof · 2019
Later among the works it cites.
Learner-aware teaching: Inverse reinforcement learning with preferences and constraints
S. Tschiatschek, A. Ghosh, L. Haug, R. Devidze, and A. Singla · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bias-reduced uncertainty estimation for deep neural classifiers
Y. Geifman, G. Uziel, and R. El-Yaniv · 2018
Cited alongside, same era.
Learning policy representations in multiagent systems
A. Grover, M. Al-Shedivat, J. Gupta, Y. Burda, and H. Edwards · 2018
Cited alongside, same era.
Teaching inverse reinforcement learners via features and demonstrations
L. Haug, S. Tschiatschek, and A. Singla · 2018
Cited alongside, same era.
Modality switching for mitigation of sensory adaptation and habituation in personal navigation systems
Kyle Kotowick and Julie Shah · 2018
Cited alongside, same era.
Consistent algorithms for multiclass classification with an abstain option
H. Ramaswamy, A. Tewari, and S. Agarwal · 2018
Cited alongside, same era.
Shared autonomy via deep reinforcement learning
Siddharth Reddy, Anca D Dragan, and Sergey Levine · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Later among the works it cites.
Open problems in cooperative ai
Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R McKee, Joel Z Leibo, Kate Larson, and Thore Graepel · 2020
Later among the works it cites.
Regression under human assistance
A. De, P. Koley, N. Ganguly, and M. Gomez-Rodriguez · 2020
Later among the works it cites.
Towards deployment of robust cooperative ai agents: An algorithmic framework for learning adaptive policies
A. Ghosh, S. Tschiatschek, H. Mahdavi, and A. Singla · 2020
Later among the works it cites.
Aligning superhuman ai with human behavior: Chess as a model system
Reid McIlroy-Young, Siddhartha Sen, Jon Kleinberg, and Ashton Anderson · 2020
Later among the works it cites.
Consistent estimators for learning to defer to an expert
Hussein Mozannar and David Sontag · 2020
Later among the works it cites.
Learning to complement humans
Bryan Wilder, Eric Horvitz, and Ece Kamar · 2020
Later among the works it cites.
Provably convergent two-timescale off-policy actor-critic with function approximation
Shangtong Zhang, Bo Liu, Hengshuai Yao, and Shimon Whiteson · 2020
Later among the works it cites.
Learning not to learn in the presence of noisy labels
Liu Ziyin, Blair Chen, Ru Wang, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency, and Masahito Ueda · 2020
Later among the works it cites.
Optimizing AI for Teamwork
Gagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz, and Daniel S. Weld · 2021
Closest in time.
Cooperative ai: machines must learn to find common ground
Allan Dafoe, Yoram Bachrach, Gillian Hadfield, Eric Horvitz, Kate Larson, and Thore Graepel · 2021
Closest in time.
Classification under human assistance
Abir De, Nastaran Okati, Ali Zarezade, and Manuel Gomez-Rodriguez · 2021
Closest in time.
Safe option-critic: Learning safety in the option-critic architecture
Arushi Jain, Khimya Khetarpal, and Doina Precup · 2021
Closest in time.
Learning to switch between machines and humans
Vahid Balazadeh Meresht, Abir De, Adish Singla, and Manuel Gomez-Rodriguez · 2021
Closest in time.
Differentiable learning under triage
Nastaran Okati, Abir De, and Manuel Gomez-Rodriguez · 2021
Closest in time.