Fetching the paper…
Reading the bibliography…
Our goal is for AI systems to correctly identify and act according to their human user's objectives.
Models of man; social and rational
Simon, H. A · 1957
Earlier work this paper cites.
The Optimal Control of Partially Observable Markov Processes
Sondik, E. J · 1971
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Tversky, A. and Kahneman, D · 1975
Earlier work this paper cites.
A Modern Approach
Russell, S. and Norvig, P · 1995
Earlier work this paper cites.
Planning and Acting in Partially Observable Stochastic Domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Ng, A. Y. and Russell, S · 2000
Earlier work this paper cites.
The Complexity of Decentralized Control of Markov Decision Processes
Bernstein, D. S., Givan, R., Immerman, N., and Zilberstein, S · 2002
Earlier work this paper cites.
Point-Based Value Iteration: An Anytime Algorithm for POMDPs
Pineau, J., Gordon, G., and Thrun, S · 2003
Earlier work this paper cites.
Dynamic Programming for Partially Observable Stochastic Games
Hansen, E. A · 2004
Cited alongside, same era.
Heuristic search value iteration for pomdps
Smith, T. and Simmons, R · 2004
Cited alongside, same era.
Bandit Based Monte-Carlo Planning
Kocsis, L. and Szepesvári, C · 2006
Cited alongside, same era.
SARSOP: Efficient Point-Based POMDP Planning by Approximating Optimally Reachable Belief Spaces
Kurniawati, H., Hsu, D., and Lee, W. S · 2008
Cited alongside, same era.
Incremental Policy Generation for Finite-Horizon DEC-POMDPs
Amato, C., Dibangoye, J. S., and Zilberstein, S · 2009
Cited alongside, same era.
Planning under uncertainty for robotic tasks with mixed observability
Ong, S. C., Png, S. W., Hsu, D., and Lee, W. S · 2010
Cited alongside, same era.
Evaluating fluency in human-robot collaboration
Hoffman, G · 2013
Later among the works it cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N · 2014
Later among the works it cites.
Faulty Reward Functions in the Wild
Amodei, D. and Clark, J · 2016
Later among the works it cites.
Concrete Problems in AI Safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Later among the works it cites.
Cooperative Inverse Reinforcement Learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Later among the works it cites.
Pragmatic-pedagogic value alignment
Fisac, J. F., Gates, M. A., Hamrick, J. B., Liu, C., Hadfield-Menell, D., Palaniappan, M., Malik, D., Sastry, S. S., Griffiths, T. L., and Dragan, A. D · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Monte Carlo Planning in Large POMDPs
Silver, D. and Veness, J · 2010
Cited alongside, same era.
DESPOT: Online POMDP Planning with Regularization
Ye, N., Somani, A., Hsu, D., and Lee, W · 2017
Later among the works it cites.