Fetching the paper…
Reading the bibliography…
This article considers the problem of diagnosing certain common errors in reward design.
Theory of games and economic behavior
Von Neumann, J., Morgenstern, O., 1944 · 1944
Earlier work this paper cites.
Problems of monetary management: the UK experience, in: Monetary theory and practice. Springer, pp. 91–121
Goodhart, C.A., 1984 · 1984
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P., 1993 · 1993
Earlier work this paper cites.
‘Improving ratings’: audit in the British university system
Strathern, M., 1997 · 1997
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping, in: Fifteenth International Conference on Machine Learning (ICML), Citeseer. pp. 463–471
Randløv, J., Alstrøm, P., 1998 · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A., Harada, D., Russell, S., 1999 · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning, in: Seventeenth International Conference on Machine Learning (ICML)
Ng, A., Russell, S., 2000 · 2000
Earlier work this paper cites.
Potential-based shaping and Q-value initialization are equivalent
Wiewiora, E., 2003 · 2003
Earlier work this paper cites.
Reinforcement learning for optimized trade execution, in: 23rd International Conference on Machine Learning (ICML), pp. 673–680
Nevmyvaka, Y., Feng, Y., Kearns, M., 2006 · 2006
Earlier work this paper cites.
Improved methods for estimating relative crash risk in a case—control study of blood alcohol levels, in: Joint Meeting of The International Council on Alcohol, Drugs & Traffic Safety (ICADTS) and The International Association of Forensic Toxicologists (TIAFT)
Peck, R., Gebers, M., Voas, R., Romano, E., 2007 · 2007
Earlier work this paper cites.
Potential-based shaping in model-based reinforcement learning, in: Twenty-third AAAI Conference on Artificial Intelligence, pp. 604–609
Asmuth, J., Littman, M.L., Zinkov, R., 2008 · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning, in: Twenty-third AAAI Conference on Artificial Intelligence, pp. 1433–1438
Ziebart, B.D., Maas, A.L., Bagnell, J.A., Dey, A.K., 2008 · 2008
Earlier work this paper cites.
Dynamic potential-based reward shaping, in: 11th International Conference on Autonomous Agents and Multiagent Systems, IFAAMAS. pp. 433–440
Devlin, S.M., Kudenko, D., 2012 · 2012
Earlier work this paper cites.
Reinforcement learning with human and MDP reward, in: 11th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), IFAAMAS
Knox, W.B., Stone, P., 2012 · 2012
Earlier work this paper cites.
Tax collections optimization for new york state
Miller, G., Weatherwax, M., Gardinier, T., Abe, N., Melville, P., Pendus, C., Jensen, D., Reddy, C.K., Thomas, V., Bennett, J., et al., 2012 · 2012
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Roijers, D.M., Vamplew, P., Whiteson, S., Dazeley, R., 2013 · 2013
Earlier work this paper cites.
Maximum entropy deep inverse reinforcement learning
Wulfmeier, M., Ondruska, P., Posner, I., 2015 · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., Mané, D., 2016 · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J.J., Schaul, T., Van Hasselt, H., Silver, D., 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences, in: Advances in Neural Information Processing Systems (NIPS), pp. 4299–4307
Christiano, P.F., Leike, J., Brown, T., Martic, M., Legg, S., Amodei, D., 2017 · 2017
Cited alongside, same era.
Carla: An open urban driving simulator, in: Conference on Robot Learning (CoRL)
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., Koltun, V., 2017 · 2017
Cited alongside, same era.
Examining accident reports involving autonomous vehicles in california
Favarò, F.M., Nader, N., Eurich, S.O., Tripp, M., Varadaraju, N., 2017 · 2017
Cited alongside, same era.
Reward shaping in episodic reinforcement learning, in: AAAI Conference on Artificial Intelligence, ACM
Grzes, M., 2017 · 2017
Cited alongside, same era.
Inverse reward design, in: Advances in Neural Information Processing Systems (NIPS), pp. 6765–6774
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S.J., Dragan, A., 2017 · 2017
Cited alongside, same era.
Dynamic input for deep reinforcement learning in autonomous driving, in: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Huegle, M., Kalweit, G., Mirchevska, B., Werling, M., Boedecker, J., 2019 · 2019
Later among the works it cites.
Learning to drive in a day, in: International Conference on Robotics and Automation (ICRA), IEEE. pp. 8248–8254
Kendall, A., Hawke, J., Janz, D., Mazur, P., Reda, D., Allen, J.M., Lam, V.D., Bewley, A., Shah, A., 2019 · 2019
Later among the works it cites.
Urban driving with multi-objective deep reinforcement learning, in: 18th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS), IFAAMAS. pp. 359–367
Li, C., Czarnecki, K., 2019 · 2019
Later among the works it cites.
Deep distributional reinforcement learning based high-level driving policy determination
Min, K., Kim, H., Huh, K., 2019 · 2019
Later among the works it cites.
Towards learning multi-agent negotiations via self-play, in: IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning end-to-end multimodal sensor policies for autonomous navigation, in: Conference on Robot Learning (CoRL), pp. 249–261
Liu, G.H., Siravuru, A., Prabhakar, S., Veloso, M., Kantor, G., 2017 · 2017
Cited alongside, same era.
Combining neural networks and tree search for task and motion planning in challenging environments, in: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE. pp. 6059–6066
Paxton, C., Raman, V., Hager, G.D., Kobilarov, M., 2017 · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., Fürnkranz, J., et al., 2017 · 2017
Cited alongside, same era.
Policy gradient based reinforcement learning approach for autonomous highway driving, in: IEEE Conference on Control Technology and Applications (CCTA), IEEE. pp. 670–675
Aradi, S., Becsi, T., Gaspar, P., 2018 · 2018
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning, in: International Conference on Learning Representations (ICLR)
Fu, J., Luo, K., Levine, S., 2018 · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., Amodei, D., 2018 · 2018
Cited alongside, same era.
CIRL: Controllable imitative reinforcement learning for vision-based self-driving, in: European Conference on Computer Vision (ECCV), pp. 584–599
Liang, X., Wang, T., Yang, L., Xing, E., 2018 · 2018
Cited alongside, same era.
Tang, Y., 2019 · 2019
Later among the works it cites.
Artificial intelligence: a modern approach
Russell, S., Norvig, P., 2020 · 2020
Later among the works it cites.
End-to-end model-free reinforcement learning for urban driving using implicit affordances, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7153–7162
Toromanoff, M., Wirbel, E., Moutarde, F., 2020 · 2020
Later among the works it cites.
Learning hierarchical behavior and motion planning for autonomous driving, in: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Wang, J., Wang, Y., Zhang, D., Yang, Y., Xiong, R., 2020 · 2020
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Dulac-Arnold, G., Levine, N., Mankowitz, D.J., Li, J., Paduraru, C., Gowal, S., Hester, T., 2021 · 2021
Closest in time.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B.R., Sobh, I., Talpaert, V., Mannion, P., Al Sallab, A.A., Yogamani, S., Pérez, P., 2021 · 2021
Closest in time.
A survey of deep rl and il for autonomous driving policy learning
Zhu, Z., Zhao, H., 2021 · 2021
Closest in time.
Faulty reward functions in the wild
Amodei, D., Clark, J., 2016 · 2022
Closest in time.
Rates of motor vehicle crashes, injuries and deaths in relation to driver age, United States, 2014-2015
Tefft, B., 2017 · 2022
Closest in time.
Traffic volume trends
U.S. Dept. of Transportation, 2020 · 2022
Closest in time.
Departmental guidance: Treatment of the value of preventing fatalities and injuries in preparing economic analyses
U.S. Dept. of Transportation, 2021 · 2022
Closest in time.
Departmental guidance on valuation of a statistical life in economic analysis
U.S. Dept. of Transportation, 2022 · 2022
Closest in time.
Navigating occluded intersections with autonomous vehicles using deep reinforcement learning, in: IEEE International Conference on Robotics and Automation (ICRA), IEEE. pp. 2034–2039
Isele, D., Rahimi, R., Cosgun, A., Subramanian, K., Fujimura, K., 2018 · 2039
Closest in time.
End-to-end race driving with deep reinforcement learning, in: IEEE International Conference on Robotics and Automation (ICRA), IEEE. pp. 2070–2075
Jaritz, M., De Charette, R., Toromanoff, M., Perot, E., Nashashibi, F., 2018 · 2075
Closest in time.