Fetching the paper…
Reading the bibliography…
As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally, using only a measure of task performance as feedback, can violate societal norms for acceptable behavior or cause harm.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R.; and Le, Q. V. 2019 · 1906
Earlier work this paper cites.
LeDeepChef: Deep Reinforcement Learning Agent for Families of Text-Based Games
Adolphs, L.; and Hofmann, T. 2019 · 1909
Earlier work this paper cites.
Learning Norms from Stories: A Prior for Value Aligned Agents
Frazier, S.; Nahian, M. S. A.; Riedl, M.; and Harrison, B. 2019 · 1912
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 1937
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P.; and Ng, A. Y. 2004 · 2004
Earlier work this paper cites.
How to Avoid Being Eaten by a Grue: Structured Exploration Strategies for Textual Worlds
Ammanabrolu, P.; Tien, E.; Hausknecht, M.; and Riedl, M. O. 2020 · 2006
Earlier work this paper cites.
The nature, importance, and difficulty of machine ethics
Moor, J. H. 2006 · 2006
Earlier work this paper cites.
Playing Text-Based Games with Common Sense
Dambekodi, S.; Frazier, S.; Ammanabrolu, P.; and Riedl, M. O. 2020 · 2012
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, S.; Subramanian, K.; Scholz, J.; Isbell, C. L.; and Thomaz, A. L. 2013 · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N. 2014 · 2014
Earlier work this paper cites.
Aligning superintelligence with human interests: A technical research agenda
Soares, N.; and Fallenstein, B. 2014 · 2014
Earlier work this paper cites.
Policy Shaping with Human Teachers
Cederborg, T.; Grover, I.; Isbell Jr, C. L.; and Thomaz, A. L. 2015 · 2015
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, K.; Kulkarni, T.; and Barzilay, R. 2015 · 2015
Cited alongside, same era.
Research priorities for robust and beneficial artificial intelligence
Russell, S.; Dewey, D.; and Tegmark, M. 2015 · 2015
Cited alongside, same era.
Deep Reinforcement Learning with a Natural Language Action Space
He, J.; Chen, J.; He, X.; Gao, J.; Li, L.; Deng, L.; and Ostendorf, M. 2016 · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J.; and Ermon, S. 2016 · 2016
Cited alongside, same era.
Model-free imitation learning with policy optimization
Ho, J.; Gupta, J.; and Ermon, S. 2016 · 2016
Cited alongside, same era.
Value alignment or misalignment-what will keep systems accountable?
Arnold, T.; Kasenberg, D.; and Scheutz, M. 2017 · 2017
Learn What Not to Learn: Action Elimination with Deep Reinforcement Learning
Zahavy, T.; Haroush, M.; Merlis, N.; Mankowitz, D. J.; and Mannor, S. 2018 · 2018
Later among the works it cites.
Using reinforcement learning to learn how to play text-based games
Zelinka, M. 2018 · 2018
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Later among the works it cites.
Human compatible: Artificial intelligence and the problem of control
Russell, S. 2019 · 2019
Later among the works it cites.
Efficient supervision for robot learning via imitation, simulation, and adaptation
Wulfmeier, M. 2019 · 2019
Later among the works it cites.
Comprehensible Context-driven Text Game Playing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Leike, J.; Martic, M.; Krakovna, V.; Ortega, P. A.; Everitt, T.; Lefrancq, A.; Orseau, L.; and Legg, S. 2017 · 2017
Cited alongside, same era.
Lin, Z.; Harrison, B.; Keech, A.; and Riedl, M. O. 2017 · 2017
Cited alongside, same era.
Third-person imitation learning
Stadie, B.; Abbeel, P.; and Sutskever, I. 2017 · 2017
Cited alongside, same era.
Textworld: A learning environment for text-based games
Côté, M.-A.; Kádár, Á.; Yuan, X.; Kybartas, B.; Barnes, T.; Fine, E.; Moore, J.; Hausknecht, M.; El Asri, L.; Adada, M.; et al. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Policy Shaping with Supervisory Attention Driven Exploration
Faulkner, T. K.; Short, E. S.; and Thomaz, A. L. 2018 · 2018
Cited alongside, same era.
Yin, X.; and May, J. 2019 · 2019
Later among the works it cites.
Graph Constrained Reinforcement Learning for Natural Language Action Spaces
Ammanabrolu, P.; and Hausknecht, M. 2020 · 2020
Later among the works it cites.
Scruples: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes
Lourie, N.; Bras, R. L.; and Choi, Y. 2020 · 2020
Later among the works it cites.
Learning Norms from Stories: A Prior for Value Aligned Agents
Nahian, M. S. A.; Frazier, S.; Riedl, M.; and Harrison, B. 2020 · 2020
Later among the works it cites.
Fine-Tuning a Transformer-Based Language Model to Avoid Generating Non-Normative Text
Peng, X.; Li, S.; Frazier, S.; and Riedl, M. 2020 · 2020
Later among the works it cites.
Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based Games
Xu, Y.; Fang, M.; Chen, L.; Du, Y.; Zhou, J. T.; and Zhang, C. 2020 · 2020
Later among the works it cites.