Fetching the paper…
Reading the bibliography…
Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (LMs) may incentivize toxicity.
A theory of human motivation
Maslow, A. H · 1943
Earlier work this paper cites.
The concept of power
Dahl, R. A · 1957
Earlier work this paper cites.
The bases of social power
French, J. R., Raven, B., and Cartwright, D · 1959
Earlier work this paper cites.
On the concept of political power
Parsons, T · 1963
Earlier work this paper cites.
Economy and society: An outline of interpretive sociology , volume 2
Weber, M · 1978
Earlier work this paper cites.
Purchasing power parity, 1985
Dornbusch, R · 1985
Earlier work this paper cites.
Relational schemas and the processing of social information
Baldwin, M. W · 1992
Earlier work this paper cites.
Glossary of industrial organisation economics and competition law
Khemani, R. S · 1993
Earlier work this paper cites.
Power: A new social analysis
Russell, B · 2004
Earlier work this paper cites.
Deception: The role of consequences
Gneezy, U · 2005
Earlier work this paper cites.
Social power. , pp. 678–692
Fiske, S. T. and Berdahl, J · 2007
Earlier work this paper cites.
Inequality: Causes and consequences
Neckerman, K. M. and Torche, F · 2007
Earlier work this paper cites.
Capital as power: A study of order and creorder
Nitzan, J. and Bichler, S · 2009
Earlier work this paper cites.
The structure of reciprocity
Molm, L. D · 2010
Earlier work this paper cites.
The power of identity
Castells, M · 2011
Earlier work this paper cites.
Linking conflict to inequality and polarization
Esteban, J. and Ray, D · 2011
Earlier work this paper cites.
Intelligence explosion: Evidence and import
Muehlhauser, L. and Salamon, A · 2012
Earlier work this paper cites.
Antifragile: Things that gain from disorder , volume 3
Taleb, N. N · 2012
Earlier work this paper cites.
Power and international relations
Baldwin, D. A · 2013
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, S., Subramanian, K., Scholz, J., Isbell, C. L., and Thomaz, A. L · 2013
Earlier work this paper cites.
Fundamentals of physics
Halliday, D., Resnick, R., and Walker, J · 2013
Earlier work this paper cites.
A legal theory of finance
Pistor, K · 2013
Earlier work this paper cites.
Capital in the twenty-first century
Piketty, T · 2014
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Earlier work this paper cites.
Deep reinforcement learning with a natural language action space
He, J., Chen, J., He, X., Gao, J., Li, L., Deng, L., and Ostendorf, M · 2016
Earlier work this paper cites.
Thomas Hobbes: Leviathan (Longman library of primary sources in philosophy)
Hobbes, T. and Missner, M · 2016
Cited alongside, same era.
Using stories to teach human values to artificial agents
Riedl, M. O. and Harrison, B · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
Safe reinforcement learning via shielding
Alshiekh, M., Bloem, R., Ehlers, R., Könighofer, B., Niekum, S., and Topcu, U · 2018
Cited alongside, same era.
Textworld: A learning environment for text-based games
Côté, M.-A., Kádár, A., Yuan, X., Kybartas, B., Barnes, T., Fine, E., Moore, J., Hausknecht, M., Asri, L. E., Adada, M., et al · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Later among the works it cites.
Training value-aligned reinforcement learning agents using a normative prior
Nahian, M. S. A., Frazier, S., Harrison, B., and Riedl, M · 2021
Later among the works it cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Shridhar, M., Yuan, X., Côté, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M · 2021
Later among the works it cites.
Pre-trained language models as prior knowledge for playing text-based games
Singh, I., Singh, G., and Modi, A · 2021
Later among the works it cites.
Long range arena : A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nail: A general interactive fiction agent
Hausknecht, M., Loynd, R., Yang, G., Swaminathan, A., and Williams, J. D · 2019
Cited alongside, same era.
The third pillar: How markets and the state leave the community behind
Rajan, R · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D · 2019
Cited alongside, same era.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Cited alongside, same era.
Learning dynamic belief graphs to generalize on text-based games
Adhikari, A., Yuan, X., Côté, M.-A., Zelinka, M., Rondeau, M.-A., Laroche, R., Poupart, P., Tang, J., Trischler, A., and Hamilton, W · 2020
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., et al · 2022
Later among the works it cites.
Aligning to social norms and values in interactive narratives
Ammanabrolu, P., Jiang, L., Sap, M., Hajizhirzi, H., and Choi, Y · 2022
Later among the works it cites.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al · 2022
Later among the works it cites.
Is power-seeking ai an existential risk?
Carlsmith, J · 2022
Later among the works it cites.
X-risk analysis for ai research
Hendrycks, D. and Mazeika, M · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Later among the works it cites.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Later among the works it cites.
Asking for knowledge (AFK): Training RL agents to query external knowledge using language
Liu, I.-J., Yuan, X., Côté, M.-A., Oudeyer, P.-Y., and Schwing, A · 2022
Later among the works it cites.
A generalist agent
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-maron, G., Giménez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., and Gabriel, I · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Closest in time.
Batch prompting: Efficient inference with large language model apis
Cheng, Z., Kasai, J., and Yu, T · 2023
Closest in time.
Natural selection favors ais over humans, 2023
Hendrycks, D · 2023
Closest in time.
Taskmatrix. ai: Completing tasks by connecting foundation models with millions of apis
Liang, Y., Wu, C., Song, T., Wu, W., Xia, Y., Liu, Y., Ou, Y., Lu, S., Ji, L., Mao, S., et al · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.