Fetching the paper…
Reading the bibliography…
Solving the AI alignment problem requires having clear, defensible values towards which AI systems can align.
Confucian analects: The great learning, and the doctrine of the mean
J. Legge et al · 1971
Earlier work this paper cites.
Plato’s Phaedo , volume 120
Plato, R. Hackforth, et al · 1972
Earlier work this paper cites.
The analects. translated by dc lau, 1979
C. A. Confucius · 1979
Earlier work this paper cites.
Ecology, meaning, and religion
R. A. Rappaport · 1979
Earlier work this paper cites.
The complete works of Aristotle , volume 2
Aristotle, J. Barnes, et al · 1984
Earlier work this paper cites.
The imperative of responsibility: In search of an ethics for the technological age
H. Jonas · 1984
Earlier work this paper cites.
Immanuel kant: Foundations of the metaphysics of morals, 1989
I. Kant and L. W. Beck · 1989
Earlier work this paper cites.
Human universals
D. Brown · 1991
Earlier work this paper cites.
Applications of ai in education
J. Beck, M. Stern, and E. Haugsjaa · 1996
Earlier work this paper cites.
Confucius: the analects
D. C. Lau · 2000
Earlier work this paper cites.
Women and human development: The capabilities approach , volume 3
M. C. Nussbaum · 2000
Earlier work this paper cites.
Young children’s rapid learning about artifacts
K. Casler and D. Kelemen · 2005
Earlier work this paper cites.
The philosophy of international law
S. Besson and J. Tasioulas · 2010
Earlier work this paper cites.
Universal human rights in theory and practice
J. Donnelly · 2013
Earlier work this paper cites.
The summa theologica: Complete edition
S. T. Aquinas et al · 2014
Earlier work this paper cites.
Superintelligence: Paths, dangers, strategies
N. Bostrom · 2014
Earlier work this paper cites.
Aristotle de anima
R. D. Hicks · 2015
Earlier work this paper cites.
New dimensions in testimony: Digitally preserving a holocaust survivor’s interactive storytelling
D. Traum, A. Jones, K. Hays, H. Maio, O. Alexander, R. Artstein, P. Debevec, A. Gainer, K. Georgila, K. Haase, et al · 2015
Earlier work this paper cites.
Moral deskilling and upskilling in a new machine age: Reflections on the ambiguous future of character
S. Vallor · 2015
Earlier work this paper cites.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Naturalizing ethics
O. Flanagan, H. Sarkissian, and D. Wong · 2016
Earlier work this paper cites.
Intergenerational justice
L. H. Meyer · 2017
Earlier work this paper cites.
Deep learning of aftershock patterns following large earthquakes
P. M. DeVries, F. Viégas, M. Wattenberg, and B. J. Meade · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Cited alongside, same era.
Impact of modern technology in education
R. Raja and P. Nagasubramani · 2018
Cited alongside, same era.
Who would destroy the world? omnicidal agents and related phenomena
P. Torres · 2018
Cited alongside, same era.
Ethics in technology practice
S. Vallor, B. Green, and I. Raicu · 2018
Cited alongside, same era.
fastmri: An open dataset and benchmarks for accelerated mri
J. Zbontar, F. Knoll, A. Sriram, T. Murrell, Z. Huang, M. J. Muckley, A. Defazio, R. Stern, P. Johnson, M. Bruno, et al · 2018
Cited alongside, same era.
Artificial intelligence, decision-making, and moral deskilling
Climate change and moral responsibility toward future generations: A confucian perspective
F. Teng · 2021
Later among the works it cites.
Computational ethics
E. Awad, S. Levine, M. Anderson, S. L. Anderson, V. Conitzer, M. Crockett, J. A. Everett, T. Evgeniou, A. Gopnik, J. C. Jamison, et al · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements, 2022
A. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, L. Campbell-Gillingham, J. Uesato, P.-S. Huang, R. Comanescu, F. Yang, A. See, S. Dathathri, R. Greig, C. Chen, D. Fritz, J. S. Elias, R. Green, S. Mokrá, N. Fernando, B. Wu, R. Foley, S. Young, I. Gabriel, W. Isaac, J. Mellor, D. Hassabis, K. Kavukcuoglu, L. A. Hendricks, and G. Irving · 2022
Later among the works it cites.
Artificial intelligence at google: Our principles, 2022
GoogleAI · 2022
Later among the works it cites.
Ai ethics, 2022
IBM · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. P. Green · 2019
Cited alongside, same era.
Ml for flood forecasting at scale, 2019
S. Nevo, V. Anisimov, G. Elidan, R. El-Yaniv, P. Giencke, Y. Gigi, A. Hassidim, Z. Moshe, M. Schlesinger, G. Shalev, A. Tirumali, A. Wiesel, O. Zlydenko, and Y. Matias · 2019
Cited alongside, same era.
Artificial intelligence in education: Challenges and opportunities for sustainable development
F. Pedro, M. Subosa, A. Rivas, and P. Valverde · 2019
Cited alongside, same era.
Artificial intelligence for law enforcement: challenges and opportunities
S. Raaijmakers · 2019
Cited alongside, same era.
Human compatible: Artificial intelligence and the problem of control
S. Russell · 2019
Cited alongside, same era.
Artificial intelligence, values, and alignment
I. Gabriel · 2020
Cited alongside, same era.
Convergences in the ethics of space exploration
B. P. Green · 2020
Cited alongside, same era.
S. Kreps, R. M. McCain, and M. Brundage · 2022
Later among the works it cites.
Our approach, 2022
MicrosoftStaff · 2022
Later among the works it cites.
Ethical use policy, 2022
Salesforce · 2022
Later among the works it cites.
Systematic review of smart health monitoring using deep learning and artificial intelligence
A. Sujith, G. S. Sajja, V. Mahalakshmi, S. Nuhmani, and B. Prasanalakshmi · 2022
Later among the works it cites.
Artificial intelligence, humanistic ethics
J. Tasioulas · 2022
Later among the works it cites.
T. Chakrabarty, V. Padmakumar, F. Brahman, and S. Muresan · 2023
Closest in time.
A multilevel framework for ai governance
H. Choung, P. David, and J. S. Seberger · 2023
Closest in time.
Ethics in the age of disruptive technologies: An operational roadmap
J. R. Flahaux, B. P. Green, and A. G. Skeet · 2023
Closest in time.
Chatgpt and the future of work: A comprehensive analysis of ai’s impact on jobs and employment
A. S. George, A. H. George, and A. G. Martin · 2023
Closest in time.
A multi-level framework for the ai alignment problem
B. L. Hou and B. P. Green · 2023
Closest in time.
Health system-scale language models are all-purpose prediction engines
L. Y. Jiang, X. C. Liu, N. P. Nejatian, M. Nasir-Moin, D. Wang, A. Abidin, K. Eaton, H. A. Riina, I. Laufer, P. Punjabi, et al · 2023
Closest in time.
The empty signifier problem: Towards clearer paradigms for operationalising "alignment" in large language models, 2023
H. R. Kirk, B. Vidgen, P. Röttger, and S. A. Hale · 2023
Closest in time.
Crossbow intruder who wanted to “kill queen” given nine-year sentence, Oct 2023
D. Lee · 2023
Closest in time.
Auditing large language models: a three-layered approach
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi · 2023
Closest in time.
Openai charter, 2023
OpenAI · 2023
Closest in time.
Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties
T. Sorensen, L. Jiang, J. Hwang, S. Levine, V. Pyatkin, P. West, N. Dziri, X. Lu, K. Rao, C. Bhagavatula, et al · 2023
Closest in time.
Algorithmic colonization of love: The ethical challenges of dating app algorithms in the age of ai
H. Wang · 2023
Closest in time.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, et al · 2023
Closest in time.